| @jim:acmegating.com | i believe the zuul jobs are actually affected by this; i will restart the executors on the revert image | 00:33 |
|---|---|---|
| @jim:acmegating.com | #status log restarted all executors on the logreceiver revert patch (in order to allow the logreceiver fix to merge) | 00:52 |
| @status:opendev.org | @jim:acmegating.com: finished logging | 00:52 |
| -@gerrit:opendev.org- Zuul merged on behalf of Takashi Kajinami: [openstack/diskimage-builder] 1001676: Declare Python 3.13 support https://review.opendev.org/c/openstack/diskimage-builder/+/1001676 | 01:09 | |
| -@gerrit:opendev.org- Zuul merged on behalf of Clark Boylan: [opendev/system-config] 988993: Set kernel.yama.ptrace_scope to 2 on executors https://review.opendev.org/c/opendev/system-config/+/988993 | 01:46 | |
| @gthiemonge:matrix.org | Hi Folks, I'm troubleshooting CI failures with the octavia centos 10 stream jobs. It looks like they always fail when they run on the vexxhost provider, the route to the rdoproject repos doesn't work: | 07:39 |
| Connecting to trunk.rdoproject.org (trunk.rdoproject.org)|38.129.56.180|:443... failed: No route to host. | ||
| https://35f57d723fb280f5d9a6-c5982ef1d1780edfccf5471d12724c1b.ssl.cf2.rackcdn.com/openstack/311f30a8482b4729be4addc0a0bb8b41/job-output.txt | ||
| I searched on opensearch, the connection to the repo works from all regions except vexxhost-ca-ymq-1 | ||
| Is it an issue that we can fix here? | ||
| @harbott.osism.tech:regio.chat | looks like trunk.rdoproject.org itself is also hosted by vexxhost, so maybe they have some internal routing issue? mnaser might be able to check or someone can check from our mirror node | 11:51 |
| @harbott.osism.tech:regio.chat | corvus: not sure if you saw that, 1002828,2 seems to also have been stuck in gate until you did the revert on the executors. given that we are very close to the end of the OpenStack release cycle, I would vote to keep zuul in this state for the next four weeks unless any further issues are detected of course. maybe we'd even need to disable the weekly restart cron to achieve that? | 11:57 |
| @jim:acmegating.com | Jens Harbott: i don't think we should freeze zuul for 4 weeks, that's an exceptionally long time. i also think we're mostly finished with the debugging of this particular change. i'm not worried about it having significant impacts. | 13:17 |
| @jim:acmegating.com | @status log restarted all zuul executors on zuul master to pick up latest logreceiver fixes | 13:21 |
| @jim:acmegating.com | #status log restarted all zuul executors on master to pick up latest logreceiver fixes | 13:22 |
| @status:opendev.org | @jim:acmegating.com: finished logging | 13:22 |
| -@gerrit:opendev.org- Stephen Finucane proposed: [openstack/project-config] 988094: Add nova-reviewers to openstack/nova ACL https://review.opendev.org/c/openstack/project-config/+/988094 | 15:22 | |
| @clarkb:matrix.org | https://github.com/orgs/community/discussions/206581 this is the upstream github discussion around requiring authentication. This has affected Gerrit CI. I am not aware of us having hit issues but it wouldn't surprise me if people have qusetions about it | 15:22 |
| @clarkb:matrix.org | gthiemonge: Jens Harbott I can reach trunk.rdoproject.org from mirror.ca-ymq-1.vexxhost.opendev.org. Makes me wonder if the issue is/was temporary or the routing trouble may be more subtle and impact specific hypervisors ? | 15:24 |
| @clarkb:matrix.org | re github `a subset of unauthenticated clone or fetch requests may now be asked to authenticate as part of GitHub’s protections against abusive traffic. If you receive a 401, update your application or script to use GitHub credentials.` | 15:28 |
| @gthiemonge:matrix.org | Clark: in opensearch, if I search "Connecting to trunk.rdoproject.org" and filter with hosts_region = "vexxhost-ca-ymq-1", all the attempts have failied since Aug 26 | 15:28 |
| @clarkb:matrix.org | gthiemonge: are there any successful jobs in that region too? Note that particular log message may only appear when it fails to connect (I'm not sure) | 15:30 |
| @clarkb:matrix.org | just wondering if we can narrow it down further | 15:30 |
| @clarkb:matrix.org | it is also theoretically possible that the firewall(s) (if any) in front of trunk.rdoproject.org are blockign our requests | 15:31 |
| @gthiemonge:matrix.org | Clark: only failures after Aug 26, I see that it worked before ("Connecting to trunk.rdoproject.org (trunk.rdoproject.org)|38.129.56.180|:443... connected") | 15:32 |
| @clarkb:matrix.org | ack thank you for confirming | 15:33 |
| @clarkb:matrix.org | so likely something changes around August 26 either in the cloud itself (routing, networks, neutron, etc) or in the configuration for trunk.rdoproject.org (firewalls, bot mitigation, etc) | 15:34 |
| @clarkb:matrix.org | and that has consistently prevented test nodes launched in the same cloud region from accessing the resources but not random curls issued from mirror.ca-ymq-1.vexxhost.opendev.org which is hosted in the same region | 15:34 |
| -@gerrit:opendev.org- Stephen Finucane proposed: | 15:44 | |
| - [openstack/project-config] 1003838: Initiate retirement of molteniron https://review.opendev.org/c/openstack/project-config/+/1003838 | ||
| - [openstack/project-config] 1003839: Retire molteniron https://review.opendev.org/c/openstack/project-config/+/1003839 | ||
| - [openstack/project-config] 1003840: Remove references to ironic-core https://review.opendev.org/c/openstack/project-config/+/1003840 | ||
| -@gerrit:opendev.org- Stephen Finucane proposed: | 15:45 | |
| - [openstack/project-config] 1003839: Retire molteniron https://review.opendev.org/c/openstack/project-config/+/1003839 | ||
| - [openstack/project-config] 1003840: Remove references to ironic-core https://review.opendev.org/c/openstack/project-config/+/1003840 | ||
| -@gerrit:opendev.org- Stephen Finucane proposed: [openstack/project-config] 1003840: Remove references to ironic-core https://review.opendev.org/c/openstack/project-config/+/1003840 | 15:51 | |
| -@gerrit:opendev.org- Stephen Finucane proposed: [openstack/project-config] 1003840: Remove references to ironic-core https://review.opendev.org/c/openstack/project-config/+/1003840 | 15:55 | |
| -@gerrit:opendev.org- Stephen Finucane proposed: [openstack/project-config] 1003848: Add nova-approvers to openstack/nova ACL https://review.opendev.org/c/openstack/project-config/+/1003848 | 16:00 | |
| @stephenfin:matrix.org | How does one get a new Gerrit group created? Context being https://lists.openstack.org/archives/list/openstack-discuss@lists.openstack.org/thread/EI44ZQG6UHIBNH3O26KHUVHIE576SCOC/ | 16:11 |
| @stephenfin:matrix.org | And additionally, how does one get one *deleted* (assuming such a thing is possible/a good idea). Context for that being https://review.opendev.org/c/openstack/project-config/+/1003840 | 16:12 |
| @clarkb:matrix.org | stephenfin: when the change (988094) merges there is a deployment step to update the acl in gerrit. That step will create any missing groups before applying acls that use the group | 16:12 |
| @stephenfin:matrix.org | TIL | 16:12 |
| @clarkb:matrix.org | the group will be created empty so we will need to add an initial seed member who can fill in everyone else (usually that is given to the PTL of the project) | 16:12 |
| @clarkb:matrix.org | deleting groups is something we haven't done historically because it wasn't well supported. You just stop using the group in the acls and remove everyone from them. I think in like the last gerrit release or two we may actually be able to delete them safely now (it checks they aren't used etc) but we haven't done any work around that yet | 16:13 |
| @stephenfin:matrix.org | Okay, on deletion, I'll tell the ironic folks just to remove themselves | 16:14 |
| @stephenfin:matrix.org | On creation, Uggla is PTL but not in nova-core (thus I had to make the changes I announced earlier today). I suspect the best thing to do is include nova-core? That's what ironic has done in ironic-reviewers https://review.opendev.org/admin/groups/cfda7dc8666aec43e442acd9c7a586c8d2c93895,members | 16:14 |
| @clarkb:matrix.org | ya I think that would work | 16:15 |
| @stephenfin:matrix.org | I will ask him if he can +1 the project-config change though | 16:15 |
| @clarkb:matrix.org | ++ | 16:15 |
| -@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [zuul/zuul-jobs] 1003852: Add buildset registry password to redactions https://review.opendev.org/c/zuul/zuul-jobs/+/1003852 | 16:18 | |
| -@gerrit:opendev.org- Stephen Finucane proposed: [openstack/project-config] 1003839: Retire molteniron https://review.opendev.org/c/openstack/project-config/+/1003839 | 16:23 | |
| -@gerrit:opendev.org- Stephen Finucane proposed: [openstack/project-config] 1003840: Remove references to ironic-core https://review.opendev.org/c/openstack/project-config/+/1003840 | 16:23 | |
| -@gerrit:opendev.org- Zuul merged on behalf of Roja Eswaran: [openstack/diskimage-builder] 999989: debootstrap: add DIB_MMDEBSTRAP_EXTRA_ARGS support https://review.opendev.org/c/openstack/diskimage-builder/+/999989 | 16:51 | |
| @fungicide:matrix.org | stephenfin: an alternative is that if an openstack project declares tact sig liaisons we add those initially instead | 17:27 |
| @fungicide:matrix.org | on the expectations that they're the ones delegated by the ptl or under dpl to handle adding other members | 17:27 |
| @harbott.osism.tech:regio.chat | https://review.opendev.org/c/opendev/zuul-providers/+/1000965 keeps failing in gate with swift upload failures :( | 18:23 |
| @fungicide:matrix.org | i think most zuul-providers changes that end up uploads to all providers are failing on enough random providers that they can never merge. i have one i've been rechecking off and on for months | 18:25 |
| @harbott.osism.tech:regio.chat | maybe we need to force-merge and then have the regular rebuilds roll out the changes? or make all those jobs non-voting in gate? | 18:26 |
| @fungicide:matrix.org | yeah, https://review.opendev.org/c/opendev/zuul-providers/+/989136 was proposed in may and still hasn't merged | 18:26 |
| @harbott.osism.tech:regio.chat | hmm, seems like V-2 doesn't trigger the attention set, so I missed that one | 18:27 |
| @fungicide:matrix.org | both options sound reasonable if we can't figure out why uploads are so unreliable | 18:27 |
| @jim:acmegating.com | in gate, we're only uploading to a single location: the new rax-flex swift in dfw | 18:41 |
| @jim:acmegating.com | we could choose a different location if we think it's more reliable, or we could do some fallback stuff | 18:41 |
| @jim:acmegating.com | i think it's worth keeping the uploads in gate though; it's been very beneficial to have changes roll out immediately | 18:42 |
| @jim:acmegating.com | https://review.opendev.org/c/opendev/zuul-providers/+/989136?tab=change-view-tab-header-zuul-results-summary -- that change had no upload failures, those were all build failures | 18:44 |
| @jim:acmegating.com | for https://review.opendev.org/c/opendev/zuul-providers/+/1000965?tab=change-view-tab-header-zuul-results-summary -- 2 of those were build failures 4 were upload failures | 18:45 |
| @jim:acmegating.com | so here's my proposal for action for the upload failures: | 18:51 |
| 1) if someone wants to talk with the rax flex folks about those 500 errors, that could be beneficial | ||
| 2) if we suspect those errors may be transient (the other successful uploads suggest they may be), someone could propose a change to wrap that in a retry loop. i would be happy to advise and review that | ||
| 3) if someone has a suggestion for a different location to upload to, i would be happy to set up the swift container and write the change to switch to that location | ||
| @jim:acmegating.com | meanwhile, identifying the cause of those build failures (looks like there may be some disk space issues) is another thing that folks could do to improve things | 18:51 |
| @clarkb:matrix.org | fungi: do you still have a line to James? | 18:53 |
| @clarkb:matrix.org | (maybe we should ask if he/they have a preferred escalation method eg using the ticket system or contacting directly etc) | 18:53 |
| @fungicide:matrix.org | yeah, i think we tipped over the reliability threshhold to completely blocked when the disk space errors started | 18:53 |
| @jim:acmegating.com | here's an example swift upload 500 error: https://zuul.opendev.org/t/opendev/build/7ac0e6e2fd0e4350ae708a5895e379c5 | 18:54 |
| @jim:acmegating.com | better link: https://zuul.opendev.org/t/opendev/build/7ac0e6e2fd0e4350ae708a5895e379c5/console#4/0/40/ubuntu-noble | 18:55 |
| @fungicide:matrix.org | i was talking to james denton via irc privmsg but it seems like he only appeared when someone else (doug?) prodded, and then stopped responding again once the immediate problem with the mirror servers in error mode was resolved | 18:55 |
| @jim:acmegating.com | that actually has our container name in it | 18:55 |
| @clarkb:matrix.org | maybe we should file a couple of tickets then? One for the auth thing and another for this swift upload thing | 18:55 |
| @clarkb:matrix.org | not to distract from this conversation, but it has been ~24 hours since I shutdown the old backup02 server. I don't see any complaints in the infra root inbox. We good with deleting that backup02 server now? | 18:56 |
| @fungicide:matrix.org | i'll go ahead and try resetting our api token for the control plane account first and see if that solves the flex keystone auth problem, then open a support ticket if it doesn't | 18:58 |
| @fungicide:matrix.org | mainly because i suspect that's the first thing they'll ask us to do before they look into their end | 18:58 |
| @clarkb:matrix.org | sounds good | 18:59 |
| @fungicide:matrix.org | it will invalidate the auth for that tenant in all classic and flex regions until we get clouds.yaml files updated, fair warning | 18:59 |
| @clarkb:matrix.org | it is the control plane side so shouldn't affect zuul though | 18:59 |
| @clarkb:matrix.org | the impact is likely to be small if any | 19:00 |
| @fungicide:matrix.org | so i want to make sure that if anyone is in the middle of trying to bring up new servers anywhere in our control plane in rackspace i don't step on toes | 19:00 |
| @clarkb:matrix.org | fungi: you didn't see any backup complaints in the root inbox did you? Just to make sure my email filtering foo isn't missing anything obvious there | 19:00 |
| @fungicide:matrix.org | lemme check | 19:00 |
| @fungicide:matrix.org | i'm a bit behind on e-mail this week, it's been busy | 19:01 |
| @clarkb:matrix.org | and I am in the middle of talking to vexxhost not rax so I don't expect any issues with the plan to change the token | 19:01 |
| @fungicide:matrix.org | the ze11 replacement or the backup server removal? | 19:01 |
| @clarkb:matrix.org | backup server removal is vexxhost | 19:02 |
| @clarkb:matrix.org | I haven't looked at ze11 replacement. That would be rax classic if sticking to where the other executors are though | 19:02 |
| @clarkb:matrix.org | (which would be affected) | 19:02 |
| @fungicide:matrix.org | okay yeah ze11 work would get impacted if i'm doing this at the same time | 19:02 |
| @jim:acmegating.com | i'm planning on doing ze11 work right after lunch actually, but don't worry, if it breaks, that's no big deal | 19:03 |
| @fungicide:matrix.org | i see https://review.opendev.org/c/opendev/system-config/+/988993 deploy failed on infra-prod-service-zuul early this morning utc, i'm assuming that's already known | 19:04 |
| @clarkb:matrix.org | nope. But we run the zuul deployment hourly so let me check if it caught up later | 19:05 |
| @clarkb:matrix.org | https://zuul.opendev.org/t/openstack/builds?job_name=infra-prod-service-zuul&skip=0 it did catch up | 19:05 |
| @fungicide:matrix.org | good enough | 19:05 |
| @jim:acmegating.com | i'll look into that anyway; there's a good chance i'm responsible for whatever broke it | 19:06 |
| @fungicide:matrix.org | timing seems like it might have coincided with something getting restarted | 19:06 |
| @clarkb:matrix.org | looks like docker compose pull failed on ze03 | 19:07 |
| @clarkb:matrix.org | `service-zuul.yaml.log.2026-09-03T03:43:34` is the log file on bridge I think | 19:07 |
| @clarkb:matrix.org | corvus: `no space left on device` | 19:08 |
| @clarkb:matrix.org | corvus: do we need to prune docker stuff there after the work you'ev done on it? | 19:08 |
| @jim:acmegating.com | yeah, also, we seem to be emitting a lot of log lines to docker; i think something weird is going on there | 19:08 |
| @clarkb:matrix.org | df says we have a little headroom there right now but not a ton | 19:08 |
| @jim:acmegating.com | are those logs going to the journal? | 19:10 |
| @clarkb:matrix.org | we don't appear to have any special journal logging config in the docker compose file on ze03 | 19:10 |
| @clarkb:matrix.org | so I think anything going to stdout/stderr would be captured by the container runtime and not forwarded further | 19:11 |
| @clarkb:matrix.org | I suspect we did that because we configure zuul to log to files directly and don't expect anything out stdout/stderr | 19:11 |
| @jim:acmegating.com | "LogConfig": { | 19:12 |
| "Type": "journald", | ||
| "Config": null | ||
| }, | ||
| @jim:acmegating.com | that's from docker inspect | 19:12 |
| @clarkb:matrix.org | oh maybe its a default? and without our extra config it just doesn't go to /var/log/containers because it doesn't get the extra tagging. So it goes into the main journal? | 19:12 |
| @jim:acmegating.com | i think so, that's probably why it recovered | 19:13 |
| @jim:acmegating.com | maybe the main thing to do here is to track down why we're getting so many messages to the container log; i conisdered that low priority, but i'll bump that to the top of the list. | 19:13 |
| @clarkb:matrix.org | ya I am surprised we get anything to stdout/stderr since we expect things in /var/log/zuul/ files via the python logging config. But maybe we've missed something in that config or its giong to stdout too? | 19:14 |
| @jim:acmegating.com | (i mean, i know the mechanism, i just haven't collected all the info to propose a solution yet) | 19:14 |
| @clarkb:matrix.org | ack | 19:15 |
| @jim:acmegating.com | i'll do that first after lunch, then ze11 :) | 19:15 |
| @clarkb:matrix.org | ok last call on backup02 deletion | 19:16 |
| @clarkb:matrix.org | has anyone seen any raeson we should delay that server delete? | 19:16 |
| @fungicide:matrix.org | still checking infra-root e-mail | 19:17 |
| @clarkb:matrix.org | got it I'll wait for your report before proceeding | 19:17 |
| @fungicide:matrix.org | clearing out the noise takes time | 19:17 |
| @fungicide:matrix.org | mainly my mua is waiting on gmail's imap to respond to the request to process 22k deletes based on a pattern match | 19:18 |
| @fungicide:matrix.org | that's several days worth | 19:18 |
| @fungicide:matrix.org | usually takes a few minutes | 19:18 |
| @fungicide:matrix.org | with batches this large it frequently times out and eventually has to be retried/continued | 19:20 |
| @fungicide:matrix.org | done. checking the infra-root inbox (after deleting a few thousand backscatter from the weekend) i see no notifications about backup failures. also checked the spam folder since they sometimes land there (after deleting tens of thousands of noise messages) and nothing there either | 19:21 |
| @fungicide:matrix.org | should be safe to proceed! | 19:21 |
| @clarkb:matrix.org | thanks I am proceeding now | 19:24 |
| @clarkb:matrix.org | #status log Deleted backup02.ca-ymq-1.vexxhost.opendev.org as it has been replaced by backup03.ca-ymq-1.vexxhost.opendev.org | 19:31 |
| @status:opendev.org | @clarkb:matrix.org: finished logging | 19:31 |
| @jim:acmegating.com | remote: https://review.opendev.org/c/zuul/zuul/+/1003883 Reduce "Ansible output exceeds max. line size" log entries [NEW] | 21:02 |
| @jim:acmegating.com | Clark: fungi i think that our log config is fine and we need to change zuul ^ | 21:03 |
| @fungicide:matrix.org | neat | 21:08 |
| @clarkb:matrix.org | I guess this is a side effect of the changes then too and not an outstanding issue | 21:16 |
| @jim:acmegating.com | Clark: actually no! | 21:16 |
| @jim:acmegating.com | i really dug into it, and this has been around for a long time | 21:17 |
| @fungicide:matrix.org | just think of the rust we spun unnecessarily all those years | 21:18 |
| @jim:acmegating.com | my only guess as to why we noticed it on ze03: my temporary log files from debugging, more restarts means more jobs and more logs | 21:18 |
| @clarkb:matrix.org | ah got it | 21:23 |
| @jim:acmegating.com | do we expect the "Ubuntu Noble OpenDev 20250110" image in rax-dfw to work? | 21:39 |
| @jim:acmegating.com | i ask because it looks like the launch script is failing at the ssh connection stage | 21:39 |
| @jim:acmegating.com | oh hey there's "Ubuntu 24.04 LTS (Cloud)" | 21:40 |
| @jim:acmegating.com | i thought i searched for that | 21:40 |
| @fungicide:matrix.org | i would expect it to work, but probably for the same reasons you expected it to work | 21:40 |
| @jim:acmegating.com | i'll try that next :) | 21:40 |
| @clarkb:matrix.org | corvus: yes our image works, but you have to tell it to use config drive | 21:41 |
| @jim:acmegating.com | ooh | 21:41 |
| @clarkb:matrix.org | it only works with config drive because rax doesn't do regular metadata service and the image we upload uses normal config drive not their patched thing | 21:41 |
| @jim:acmegating.com | Clark: would you recommend i use ours or try theirs? | 21:41 |
| @jim:acmegating.com | not sure what we did for the other executors | 21:41 |
| @clarkb:matrix.org | * it only works with config drive because rax doesn't do regular metadata service and the image we upload uses normal cloud init not their patched thing | 21:41 |
| @jim:acmegating.com | (but would probably be good to match that) | 21:42 |
| @clarkb:matrix.org | corvus: I think most of our stuff initially was based on our image because they didn't haev a noble image available for some time | 21:42 |
| @clarkb:matrix.org | I agree that matching what the others are is probably best for consistency | 21:42 |
| @clarkb:matrix.org | openstack server show ze01.opendev.org should tell you what the image was | 21:42 |
| @jim:acmegating.com | ours | 21:42 |
| @clarkb:matrix.org | launch node script has a --config-drive flag to enable config drive usage | 21:43 |
| @jim:acmegating.com | yep, running `/usr/launcher-venv/bin/launch-node $FQDN --flavor "$FLAVOR" --cloud=$OS_CLOUD --region=$OS_REGION_NAME --image 8e6a5ba1-0447-4652-bae3-04d016746a9f --config-drive` now | 21:44 |
| @jim:acmegating.com | couple more times i might remember this | 21:44 |
| -@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/zone-opendev.org] 1003887: Add new ze11 https://review.opendev.org/c/opendev/zone-opendev.org/+/1003887 | 21:57 | |
| -@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/system-config] 1003888: Add ze11 https://review.opendev.org/c/opendev/system-config/+/1003888 | 21:58 | |
| -@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/system-config] 1003889: Add ze11 to cacti https://review.opendev.org/c/opendev/system-config/+/1003889 | 21:59 | |
| -@gerrit:opendev.org- Zuul merged on behalf of James E. Blair https://matrix.to/#/@jim:acmegating.com: [opendev/zone-opendev.org] 1003887: Add new ze11 https://review.opendev.org/c/opendev/zone-opendev.org/+/1003887 | 22:15 | |
| -@gerrit:opendev.org- Zuul merged on behalf of James E. Blair https://matrix.to/#/@jim:acmegating.com: | 22:53 | |
| - [opendev/system-config] 1003888: Add ze11 https://review.opendev.org/c/opendev/system-config/+/1003888 | ||
| - [opendev/system-config] 1003889: Add ze11 to cacti https://review.opendev.org/c/opendev/system-config/+/1003889 | ||
| @clarkb:matrix.org | I think ze11 should be deployed now | 23:41 |
Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!