| -@gerrit:opendev.org- OpenStack Proposal Bot proposed: [openstack/project-config] 1003397: Normalize projects.yaml https://review.opendev.org/c/openstack/project-config/+/1003397 | 02:10 | |
| -@gerrit:opendev.org- Zuul merged on behalf of Clark Boylan: [opendev/irc-meetings] 1003393: Switch opendev office hour to biweekly-even https://review.opendev.org/c/opendev/irc-meetings/+/1003393 | 05:17 | |
| -@gerrit:opendev.org- Zuul merged: [openstack/project-config] 1003397: Normalize projects.yaml https://review.opendev.org/c/openstack/project-config/+/1003397 | 13:42 | |
| @jim:acmegating.com | ze03 memory looks good with the latest zuul changes, so i plan on restarting the rest of the executors with that today | 14:57 |
|---|---|---|
| -@gerrit:opendev.org- Monty Taylor https://matrix.to/#/@mordred:inaugust.com proposed: [zuul/zuul-jobs] 1001597: Get artifacts from builds via buildset https://review.opendev.org/c/zuul/zuul-jobs/+/1001597 | 15:00 | |
| @mordred:waterwanders.com | fungi, Clark if either or you are around and have a sec, could you nudge https://review.opendev.org/c/openstack/project-config/+/1002832 over the line for me? | 15:14 |
| @fungicide:matrix.org | sure, lookin' | 15:15 |
| @mordred:waterwanders.com | thanks! | 15:19 |
| @fungicide:matrix.org | any time | 15:19 |
| -@gerrit:opendev.org- Zuul merged on behalf of Monty Taylor https://matrix.to/#/@mordred:inaugust.com: [openstack/project-config] 1002832: Add a few more wandertracks and inaugust repos https://review.opendev.org/c/openstack/project-config/+/1002832 | 15:26 | |
| -@gerrit:opendev.org- Monty Taylor https://matrix.to/#/@mordred:inaugust.com proposed: [zuul/zuul-jobs] 1001597: Get artifacts from builds via buildset https://review.opendev.org/c/zuul/zuul-jobs/+/1001597 | 15:43 | |
| @harbott.osism.tech:regio.chat | also still waiting for reviews https://review.opendev.org/c/openstack/project-config/+/993691 https://review.opendev.org/c/opendev/zuul-providers/+/1000965 | 16:16 |
| @clarkb:matrix.org | both have my +2 now. I think you can probably single core approve them if no one else gets to them soon | 16:19 |
| @clarkb:matrix.org | * Jens Harbott: both have my +2 now. I think you can probably single core approve them if no one else gets to them soon | 16:20 |
| @clarkb:matrix.org | Gerrit's mailing list is reporting that their CI system is broken because Github is rejecting all unauthenticated requests | 16:20 |
| @clarkb:matrix.org | I have no idea if that is affecting us, but if true seems like it would be problematic and something we should be aware of | 16:20 |
| @harbott.osism.tech:regio.chat | the openstack integrated gate seems blocked by a swift change that is 19h old, not sure if it got somehow muddled by a ze restart? | 16:28 |
| @harbott.osism.tech:regio.chat | corvus: I'll wait a bit to have you take a look if possible, else I'd just dequeue it I guess | 16:30 |
| @fungicide:matrix.org | Jens Harbott: hrm, yeah the console log for the remaining build ends at 23:15z | 16:30 |
| @clarkb:matrix.org | The stream loads implying the build is still running | 16:30 |
| @clarkb:matrix.org | But ya the timestamp is much older than our timeouts | 16:30 |
| @fungicide:matrix.org | so not blocked for 19 hours (that's when it was enqueued), but still about 17 | 16:31 |
| -@gerrit:opendev.org- Zuul merged on behalf of Dr. Jens Harbott: [openstack/project-config] 993691: Clean up log-classify jobs https://review.opendev.org/c/openstack/project-config/+/993691 | 16:32 | |
| @jim:acmegating.com | let's just dequeue it; i doubt it's a new issue and i won't have an opportunity to debug it right now. | 16:32 |
| @harbott.osism.tech:regio.chat | the stuck job says ze03, so that would match. let me just take a quick look at the node to see if it is still doing something | 16:34 |
| @fungicide:matrix.org | yeah, if you could somehow force that job to stop then there's a ton of stuff that can merge | 16:34 |
| @fungicide:matrix.org | since it's non-voting and all the voting jobs for that item succeeded already | 16:35 |
| @jim:acmegating.com | oh in that case... let me dequeue it | 16:35 |
| @jim:acmegating.com | i'm ready to restart the rest of the executors, so i'll sequence this properly | 16:35 |
| @fungicide:matrix.org | there's like 11 items in the shared queue that succeeded on all their voting | 16:36 |
| @fungicide:matrix.org | jobs | 16:36 |
| @harbott.osism.tech:regio.chat | oops, sorry, I saw that response too late, dequeued already | 16:38 |
| @jim:acmegating.com | no prob | 16:38 |
| @jim:acmegating.com | #status log restarted zuul executors with latest log receiving memory improvements | 16:39 |
| @status:opendev.org | @jim:acmegating.com: finished logging | 16:39 |
| @harbott.osism.tech:regio.chat | I think we can re-enqueue the swift change into gate then, too? since the dequeue doesn't report anything back to gerrit? | 16:40 |
| @clarkb:matrix.org | yes that seems fine | 16:41 |
| @jim:acmegating.com | we might end up with some erroneous merge conflict errors due to the restart combined with dequeue; just fyi if you see them just recheck/re-enqueue | 16:41 |
| @jim:acmegating.com | looking at our graphs, we could use another executor... we should repair ze11 | 16:47 |
| @jim:acmegating.com | i think that needs to be rebuilt | 16:48 |
| @clarkb:matrix.org | replacing it does seem like the simplest way to address the network connectivity issues | 16:50 |
| @clarkb:matrix.org | fungi: I'm starting to look at backup02 cleanup. I think my idea is to shutdown backup02 (but don't delete it). Then detach its volume and attach it on backup03. Any idea what will happen if we then try to boot backup02 for some reason? Should I comment out the fstab entry for its backup volume mount to avoid failures there? | 17:16 |
| But then also after I attach it to backup03 I will need to add it to lvm as a new separate PV, VG, LV set then mount it to /opt/backups-202010? Is there anything special I need to do to avoid having it interact with the current /opt/backups-202605 lvm volume stuff? | ||
| @clarkb:matrix.org | The main reason for shutting it down and not deleting 02 to start is to avoid any unexpected cascade on delete behaviors with the volume | 17:17 |
| @fungicide:matrix.org | what happens at boot may depend on the fstab options for that device | 17:18 |
| @fungicide:matrix.org | e.g. booting may hang indefinitely waiting for it to be hot-plugged | 17:18 |
| @fungicide:matrix.org | also depends on vintage of the operating system and what init it's using, so hard to say without trying | 17:19 |
| @clarkb:matrix.org | in that case do you think commenting it out before shutting down is a good idea? | 17:20 |
| @clarkb:matrix.org | `/opt/backups-202010 ext4 defaults 0 0` so I guess the value of defaults on that platform is probably what determines the behavior | 17:21 |
| @clarkb:matrix.org | looks like defaults will likely try up to a timeout then fail and drop to recovery shell? | 17:22 |
| @fungicide:matrix.org | it's a good idea, sure | 17:22 |
| @jim:acmegating.com | there are a bunch of errors in the openstack integrated queue likely due to the admin actions we took earlier, so i promoted one of the changes in order to reset the queue | 17:25 |
| @clarkb:matrix.org | fungi: googling goes so far as to suggest unmounting the volume on the running system, running vgchange -an then vgexport against the volume group before moving the disk too | 17:26 |
| @clarkb:matrix.org | maybe that is for running systems? seems to imply that may be the case reading some manpages | 17:27 |
| @harbott.osism.tech:regio.chat | corvus: seems like this triggered more errors like https://zuul.opendev.org/t/openstack/build/bda7dc28d2c44732b5c40a9db2b9c007 ? | 17:29 |
| @jim:acmegating.com | yeah, same error as earlier, so i guess it's not transient; we may have a broken repo on one of the executors (and perhaps broken in a way that zuul doesn't think it should delete/recreate) | 17:30 |
| @jim:acmegating.com | i'll track that down | 17:30 |
| @clarkb:matrix.org | red hat docs imply I should do the extra steps and also shutdown the server before moving things so I'll probably go that route unless there is any feedback to the contrary in the next little bit | 17:35 |
| @clarkb:matrix.org | https://access.redhat.com/solutions/4123 this doc fwiw | 17:39 |
| @jim:acmegating.com | this should fix the issue that let the corrupt glance repo slip through: https://review.opendev.org/c/zuul/zuul/+/1003585 Improve corrupt repo detection [NEW] | 17:39 |
| @jim:acmegating.com | i'm going to stop ze09, delete the repo, then start again | 17:39 |
| @fungicide:matrix.org | Clark: if the server is running, terminate any running processes that `lsof` reports using paths under it if any, `umount ...` the volume, `vgchange -an` to deactivate the volume group, comment out the corresponding line in `/etc/fstab` then i forget the detach subcommand for openstackclient off the top of my head but whatever that is | 17:40 |
| @clarkb:matrix.org | fungi: ya I can figure out the openstackclient commands | 17:41 |
| @clarkb:matrix.org | and that matches the red hat doc I linked above | 17:41 |
| @fungicide:matrix.org | generally, if you forget a step, most of the subsequent steps will fail | 17:41 |
| @jim:acmegating.com | #status log deleted corrupt glance repo on ze09 and restarted | 17:41 |
| @status:opendev.org | @jim:acmegating.com: finished logging | 17:41 |
| @fungicide:matrix.org | including nova refusing to detach the cinder volume if lvm on the guest is still using it | 17:42 |
| @clarkb:matrix.org | fungi: and fail safely without data corruption right? | 17:42 |
| @fungicide:matrix.org | yes | 17:42 |
| @fungicide:matrix.org | basically those commands will refuse to take any action | 17:42 |
| @clarkb:matrix.org | cool I'll get started on this following your process above as well as the red hat doc (they align for the most part on the shutdown/removal side of things) | 17:42 |
| @fungicide:matrix.org | also worth noting, i think in the past i've tried detaching a volume from a shutoff guest and the cloud refuses | 17:43 |
| @fungicide:matrix.org | so while shutting down the vm seems like an easy way to avoid extra steps, it's not | 17:43 |
| @clarkb:matrix.org | oh interesting you think I should try to detach after the deactivation steps are completed and before shutting it down then? | 17:44 |
| @jim:acmegating.com | okay, i repromoted that poor cinder change to reset the queue again | 17:44 |
| @clarkb:matrix.org | `sudo lsof | grep backups` has no results so now I will attempt to umount | 17:44 |
| @fungicide:matrix.org | Clark: yeah, detach before stopping the vm | 17:46 |
| @fungicide:matrix.org | it basically signals to the guest kernel to "eject" that hotpluggable/removable volume | 17:47 |
| @clarkb:matrix.org | its interesting that the `vgs` output doesn't seem to chagne before and after `sudo vgchange -an main-202010` but that command did report `0 logical volume(s) in volume group "main-202010" now active` | 17:47 |
| @clarkb:matrix.org | fungi: ok I have umounted, commented out fstab, run vgchange -an and vgexport I should be ready to detach via the openstackclient now? No pv or lv commands to run first? | 17:49 |
| @clarkb:matrix.org | the openstackclient command is `server remove volume serverid volumeid` fwiw | 17:50 |
| @fungicide:matrix.org | correct | 17:50 |
| @fungicide:matrix.org | no other commands | 17:50 |
| @clarkb:matrix.org | thanks running that server remove volume command next then | 17:50 |
| @fungicide:matrix.org | basically the lv activity is being stopped by umount, the vg activity is stopped by vgchange, and there will be no pv activity because a pv is really just a label on a block device not really any sort of translation layer | 17:51 |
| @clarkb:matrix.org | the removal worked. I will proceed with attaching it to 03 and then doing a pvscan, lvmdevices --adddev, vgimport, vgchange -ay, create a mount mount add it to fstab them mount -a | 17:53 |
| @fungicide:matrix.org | awesome | 17:53 |
| @fungicide:matrix.org | note that most of those steps probably happen automagically on newer distros due to udev hotplugging events | 17:54 |
| @fungicide:matrix.org | after you see it attached successfully (check dmesg for report of the new block device appearing), udevd will likely fire off a vgscan and then active any volume groups it discovers | 17:55 |
| @clarkb:matrix.org | yes `pvs` and `vgs` shows it | 17:56 |
| @fungicide:matrix.org | i would proceed to just check `lvs` and if it reports existence of the logical volume then mount it | 17:56 |
| @clarkb:matrix.org | fungi: `lvs` does not show it possibly because I did vgexport so need to explciitly vgimport it? | 17:56 |
| @clarkb:matrix.org | and maybe vgchange -ay ? | 17:56 |
| @fungicide:matrix.org | if vgs indicates the vg isn't active then you want `vgchange -ay` | 17:57 |
| @fungicide:matrix.org | i don't usually bother with export and import | 17:57 |
| @clarkb:matrix.org | ya attr says exported so I suspect I need to go through these steps to make it imported | 17:58 |
| @fungicide:matrix.org | that makes sense. i'm not even clear on what export/import do since i haven't relied on them. maybe it helps when moving between significantly different driver versions or something | 17:59 |
| @clarkb:matrix.org | fungi: the docs i found indicate its to avoid doing too much automatically | 17:59 |
| @clarkb:matrix.org | basically forces manual intervention before proceeding past important points? | 18:00 |
| @fungicide:matrix.org | aha, from the manpage it looks like `vgexport` is for when you don't want it to get activated automatically | 18:00 |
| @clarkb:matrix.org | I have imported then activated it. `lvs` show it with one attr difference to the other so now i'm looking into that | 18:01 |
| @fungicide:matrix.org | also export/import clears and then resets the system id, though i've never found that to be a problem | 18:01 |
| @clarkb:matrix.org | its bit 6. The old volume says `o` and the new one says `-`. `o` apparently means `device open` | 18:02 |
| @clarkb:matrix.org | anyway it shows up in lsblk. The old uuid is still there. I think I'm ready to create the mount point, edit fstab and mount -a | 18:03 |
| @fungicide:matrix.org | if it's not mounted i think that will be the difference | 18:03 |
| @clarkb:matrix.org | aha | 18:03 |
| @fungicide:matrix.org | once it's mounted `o` will indicate it's open | 18:03 |
| @clarkb:matrix.org | Any concern with reusing the old fstab entry? it was uuid based so device paths shouldn't matter and are we good with defaults? | 18:04 |
| @fungicide:matrix.org | no concern at all as long as it's correct | 18:04 |
| @clarkb:matrix.org | ha | 18:05 |
| @clarkb:matrix.org | the uuid appears to match. its still ext4 and I mkdir'd the old mount point in /opt | 18:06 |
| @clarkb:matrix.org | so I think the main question is if defaults are sufficient/appropriate here | 18:06 |
| @fungicide:matrix.org | the uuid is part of the label written to the device, so yes should match | 18:08 |
| @fungicide:matrix.org | i don't see any reason not to stick with defaults, unless we maybe want to mount it read-only? | 18:09 |
| @clarkb:matrix.org | yup it is mounted now. `mount -a` warned that systemd is out of sync and suggseted a `systemctl daemon-reload` too so I did that | 18:09 |
| @clarkb:matrix.org | fungi: I thought about mounting it ro but I think its ok? | 18:09 |
| @clarkb:matrix.org | my main concern is that if we need to mount backups out of there then ro may make tools unhappy | 18:09 |
| @clarkb:matrix.org | its possible they use backup dir local caches etc | 18:09 |
| @clarkb:matrix.org | anyway its mounted now and systemd is reloaded. I'm going to shutdown the old server momentarily. THen we can delete the old server if backups continue to work over the next day or so | 18:10 |
| @fungicide:matrix.org | yeah, makes sense. i have no real preference there | 18:10 |
| @clarkb:matrix.org | #status log Moved backup02.ca-ymq-1.vexxhost.opendev.org's data volume to backup03.ca-ymq-1.vexxhost.opendev.org and mounted it there. backup02 is now in a SHUTOFF state and can be deleted if backups look happy for ~24 hours. | 18:12 |
| @status:opendev.org | @clarkb:matrix.org: finished logging | 18:12 |
| @fungicide:matrix.org | thanks! | 18:13 |
| @clarkb:matrix.org | and thank you for guiding me through that | 18:14 |
| @fungicide:matrix.org | of course | 18:17 |
| @fungicide:matrix.org | my pleasure as always | 18:17 |
| @clarkb:matrix.org | oh and `lvs` attr bits match now after mounting | 18:18 |
| @fungicide:matrix.org | perfect | 18:24 |
| @jim:acmegating.com | i don't see any new error results in gate; so i think that's fixed. if anyone asks about old error results: just recheck | 18:24 |
| @clarkb:matrix.org | corvus: ack thanks for fixing it. | 18:24 |
| @jim:acmegating.com | full service here: break and fix | 18:24 |
| @mordred:waterwanders.com | I've got a few multi-hour buildsets. I'm assuming that's just fallout from the earlier fun., but they all seem to be in a waiting for nodes holding pattern and I dont' see any node requests. should I just rekick those jobs because stuck across restarts? Or should I be patient because the system is catching up (I honestly can't tell) | 19:25 |
| @clarkb:matrix.org | They look similar to the swift change that started the earlier debugging | 19:33 |
| @mordred:waterwanders.com | I'm leaning towards stuck. I did an experiment and rekicked 1 of them, and it now shows node requests like I'd expect | 19:33 |
| @clarkb:matrix.org | Jobs stuck with streaming logs available but going nowhere. I think restarting them is appropriate. You can do that with new patch sets or an admin can kick them out and reenque ie | 19:33 |
| @clarkb:matrix.org | * Jobs stuck with streaming logs available but going nowhere. I think restarting them is appropriate. You can do that with new patch sets or an admin can kick them out and reenqueue | 19:34 |
| @mordred:waterwanders.com | -A/+A works | 19:34 |
| @clarkb:matrix.org | Ah cool that's easy | 19:34 |
| @jim:acmegating.com | seems like there's probably a cleanup bug somewhere, but that can probably wait for a better time | 19:50 |
| @mordred:waterwanders.com | hrm. I restarted all four, and all four are back to waiting on the same jobs they were waiting on before, and I once again don't see any node requests | 20:25 |
| @jim:acmegating.com | hrm lemme check | 20:28 |
| @jim:acmegating.com | i see waiting on opendev-buildset-registry | 20:29 |
| @mordred:waterwanders.com | yeah - buildset-registry is paused, so they _should_ be good to go at this point? | 20:30 |
| @jim:acmegating.com | but it doesn't say it's paused | 20:31 |
| @jim:acmegating.com | maybe that's the issue | 20:31 |
| @mordred:waterwanders.com | oh, yeah. good point. So like - the job content is paused, the job isn't paused. thus sadness | 20:31 |
| @jim:acmegating.com | mordred: i wonder if what is unique here is that the playbook does nothing other than pause the job | 20:42 |
| @jim:acmegating.com | 2026-09-02 19:37:22,486 DEBUG zuul.AnsibleJob: [e: c361034ba2ad4cb28d292940b2aff402] [build: 2e822bf119204081b5871d031b29df85] Ansible complete, result RESULT_NORMAL code 0 | 20:43 |
| @jim:acmegating.com | i see that log entry, which means the ansible process has terminated | 20:43 |
| @jim:acmegating.com | i don't see the "END" entry | 20:44 |
| @jim:acmegating.com | which should show up in the console output | 20:44 |
| @jim:acmegating.com | between those two, we release semaphores | 20:44 |
| @jim:acmegating.com | does this job have any semaphores? | 20:44 |
| @jim:acmegating.com | if not, then the next thing we do is a sync point with the log receiver. that's new code and could conceivably be broken by a weird playbook. | 20:46 |
| @jim:acmegating.com | i'll work on a test to try to reproduce this. in the mean time, i don't have a workable suggestion to avoid the issue | 20:46 |
| @mordred:waterwanders.com | it _shouldn't_ have any semaphores | 20:47 |
| @mordred:waterwanders.com | and that's just the normal opendev-buildset-registry job | 20:47 |
| @mordred:waterwanders.com | so - yeah, my hunch would start to be something about how taht's interacting with the new log receiver? | 20:48 |
| @jim:acmegating.com | yeah, i think zuul's registry job does more, so that's why we didn't see it there | 20:48 |
| @mordred:waterwanders.com | ah - nod | 20:48 |
| @mordred:waterwanders.com | oh - cause you're not using a separate buidlset-registry on zuul at all | 20:49 |
| @mordred:waterwanders.com | I seem to be the largest user of raw opendev-buildset-registry - but dib and system-config might trip over this too. not urgent on my end, I'll be fine :) | 20:53 |
| @fungicide:matrix.org | i'm not where i can troubleshoot at the moment but i think 1002828,2 in the openstack tenant gate pipeline may have hit that condition, it's waitint to report for a build that has its last console log line timestamped 19:18z | 21:25 |
| @jim:acmegating.com | oh that's interesting; i don't think that job pauses | 21:30 |
| @mordred:waterwanders.com | fungi: that one looks a little different, although I can't 100% say for certain. it's stuck on switf-ulload-image in a non-voting job. It *is* sort of stuck at around the end of the buildset registry steps though | 21:30 |
| @jim:acmegating.com | okay i found the bug | 21:32 |
| @jim:acmegating.com | it's kind of funny because Clark asked "what if the secret is an integer" and i was like 'that makes no sense, it will fail no matter what so it's fine if it fails early". | 21:34 |
| @jim:acmegating.com | it looks like we have secrets that are integers somehow | 21:34 |
| @clarkb:matrix.org | oh hey my comment opinted out a thing | 21:34 |
| @jim:acmegating.com | and it failed and it took out the management thread which is why it sever syncs | 21:34 |
| @jim:acmegating.com | * and it failed and it took out the management thread which is why it never syncs | 21:34 |
| @clarkb:matrix.org | I'm in the middle of putting clothes in a suitcase to make sure I don't need to go shopping tomorrow, but I'll do my best to review a fix when there is one | 21:35 |
| @jim:acmegating.com | so we have to figure out why we have integers (maybe we're passing in secret indexes instead of dereferencing them in some cases) | 21:35 |
| @jim:acmegating.com | we could make the management thread robust against this, but -- this has helped us catch a case where we almost certainly would have failed to actually redact secrets, so... | 21:35 |
| @mordred:waterwanders.com | oh - *fascinating* | 21:36 |
| @jim:acmegating.com | do you use oidc tokens in this job? | 21:37 |
| @mordred:waterwanders.com | no, not in this one | 21:38 |
| @mordred:waterwanders.com | the only secret involved in this for me is the one from base-jobs | 21:39 |
| @mordred:waterwanders.com | but | 21:39 |
| @mordred:waterwanders.com | ``` | 21:39 |
| - secret: | ||
| name: opendev-intermediate-registry | ||
| data: | ||
| host: insecure-ci-registry.opendev.org | ||
| port: 5000 | ||
| username: zuul | ||
| ``` | ||
| @mordred:waterwanders.com | there's an integer | 21:39 |
| @mordred:waterwanders.com | (obviously that secret data is not actually secret, but is a convenient use of the secret payload in this case) | 21:41 |
| @jim:acmegating.com | we should only send over the ones that were encrypted... but maybe there's return data? | 21:42 |
| @mordred:waterwanders.com | well good - I was actually going to suggest "maybe we shouldn't attempt to redact things that weren't encrypted" :) | 21:43 |
| @jim:acmegating.com | yeah, anticipated that case :) | 21:43 |
| @mordred:waterwanders.com | do you mean return data as in in zuul_return? or something else. I dont see anything related to this job in a zuul_return other than pause. | 21:44 |
| @jim:acmegating.com | yeah that's what i meant | 21:45 |
| @mordred:waterwanders.com | fwiw - I tried reworking the wandertracks buildset registry setup to work like zuul's, not running a dedicated one, but instead just using the image build job as the buildset registry. It's (I think now given you found the integer bug) broken in the same way. I'll keep it- honestly, it's a more efficient model anyway. but just wanted to report that path | 21:48 |
| @jim:acmegating.com | okay here's another clue; we got the PRE-RUN END but no RUN START; so that means things were working after the pre-run but broke with the start of the run, so it should be the set of secrets it sent over for that playbook | 21:49 |
| @mordred:waterwanders.com | (there was supposed to be an "expectedly" in that parenthetical) | 21:49 |
| @mordred:waterwanders.com | corvus: I see run playbooks running. https://zuul.opendev.org/t/opendev/stream/42213ed6bb3e44dbada07be6bfad2de0?logfile=console.log ... that's from playbooks/container-image/run.yaml in opendev/base-jobs | 21:55 |
| @mordred:waterwanders.com | but - it sure does get to the end of the first run play: https://opendev.org/opendev/base-jobs/src/branch/master/playbooks/container-image/run.yaml line 3 - and we never see any starting lines for the second play that runs the pause. so, that's nice and weird | 21:57 |
| @jim:acmegating.com | sorry i meant specifically the message "RUN START" -- that goes through the same management thread that died | 21:57 |
| @jim:acmegating.com | so it tells us when it died | 21:57 |
| @mordred:waterwanders.com | ah - NOD | 21:57 |
| @mordred:waterwanders.com | yeah | 21:57 |
| @jim:acmegating.com | i repled and got the secrets, the integer is 5000 | 21:57 |
| @mordred:waterwanders.com | well, that secret from earlier would definitely be involved. heh. yeah | 21:57 |
| @mordred:waterwanders.com | \o/ | 21:57 |
| @jim:acmegating.com | the buildset registry role does a zuul return with the buildset registry connection data | 22:20 |
| @jim:acmegating.com | we don't apply the same filtering to zuul-return secrets as we do normal ones | 22:20 |
| @jim:acmegating.com | (that's the bug -- we should do that) | 22:20 |
| @jim:acmegating.com | so it was a slightly different port 5000 from a secret | 22:21 |
| @jim:acmegating.com | er rather, secret return data | 22:21 |
| @mordred:waterwanders.com | ah - nod. I was trying to walk through and we where we were missing it, and I had not yet found the place | 22:21 |
| @mordred:waterwanders.com | ++ | 22:21 |
| @mordred:waterwanders.com | it just so happens that the buildset registry also creates a thing that runs on oport 5000 :) | 22:22 |
| @jim:acmegating.com | oh well.. actually... | 22:26 |
| @jim:acmegating.com | there's no such thing as unencrypted secret return data | 22:26 |
| @jim:acmegating.com | i think we may need to just exempt the secret return data from redaction... there is a facility where a user can opt into that explicitly, so if they want to do that, then they can, say, return a credential as secret return data, and then also return it as redactable data | 22:28 |
| @jim:acmegating.com | Clark: mordred fungi remote: https://review.opendev.org/c/zuul/zuul/+/1003619 Ensure string secrets for redaction in log receiver [NEW] | 22:36 |
| that's what we need to do in zuul to fix the immediate problem | ||
| @jim:acmegating.com | the issue of what to do with unimportant secret return data is a lower-priority zuul design issue | 22:36 |
| @mordred:waterwanders.com | corvus: lgtm | 22:46 |
Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!