Tuesday, 2026-09-29

-@gerrit:opendev.org- Zuul merged on behalf of Steve Baker: [openstack/diskimage-builder] 983813: Refactor 02-set-machine-id into an element https://review.opendev.org/c/openstack/diskimage-builder/+/98381300:19
-@gerrit:opendev.org- Stephen Finucane proposed:14:17
- [openstack/project-config] 1007878: gerrit: Sync ACLs across all oslo deliverables https://review.opendev.org/c/openstack/project-config/+/1007878
- [openstack/project-config] 1007879: gerrit: Remove Backport-Candidate label from oslo projects https://review.opendev.org/c/openstack/project-config/+/1007879
@fungicide:matrix.orgmmm, seems like static.o.o is having trouble serving content to me suddenly15:00
@fungicide:matrix.org`load average: 2201.01, 2232.10, 1393.90`15:00
@fungicide:matrix.orgthat seems mildly high15:01
@clarkb:matrix.orgcorvus: ^ will that impact zuul release?15:01
@clarkb:matrix.orgfungi: only slightly above normal :)15:01
@fungicide:matrix.orgi finally got content, but it took 30 seconds or more to load15:02
@fungicide:matrix.orgbasically all worker slots are in "sending reply" (W) state15:04
@clarkb:matrix.orgthe server is responsive so ya I don't think we're spinning the CPU and instead must be IO?15:04
@fungicide:matrix.orgunsurprisingly, the vast majority of the requests currently being handled by workers are for docs.openstack.org15:05
@clarkb:matrix.orgbased on some non scientific log tailing docs.openstack.org is the majority ya that15:05
@fungicide:matrix.orgi was looking at server-status instead of logs, but good correlation then15:06
@clarkb:matrix.orglooks like significant numbers of 301s confusing the crawlers15:06
@clarkb:matrix.orgso I think this is "the content with its redirect rules make this a particularly problematic set of content for crawlers" again15:07
@clarkb:matrix.orgload is falling though15:07
@jim:acmegating.com> <@clarkb:matrix.org> corvus: ^ will that impact zuul release?15:08
i don't think so
@clarkb:matrix.orgack15:08
@clarkb:matrix.orgI wonder if the crawlers are self moderating on the infinite recursion these days15:08
@clarkb:matrix.orgbecause if load is the metric they seem to be backing off on their own15:09
@clarkb:matrix.orgfungi: I wonder if this would make a good PTG topic for openstack15:09
@fungicide:matrix.orgit would have to be a tc topic, there is no team managing that stuff15:09
@clarkb:matrix.orgessentially the redirect rules are too generic/not specific enough and lead to crawlers redirecting themselves in a frenzy15:10
@clarkb:matrix.orgwhich makes the content far more prone to these issues than other content15:10
@clarkb:matrix.orgbut I don't even know what would break if we start to try and make the rules less generic15:10
@clarkb:matrix.orgif you grep for `" 301 " ` there should be plenty of examples in the docs.openstack.org access log15:13
-@gerrit:opendev.org- Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org marked as active: [openstack/project-config] 1007491: Temporarily remove release docs semaphores https://review.opendev.org/c/openstack/project-config/+/100749115:14
@clarkb:matrix.org`/security-guide/content/api-endpoints/dashboard/compute/dashboard/introduction/introduction/networking/image-storage/identity/compute.html` its like the crawlers are just doing matrix multiplication on valid file path tokens and hoping they find secret data15:14
@fungicide:matrix.orgyes, we observed that behavior, like, last year even15:14
@clarkb:matrix.orgyup definitely not new15:15
@fungicide:matrix.organd then that ends up redirecting to just '/security-guide/'15:15
@fungicide:matrix.organd the bots aren't smart enough to realize that the content they were redirected to is something they've already fetched15:16
@clarkb:matrix.organd that everything they try to fetch at /security-guide/random/cross/product/multiplication/of/token/terms will lead right back there too15:17
@clarkb:matrix.orgexcept load is back to what I would consider normal so maybe they do realize that now after an initial burst of sillyness15:17
@clarkb:matrix.orgI do still see 301s that looks like crawlers but rather than being all of the requests they are now a subset of requests15:20
@clarkb:matrix.orgFollowing up on the zuul db issue from yesterday: disk consumption appears to be stable. No crazy new growth. I think it was corvus who mentioned we have been right on the edge for some time and we simply tipped over15:34
@fungicide:matrix.orgmakes sense, thanks for circling back around to check in on it15:35
-@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/base-jobs] 1007900: Don't check for wheel in release jobs https://review.opendev.org/c/opendev/base-jobs/+/100790015:48
@jim:acmegating.comClark: fungi ^ that should address the issue in zuul's release job15:48
@clarkb:matrix.orgI went ahead and approved that. The var name is correct according to git grep and I expect we'll be able to check this works shortly15:51
-@gerrit:opendev.org- Zuul merged on behalf of James E. Blair https://matrix.to/#/@jim:acmegating.com: [opendev/base-jobs] 1007900: Don't check for wheel in release jobs https://review.opendev.org/c/opendev/base-jobs/+/100790016:02
-@gerrit:opendev.org- Zuul merged on behalf of Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org: [openstack/project-config] 1007491: Temporarily remove release docs semaphores https://review.opendev.org/c/openstack/project-config/+/100749116:37
-@gerrit:opendev.org- Clark Boylan proposed: [openstack/project-config] 1007930: Add a prometheus node exporter dashboard to grafana https://review.opendev.org/c/openstack/project-config/+/100793016:53
@clarkb:matrix.orghow do we feel about adapting existing grafana dashboards on their public dashboards listing/library to our needs? I've done that in ^ to see how it goes16:54
@clarkb:matrix.orgwill be interesting to see if I properly edited the template to not be a template16:54
@jim:acmegating.comit's an unreviewable change :)16:55
@jim:acmegating.comhttps://review.opendev.org/c/openstack/project-config/+/1007930/1/grafana/prometheus-node-exporter.json16:55
@clarkb:matrix.orgya I think the idea ianw had in the past was that we could hold a node and review the results? I agree 15k lines of json is not human friendly16:55
@clarkb:matrix.orgthough I don't know that anything we write would be signficantly smaller?16:55
@clarkb:matrix.orgnode exporter has a lot of data. Maybe the way to make it reviewable is to add panels in iteratiev changes?16:56
@jim:acmegating.comeh, we can look at the results16:56
@mnasiadka:matrix.orgWell, screenshots of panels would not help a lot16:58
@clarkb:matrix.orgmnasiadka: ya I think we would need to hold a node once we think we're at that point16:58
@clarkb:matrix.orgalso chances are that I got the untemplating process wrong and this will need a few edits before it is even usable16:59
@mnasiadka:matrix.orgWell, there's Grafana Foundation SDK (Python), but for OpenDev needs sounds like an overkill17:03
@clarkb:matrix.orgI think zuul missed the event for that change and hasn't enqueued it17:04
@clarkb:matrix.orgI was going to bring this up in the pre ptg today but maybe we should drop the AAAA dns record for review?17:04
@mnasiadka:matrix.orgOther option is just exposing a short guide/script how to start grafana in a container and using the in-git configuration for it - we don't do prometheus authentication - so that should work?17:05
@clarkb:matrix.orgmnasiadka: yes you can run the container locally and then run grafyaml against it17:05
@fungicide:matrix.orgif we drop the aaaa for review, should we also do it for mirror.ca-ymq-1.vexxhost?17:06
@clarkb:matrix.orgit is fairly straightforward compared to many of the other systems as it is basically stateless other than the grafyaml application17:06
@clarkb:matrix.orgfungi: possibly, but mirror.ca-ymq-1.vexxhost will still try to proxy via ipv6 so it may not be as clear cut there17:06
@clarkb:matrix.orgI'm more worried about the openstack release losing half of its tag created events tomorrow17:06
@clarkb:matrix.orghrm the project-config change enqueued before my recheck I think. As zuul says it is 5 minutes old and my recheck is more like 3 minutes ago17:07
@clarkb:matrix.orgbut that is still ~10 minutes after I pushed the change. Maybe queues were backed up (Zuul has been busy)17:07
@jim:acmegating.com14148306502e4b0eba21aa8c72c0fee1 is the event for your patchset-created17:10
@clarkb:matrix.orgack thanks. Looks like zuul saw the event without issues but then there was an 8.5 ish minute delay in processing it on zuul0117:11
@clarkb:matrix.orgwhich lines up with my rough numbers above17:11
@jim:acmegating.comit enqueued the change about 2 seconds after it received it17:11
@fungicide:matrix.orglikely the result of 1007491 merging before it17:12
@clarkb:matrix.orgcorvus: it didn't render it in the UI that quickly. I agree it seems to say check needs to handle this very quickly17:12
@clarkb:matrix.org`2026-09-29 17:02:11,991 DEBUG zuul.Scheduler: [e: 14148306502e4b0eba21aa8c72c0fee1] Processing trigger event <GerritTriggerEvent patchset-created opendev.org/openstack/project-config master 1007930,1>` prior to this log entry there is a >8 minute lack of entries with that event id17:13
@clarkb:matrix.orgoh processing trigger event is where it actually enqueues it to what the UI sees?17:14
@clarkb:matrix.orgbefore that its in the queue to be processed?17:14
@jim:acmegating.comthe ui is a snapshot, only updated after pipeline processing17:14
@jim:acmegating.comit's always going to be delayed17:14
@clarkb:matrix.orgya so its in the queue but zuul is busy and once it does Processing trigger event it enters a state where it is directly visible in the UI?17:15
@jim:acmegating.comi guess i'm trying to say that "visible in the ui" does not correspond to a log entry17:15
@jim:acmegating.com(i mean, we could come up with one if that was important, but it's like 3 degrees of separation from what we have now)17:15
@clarkb:matrix.orggot it17:16
@jim:acmegating.comthe event arrived at the pipeline 2 seconds after it was received; looks like the pipeline was busy for 8.5 minutes before it got to it17:16
@clarkb:matrix.orgIIRC we changed what the queue values mean on the dashboard at some point. When I checked the two numbers were 0/0 so I thought zuul was cuaght up. Thinking out loud here I wonder if a "timestamp of last trigger event processed" value would help there?17:17
@jim:acmegating.comdo we need to know what it was doing?  or are we okay with "zuul didn't miss the event, it was busy"17:17
@clarkb:matrix.orgI think I'm ok with zuul didn't miss the event it was busy. Just thinking out loud now how to potentially expose the "it is busy be a little patient" in the web ui17:18
@jim:acmegating.comClark: the pipeline-level event queues are on the pipeline page: https://zuul.opendev.org/t/openstack/status/pipeline/check17:18
@jim:acmegating.comthey were deemed too noisy for the tenant overview page17:18
@jim:acmegating.com(so the tenant overview only shows the tenant queues)17:19
@jim:acmegating.comi'd probably be okay putting them back if you think it's important17:19
@clarkb:matrix.orgcorvus: the three queue values top right being relevant here?17:20
@jim:acmegating.comyep17:20
@clarkb:matrix.orgI'm not sure I event knew that there were different pipeline specific pages like that (I mean I may have reviewed the change at some point it just never registered in my brain what that means for me as a zuul user)17:20
@clarkb:matrix.orgI don't know that it needs to go on the front page for status.17:21
@clarkb:matrix.orghttps://3d0ae5a1f74faf2e3f94-8cc229e52defecc7a0c44a815309c1d1.ssl.cf5.rackcdn.com/openstack/d53a439f80e74e61be6196c6c39fb8ad/screenshots/node-exporter-full.png neat it seems to work. I guess I should work on holding a node then17:30
@clarkb:matrix.orgunfortunately the disk space graph didn't render. But I don't think we need everything working upfront for this to be valuable17:31
@mnasiadka:matrix.orgClark: wonder if we can uncollapse all the sections somehow17:31
@mnasiadka:matrix.org(unless they are empty and I miss the point) :)17:32
@clarkb:matrix.orgmnasiadka: those panels have a collapsed: true attribute we could change17:32
@clarkb:matrix.orgno they are collapsed by default in the config but it appears to eb editable17:32
@mnasiadka:matrix.orgI think for user experience we might want them to be collapsed, but for the screenshot - it would be nice to have them all visible17:32
@clarkb:matrix.orgnext steps are probably: land the minimally viable thing, fix any issues like with disk consumption, then tune to what we want by uncollapsing etc17:32
@fungicide:matrix.orgwe're using selenium for the screenshots, doesn't it have the ability to "click on" certain elements in the dom?17:33
@clarkb:matrix.orgI'll work on holding a node now so that we can decide if this is minimally viable17:33
@clarkb:matrix.orgfungi: it does17:33
@clarkb:matrix.orgbut the screenshot sysstem is very generic. It lists all the dashboards from the grafana api then opens each one. It doesn't have per dashboard behaviors right now17:33
@fungicide:matrix.orgah, so it's more our plumbing that would need to be extended17:34
@fungicide:matrix.orgwhich, yes, is more work17:34
-@gerrit:opendev.org- Clark Boylan proposed: [opendev/system-config] 991512: DNM intentional grafana failure to hold a node https://review.opendev.org/c/opendev/system-config/+/99151217:35
@clarkb:matrix.orgok ^ hold request is in place for that now. zuul02 has pulled the zuul-client image as part of that process (yesterday I did it on zuul01 and pulled hte image there)17:36
@fungicide:matrix.org5 minutes to the opendevent!17:55
@mnasiadka:matrix.orgComing back to the AAAA entries - I think we might need to tune gai.conf on the mirror to prefer ipv4 until ipv6 gets fixed17:55
@mnasiadka:matrix.orgfungi: are you going to devent?17:55
@clarkb:matrix.orgmnasiadka: ya that would solve the proxying problem on the backend17:56
@clarkb:matrix.orgbut lets talk about it in a few minuets as this is topical17:56
@mnasiadka:matrix.orgI won't join today, have some other plans - but will be there tomorrow ;-)17:57
@clarkb:matrix.orgcool I know today is a bit later for you too. See you tomorrow17:57
@fungicide:matrix.orgmnasiadka: i didn't even know it was a thing17:57
@mnasiadka:matrix.orgI was just laughing, opendevent reminds me of venting (negative emotions stuff)17:58
@mnasiadka:matrix.orgBut it seems it's a thing, removing trapped air from closed system ;-)17:58
@fungicide:matrix.orgoh that too, yep!17:59
@clarkb:matrix.orgwe can vent and event17:59
@clarkb:matrix.orgits time. I made it into the meetpad first so I am moderator!18:00
@clarkb:matrix.orgheld grafana node with node exporter dashboard here: https://217.182.140.216/d/49af61ae98/node-exporter-full?orgId=1&from=now-24h&to=now&timezone=utc&var-job=node&var-nodename=prometheus01&var-node=prometheus01.opendev.org&refresh=1m19:04
@clarkb:matrix.orgThere are definitely some improvemenst that can be made but I think this may be useable as is and maybe we land the change and then improve from there? In particular there is a systemd panel with graphs that don't work as we don't export systemd info so that can be removed. The disk info isn't working for some reason and that is useful info so should be corrected19:55
-@gerrit:opendev.org- Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org proposed: [opendev/zone-opendev.org] 1007953: Temporarily remove IPv6 for review and ymq mirror https://review.opendev.org/c/opendev/zone-opendev.org/+/100795319:57
@jim:acmegating.comalso i really hope we never need the "sytem timesync" panels20:03
@clarkb:matrix.orgfungi: should I go ahead and approve https://review.opendev.org/c/opendev/zone-opendev.org/+/1007953 ?20:12
@clarkb:matrix.orgI'm finishing some food then will go for a walk between meetings. But otherwise I'm around to keep an eye on it20:12
@fungicide:matrix.orgClark: yeah, i'm cooking/eating dinner but also around20:15
@fungicide:matrix.orgapprove at will20:15
-@gerrit:opendev.org- Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org proposed wip: [opendev/zone-opendev.org] 1007956: Revert "Temporarily remove IPv6 for review and ymq mirror" https://review.opendev.org/c/opendev/zone-opendev.org/+/100795620:17
@clarkb:matrix.orgdone20:16
@fungicide:matrix.organd that's ^ just to remind us later20:17
-@gerrit:opendev.org- Zuul merged on behalf of Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org: [opendev/zone-opendev.org] 1007953: Temporarily remove IPv6 for review and ymq mirror https://review.opendev.org/c/opendev/zone-opendev.org/+/100795320:20
@clarkb:matrix.orgmnasiadka: I think the reason that we don't have filesystem/disk info is that it isn't being collected. Your docker compose file does bind mount / and tells node exporter where that filesystem is per the documentation, but curling the metrics endpoint directly I'm not seeing the data and I don't see it in prometheus either21:13
@clarkb:matrix.orgI do see: `node_filesystem_device_error{device="/dev/vda1",device_error="no such file or directory",fstype="ext4",mountpoint="/host"} 1` I wonder if we need /dev bind mounted too so that it can see the devices as well?21:14
@clarkb:matrix.orghttps://oneuptime.com/blog/post/2026-07-31-node-filesystem-device-error-container/view#check-the-container-s-host-view this says that using the host pid namespace may be necessary?21:17
@clarkb:matrix.organyway I think this is a problem on the collection side not the dashboard side of things so we can probably treat these as independent issues and work them separately21:17
@clarkb:matrix.orgI'm also realizing that systems with volumes will make this more complicated. We may need to think about solutions within that context21:18
@fungicide:matrix.orgwhat we really want is to enumerate and track filesystems of specific types, not devices21:23
@fungicide:matrix.orgutilization occurs at the filesystem layer anyway, not the block device layer21:23
@fungicide:matrix.orghopefully it can interrogate e.g. `/proc/mounts` rather that digging through `/dev` looking for blockdevs21:26
@fungicide:matrix.orgbut yeah, i can see where doing that from inside a container is challenging21:28
@fungicide:matrix.orgjust like running snmpd inside a container probably would be21:28
@clarkb:matrix.orgya the problem I think is that the container gets a selective view of filesystems due to namespacing and all that so we probably need to figure out the right incantation to expose what we want. This should be testable too. Basically update the node exporter config and possibly bind mounts then have testinfra fetch the metrics and check for filesystem metrics that aren't errors like the one I posted above21:29
@clarkb:matrix.orghttps://github.com/prometheus/node_exporter#docker says pid: host and we aren't doing that so I suspect that will solve it21:30
@clarkb:matrix.orgthat won't fix the random volumes on various servers but should handle /21:30
@clarkb:matrix.orgI'll work on a change to get that going and test it21:31
@fungicide:matrix.orgi do worry that individually bind-mounting the filesystems into a container will lead to us forgetting to do it when we add something, so better if we don't need to21:31
@clarkb:matrix.orgwell to do that we'd need to stop running node exporter in a container21:32
@clarkb:matrix.orgwhich runs into potential problems of building node exporter for various systems (probably solvable as it is go right?)21:32
@fungicide:matrix.orgmaybe we need something to periodically scan for ext4 filesystems on our servers and update the agent container configs with the correct set of bindmounts, once we've got things generally working21:34
-@gerrit:opendev.org- Clark Boylan proposed: [opendev/system-config] 1007977: Add host pid namespace to node exporter docker config https://review.opendev.org/c/opendev/system-config/+/100797721:39
@clarkb:matrix.orgmaybe that will do it. The test case should help too21:39
@clarkb:matrix.orgoh I forget to add the documentation link21:40
-@gerrit:opendev.org- Clark Boylan proposed: [opendev/system-config] 1007977: Add host pid namespace to node exporter docker config https://review.opendev.org/c/opendev/system-config/+/100797721:40
@clarkb:matrix.orgsecond pre ptg block of the day is start now22:01
@clarkb:matrix.orgdetails are in https://etherpad.opendev.org/p/opendev-preptg-20260922:01
-@gerrit:opendev.org- Steve Baker proposed:23:03
- [openstack/diskimage-builder] 1007797: epel: allow epel-release to not be installed at all https://review.opendev.org/c/openstack/diskimage-builder/+/1007797
- [openstack/diskimage-builder] 1007798: Check for python3-pyyaml before installing https://review.opendev.org/c/openstack/diskimage-builder/+/1007798
- [openstack/diskimage-builder] 1007799: Handle passwd coming from shadow-utils on rhel-10 https://review.opendev.org/c/openstack/diskimage-builder/+/1007799
- [openstack/diskimage-builder] 1007800: pip-and-virtrualenv don't depend on epel https://review.opendev.org/c/openstack/diskimage-builder/+/1007800
- [openstack/diskimage-builder] 984486: Skip local loop device creation for no-final-image builds https://review.opendev.org/c/openstack/diskimage-builder/+/984486
- [openstack/diskimage-builder] 1007795: Add DIB_BIND_MOUNTS option for container environments https://review.opendev.org/c/openstack/diskimage-builder/+/1007795
- [openstack/diskimage-builder] 983815: Add tarball element for unprivileged container builds https://review.opendev.org/c/openstack/diskimage-builder/+/983815
- [openstack/diskimage-builder] 1007796: Add dnf-assert element https://review.opendev.org/c/openstack/diskimage-builder/+/1007796

Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!