| -@gerrit:opendev.org- Zuul merged on behalf of Steve Baker: [openstack/diskimage-builder] 983813: Refactor 02-set-machine-id into an element https://review.opendev.org/c/openstack/diskimage-builder/+/983813 | 00:19 | |
| -@gerrit:opendev.org- Stephen Finucane proposed: | 14:17 | |
| - [openstack/project-config] 1007878: gerrit: Sync ACLs across all oslo deliverables https://review.opendev.org/c/openstack/project-config/+/1007878 | ||
| - [openstack/project-config] 1007879: gerrit: Remove Backport-Candidate label from oslo projects https://review.opendev.org/c/openstack/project-config/+/1007879 | ||
| @fungicide:matrix.org | mmm, seems like static.o.o is having trouble serving content to me suddenly | 15:00 |
|---|---|---|
| @fungicide:matrix.org | `load average: 2201.01, 2232.10, 1393.90` | 15:00 |
| @fungicide:matrix.org | that seems mildly high | 15:01 |
| @clarkb:matrix.org | corvus: ^ will that impact zuul release? | 15:01 |
| @clarkb:matrix.org | fungi: only slightly above normal :) | 15:01 |
| @fungicide:matrix.org | i finally got content, but it took 30 seconds or more to load | 15:02 |
| @fungicide:matrix.org | basically all worker slots are in "sending reply" (W) state | 15:04 |
| @clarkb:matrix.org | the server is responsive so ya I don't think we're spinning the CPU and instead must be IO? | 15:04 |
| @fungicide:matrix.org | unsurprisingly, the vast majority of the requests currently being handled by workers are for docs.openstack.org | 15:05 |
| @clarkb:matrix.org | based on some non scientific log tailing docs.openstack.org is the majority ya that | 15:05 |
| @fungicide:matrix.org | i was looking at server-status instead of logs, but good correlation then | 15:06 |
| @clarkb:matrix.org | looks like significant numbers of 301s confusing the crawlers | 15:06 |
| @clarkb:matrix.org | so I think this is "the content with its redirect rules make this a particularly problematic set of content for crawlers" again | 15:07 |
| @clarkb:matrix.org | load is falling though | 15:07 |
| @jim:acmegating.com | > <@clarkb:matrix.org> corvus: ^ will that impact zuul release? | 15:08 |
| i don't think so | ||
| @clarkb:matrix.org | ack | 15:08 |
| @clarkb:matrix.org | I wonder if the crawlers are self moderating on the infinite recursion these days | 15:08 |
| @clarkb:matrix.org | because if load is the metric they seem to be backing off on their own | 15:09 |
| @clarkb:matrix.org | fungi: I wonder if this would make a good PTG topic for openstack | 15:09 |
| @fungicide:matrix.org | it would have to be a tc topic, there is no team managing that stuff | 15:09 |
| @clarkb:matrix.org | essentially the redirect rules are too generic/not specific enough and lead to crawlers redirecting themselves in a frenzy | 15:10 |
| @clarkb:matrix.org | which makes the content far more prone to these issues than other content | 15:10 |
| @clarkb:matrix.org | but I don't even know what would break if we start to try and make the rules less generic | 15:10 |
| @clarkb:matrix.org | if you grep for `" 301 " ` there should be plenty of examples in the docs.openstack.org access log | 15:13 |
| -@gerrit:opendev.org- Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org marked as active: [openstack/project-config] 1007491: Temporarily remove release docs semaphores https://review.opendev.org/c/openstack/project-config/+/1007491 | 15:14 | |
| @clarkb:matrix.org | `/security-guide/content/api-endpoints/dashboard/compute/dashboard/introduction/introduction/networking/image-storage/identity/compute.html` its like the crawlers are just doing matrix multiplication on valid file path tokens and hoping they find secret data | 15:14 |
| @fungicide:matrix.org | yes, we observed that behavior, like, last year even | 15:14 |
| @clarkb:matrix.org | yup definitely not new | 15:15 |
| @fungicide:matrix.org | and then that ends up redirecting to just '/security-guide/' | 15:15 |
| @fungicide:matrix.org | and the bots aren't smart enough to realize that the content they were redirected to is something they've already fetched | 15:16 |
| @clarkb:matrix.org | and that everything they try to fetch at /security-guide/random/cross/product/multiplication/of/token/terms will lead right back there too | 15:17 |
| @clarkb:matrix.org | except load is back to what I would consider normal so maybe they do realize that now after an initial burst of sillyness | 15:17 |
| @clarkb:matrix.org | I do still see 301s that looks like crawlers but rather than being all of the requests they are now a subset of requests | 15:20 |
| @clarkb:matrix.org | Following up on the zuul db issue from yesterday: disk consumption appears to be stable. No crazy new growth. I think it was corvus who mentioned we have been right on the edge for some time and we simply tipped over | 15:34 |
| @fungicide:matrix.org | makes sense, thanks for circling back around to check in on it | 15:35 |
| -@gerrit:opendev.org- James E. Blair https://matrix.to/#/@jim:acmegating.com proposed: [opendev/base-jobs] 1007900: Don't check for wheel in release jobs https://review.opendev.org/c/opendev/base-jobs/+/1007900 | 15:48 | |
| @jim:acmegating.com | Clark: fungi ^ that should address the issue in zuul's release job | 15:48 |
| @clarkb:matrix.org | I went ahead and approved that. The var name is correct according to git grep and I expect we'll be able to check this works shortly | 15:51 |
| -@gerrit:opendev.org- Zuul merged on behalf of James E. Blair https://matrix.to/#/@jim:acmegating.com: [opendev/base-jobs] 1007900: Don't check for wheel in release jobs https://review.opendev.org/c/opendev/base-jobs/+/1007900 | 16:02 | |
| -@gerrit:opendev.org- Zuul merged on behalf of Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org: [openstack/project-config] 1007491: Temporarily remove release docs semaphores https://review.opendev.org/c/openstack/project-config/+/1007491 | 16:37 | |
| -@gerrit:opendev.org- Clark Boylan proposed: [openstack/project-config] 1007930: Add a prometheus node exporter dashboard to grafana https://review.opendev.org/c/openstack/project-config/+/1007930 | 16:53 | |
| @clarkb:matrix.org | how do we feel about adapting existing grafana dashboards on their public dashboards listing/library to our needs? I've done that in ^ to see how it goes | 16:54 |
| @clarkb:matrix.org | will be interesting to see if I properly edited the template to not be a template | 16:54 |
| @jim:acmegating.com | it's an unreviewable change :) | 16:55 |
| @jim:acmegating.com | https://review.opendev.org/c/openstack/project-config/+/1007930/1/grafana/prometheus-node-exporter.json | 16:55 |
| @clarkb:matrix.org | ya I think the idea ianw had in the past was that we could hold a node and review the results? I agree 15k lines of json is not human friendly | 16:55 |
| @clarkb:matrix.org | though I don't know that anything we write would be signficantly smaller? | 16:55 |
| @clarkb:matrix.org | node exporter has a lot of data. Maybe the way to make it reviewable is to add panels in iteratiev changes? | 16:56 |
| @jim:acmegating.com | eh, we can look at the results | 16:56 |
| @mnasiadka:matrix.org | Well, screenshots of panels would not help a lot | 16:58 |
| @clarkb:matrix.org | mnasiadka: ya I think we would need to hold a node once we think we're at that point | 16:58 |
| @clarkb:matrix.org | also chances are that I got the untemplating process wrong and this will need a few edits before it is even usable | 16:59 |
| @mnasiadka:matrix.org | Well, there's Grafana Foundation SDK (Python), but for OpenDev needs sounds like an overkill | 17:03 |
| @clarkb:matrix.org | I think zuul missed the event for that change and hasn't enqueued it | 17:04 |
| @clarkb:matrix.org | I was going to bring this up in the pre ptg today but maybe we should drop the AAAA dns record for review? | 17:04 |
| @mnasiadka:matrix.org | Other option is just exposing a short guide/script how to start grafana in a container and using the in-git configuration for it - we don't do prometheus authentication - so that should work? | 17:05 |
| @clarkb:matrix.org | mnasiadka: yes you can run the container locally and then run grafyaml against it | 17:05 |
| @fungicide:matrix.org | if we drop the aaaa for review, should we also do it for mirror.ca-ymq-1.vexxhost? | 17:06 |
| @clarkb:matrix.org | it is fairly straightforward compared to many of the other systems as it is basically stateless other than the grafyaml application | 17:06 |
| @clarkb:matrix.org | fungi: possibly, but mirror.ca-ymq-1.vexxhost will still try to proxy via ipv6 so it may not be as clear cut there | 17:06 |
| @clarkb:matrix.org | I'm more worried about the openstack release losing half of its tag created events tomorrow | 17:06 |
| @clarkb:matrix.org | hrm the project-config change enqueued before my recheck I think. As zuul says it is 5 minutes old and my recheck is more like 3 minutes ago | 17:07 |
| @clarkb:matrix.org | but that is still ~10 minutes after I pushed the change. Maybe queues were backed up (Zuul has been busy) | 17:07 |
| @jim:acmegating.com | 14148306502e4b0eba21aa8c72c0fee1 is the event for your patchset-created | 17:10 |
| @clarkb:matrix.org | ack thanks. Looks like zuul saw the event without issues but then there was an 8.5 ish minute delay in processing it on zuul01 | 17:11 |
| @clarkb:matrix.org | which lines up with my rough numbers above | 17:11 |
| @jim:acmegating.com | it enqueued the change about 2 seconds after it received it | 17:11 |
| @fungicide:matrix.org | likely the result of 1007491 merging before it | 17:12 |
| @clarkb:matrix.org | corvus: it didn't render it in the UI that quickly. I agree it seems to say check needs to handle this very quickly | 17:12 |
| @clarkb:matrix.org | `2026-09-29 17:02:11,991 DEBUG zuul.Scheduler: [e: 14148306502e4b0eba21aa8c72c0fee1] Processing trigger event <GerritTriggerEvent patchset-created opendev.org/openstack/project-config master 1007930,1>` prior to this log entry there is a >8 minute lack of entries with that event id | 17:13 |
| @clarkb:matrix.org | oh processing trigger event is where it actually enqueues it to what the UI sees? | 17:14 |
| @clarkb:matrix.org | before that its in the queue to be processed? | 17:14 |
| @jim:acmegating.com | the ui is a snapshot, only updated after pipeline processing | 17:14 |
| @jim:acmegating.com | it's always going to be delayed | 17:14 |
| @clarkb:matrix.org | ya so its in the queue but zuul is busy and once it does Processing trigger event it enters a state where it is directly visible in the UI? | 17:15 |
| @jim:acmegating.com | i guess i'm trying to say that "visible in the ui" does not correspond to a log entry | 17:15 |
| @jim:acmegating.com | (i mean, we could come up with one if that was important, but it's like 3 degrees of separation from what we have now) | 17:15 |
| @clarkb:matrix.org | got it | 17:16 |
| @jim:acmegating.com | the event arrived at the pipeline 2 seconds after it was received; looks like the pipeline was busy for 8.5 minutes before it got to it | 17:16 |
| @clarkb:matrix.org | IIRC we changed what the queue values mean on the dashboard at some point. When I checked the two numbers were 0/0 so I thought zuul was cuaght up. Thinking out loud here I wonder if a "timestamp of last trigger event processed" value would help there? | 17:17 |
| @jim:acmegating.com | do we need to know what it was doing? or are we okay with "zuul didn't miss the event, it was busy" | 17:17 |
| @clarkb:matrix.org | I think I'm ok with zuul didn't miss the event it was busy. Just thinking out loud now how to potentially expose the "it is busy be a little patient" in the web ui | 17:18 |
| @jim:acmegating.com | Clark: the pipeline-level event queues are on the pipeline page: https://zuul.opendev.org/t/openstack/status/pipeline/check | 17:18 |
| @jim:acmegating.com | they were deemed too noisy for the tenant overview page | 17:18 |
| @jim:acmegating.com | (so the tenant overview only shows the tenant queues) | 17:19 |
| @jim:acmegating.com | i'd probably be okay putting them back if you think it's important | 17:19 |
| @clarkb:matrix.org | corvus: the three queue values top right being relevant here? | 17:20 |
| @jim:acmegating.com | yep | 17:20 |
| @clarkb:matrix.org | I'm not sure I event knew that there were different pipeline specific pages like that (I mean I may have reviewed the change at some point it just never registered in my brain what that means for me as a zuul user) | 17:20 |
| @clarkb:matrix.org | I don't know that it needs to go on the front page for status. | 17:21 |
| @clarkb:matrix.org | https://3d0ae5a1f74faf2e3f94-8cc229e52defecc7a0c44a815309c1d1.ssl.cf5.rackcdn.com/openstack/d53a439f80e74e61be6196c6c39fb8ad/screenshots/node-exporter-full.png neat it seems to work. I guess I should work on holding a node then | 17:30 |
| @clarkb:matrix.org | unfortunately the disk space graph didn't render. But I don't think we need everything working upfront for this to be valuable | 17:31 |
| @mnasiadka:matrix.org | Clark: wonder if we can uncollapse all the sections somehow | 17:31 |
| @mnasiadka:matrix.org | (unless they are empty and I miss the point) :) | 17:32 |
| @clarkb:matrix.org | mnasiadka: those panels have a collapsed: true attribute we could change | 17:32 |
| @clarkb:matrix.org | no they are collapsed by default in the config but it appears to eb editable | 17:32 |
| @mnasiadka:matrix.org | I think for user experience we might want them to be collapsed, but for the screenshot - it would be nice to have them all visible | 17:32 |
| @clarkb:matrix.org | next steps are probably: land the minimally viable thing, fix any issues like with disk consumption, then tune to what we want by uncollapsing etc | 17:32 |
| @fungicide:matrix.org | we're using selenium for the screenshots, doesn't it have the ability to "click on" certain elements in the dom? | 17:33 |
| @clarkb:matrix.org | I'll work on holding a node now so that we can decide if this is minimally viable | 17:33 |
| @clarkb:matrix.org | fungi: it does | 17:33 |
| @clarkb:matrix.org | but the screenshot sysstem is very generic. It lists all the dashboards from the grafana api then opens each one. It doesn't have per dashboard behaviors right now | 17:33 |
| @fungicide:matrix.org | ah, so it's more our plumbing that would need to be extended | 17:34 |
| @fungicide:matrix.org | which, yes, is more work | 17:34 |
| -@gerrit:opendev.org- Clark Boylan proposed: [opendev/system-config] 991512: DNM intentional grafana failure to hold a node https://review.opendev.org/c/opendev/system-config/+/991512 | 17:35 | |
| @clarkb:matrix.org | ok ^ hold request is in place for that now. zuul02 has pulled the zuul-client image as part of that process (yesterday I did it on zuul01 and pulled hte image there) | 17:36 |
| @fungicide:matrix.org | 5 minutes to the opendevent! | 17:55 |
| @mnasiadka:matrix.org | Coming back to the AAAA entries - I think we might need to tune gai.conf on the mirror to prefer ipv4 until ipv6 gets fixed | 17:55 |
| @mnasiadka:matrix.org | fungi: are you going to devent? | 17:55 |
| @clarkb:matrix.org | mnasiadka: ya that would solve the proxying problem on the backend | 17:56 |
| @clarkb:matrix.org | but lets talk about it in a few minuets as this is topical | 17:56 |
| @mnasiadka:matrix.org | I won't join today, have some other plans - but will be there tomorrow ;-) | 17:57 |
| @clarkb:matrix.org | cool I know today is a bit later for you too. See you tomorrow | 17:57 |
| @fungicide:matrix.org | mnasiadka: i didn't even know it was a thing | 17:57 |
| @mnasiadka:matrix.org | I was just laughing, opendevent reminds me of venting (negative emotions stuff) | 17:58 |
| @mnasiadka:matrix.org | But it seems it's a thing, removing trapped air from closed system ;-) | 17:58 |
| @fungicide:matrix.org | oh that too, yep! | 17:59 |
| @clarkb:matrix.org | we can vent and event | 17:59 |
| @clarkb:matrix.org | its time. I made it into the meetpad first so I am moderator! | 18:00 |
| @clarkb:matrix.org | held grafana node with node exporter dashboard here: https://217.182.140.216/d/49af61ae98/node-exporter-full?orgId=1&from=now-24h&to=now&timezone=utc&var-job=node&var-nodename=prometheus01&var-node=prometheus01.opendev.org&refresh=1m | 19:04 |
| @clarkb:matrix.org | There are definitely some improvemenst that can be made but I think this may be useable as is and maybe we land the change and then improve from there? In particular there is a systemd panel with graphs that don't work as we don't export systemd info so that can be removed. The disk info isn't working for some reason and that is useful info so should be corrected | 19:55 |
| -@gerrit:opendev.org- Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org proposed: [opendev/zone-opendev.org] 1007953: Temporarily remove IPv6 for review and ymq mirror https://review.opendev.org/c/opendev/zone-opendev.org/+/1007953 | 19:57 | |
| @jim:acmegating.com | also i really hope we never need the "sytem timesync" panels | 20:03 |
| @clarkb:matrix.org | fungi: should I go ahead and approve https://review.opendev.org/c/opendev/zone-opendev.org/+/1007953 ? | 20:12 |
| @clarkb:matrix.org | I'm finishing some food then will go for a walk between meetings. But otherwise I'm around to keep an eye on it | 20:12 |
| @fungicide:matrix.org | Clark: yeah, i'm cooking/eating dinner but also around | 20:15 |
| @fungicide:matrix.org | approve at will | 20:15 |
| -@gerrit:opendev.org- Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org proposed wip: [opendev/zone-opendev.org] 1007956: Revert "Temporarily remove IPv6 for review and ymq mirror" https://review.opendev.org/c/opendev/zone-opendev.org/+/1007956 | 20:17 | |
| @clarkb:matrix.org | done | 20:16 |
| @fungicide:matrix.org | and that's ^ just to remind us later | 20:17 |
| -@gerrit:opendev.org- Zuul merged on behalf of Jeremy Stanley https://matrix.to/#/@fungicide:matrix.org: [opendev/zone-opendev.org] 1007953: Temporarily remove IPv6 for review and ymq mirror https://review.opendev.org/c/opendev/zone-opendev.org/+/1007953 | 20:20 | |
| @clarkb:matrix.org | mnasiadka: I think the reason that we don't have filesystem/disk info is that it isn't being collected. Your docker compose file does bind mount / and tells node exporter where that filesystem is per the documentation, but curling the metrics endpoint directly I'm not seeing the data and I don't see it in prometheus either | 21:13 |
| @clarkb:matrix.org | I do see: `node_filesystem_device_error{device="/dev/vda1",device_error="no such file or directory",fstype="ext4",mountpoint="/host"} 1` I wonder if we need /dev bind mounted too so that it can see the devices as well? | 21:14 |
| @clarkb:matrix.org | https://oneuptime.com/blog/post/2026-07-31-node-filesystem-device-error-container/view#check-the-container-s-host-view this says that using the host pid namespace may be necessary? | 21:17 |
| @clarkb:matrix.org | anyway I think this is a problem on the collection side not the dashboard side of things so we can probably treat these as independent issues and work them separately | 21:17 |
| @clarkb:matrix.org | I'm also realizing that systems with volumes will make this more complicated. We may need to think about solutions within that context | 21:18 |
| @fungicide:matrix.org | what we really want is to enumerate and track filesystems of specific types, not devices | 21:23 |
| @fungicide:matrix.org | utilization occurs at the filesystem layer anyway, not the block device layer | 21:23 |
| @fungicide:matrix.org | hopefully it can interrogate e.g. `/proc/mounts` rather that digging through `/dev` looking for blockdevs | 21:26 |
| @fungicide:matrix.org | but yeah, i can see where doing that from inside a container is challenging | 21:28 |
| @fungicide:matrix.org | just like running snmpd inside a container probably would be | 21:28 |
| @clarkb:matrix.org | ya the problem I think is that the container gets a selective view of filesystems due to namespacing and all that so we probably need to figure out the right incantation to expose what we want. This should be testable too. Basically update the node exporter config and possibly bind mounts then have testinfra fetch the metrics and check for filesystem metrics that aren't errors like the one I posted above | 21:29 |
| @clarkb:matrix.org | https://github.com/prometheus/node_exporter#docker says pid: host and we aren't doing that so I suspect that will solve it | 21:30 |
| @clarkb:matrix.org | that won't fix the random volumes on various servers but should handle / | 21:30 |
| @clarkb:matrix.org | I'll work on a change to get that going and test it | 21:31 |
| @fungicide:matrix.org | i do worry that individually bind-mounting the filesystems into a container will lead to us forgetting to do it when we add something, so better if we don't need to | 21:31 |
| @clarkb:matrix.org | well to do that we'd need to stop running node exporter in a container | 21:32 |
| @clarkb:matrix.org | which runs into potential problems of building node exporter for various systems (probably solvable as it is go right?) | 21:32 |
| @fungicide:matrix.org | maybe we need something to periodically scan for ext4 filesystems on our servers and update the agent container configs with the correct set of bindmounts, once we've got things generally working | 21:34 |
| -@gerrit:opendev.org- Clark Boylan proposed: [opendev/system-config] 1007977: Add host pid namespace to node exporter docker config https://review.opendev.org/c/opendev/system-config/+/1007977 | 21:39 | |
| @clarkb:matrix.org | maybe that will do it. The test case should help too | 21:39 |
| @clarkb:matrix.org | oh I forget to add the documentation link | 21:40 |
| -@gerrit:opendev.org- Clark Boylan proposed: [opendev/system-config] 1007977: Add host pid namespace to node exporter docker config https://review.opendev.org/c/opendev/system-config/+/1007977 | 21:40 | |
| @clarkb:matrix.org | second pre ptg block of the day is start now | 22:01 |
| @clarkb:matrix.org | details are in https://etherpad.opendev.org/p/opendev-preptg-202609 | 22:01 |
| -@gerrit:opendev.org- Steve Baker proposed: | 23:03 | |
| - [openstack/diskimage-builder] 1007797: epel: allow epel-release to not be installed at all https://review.opendev.org/c/openstack/diskimage-builder/+/1007797 | ||
| - [openstack/diskimage-builder] 1007798: Check for python3-pyyaml before installing https://review.opendev.org/c/openstack/diskimage-builder/+/1007798 | ||
| - [openstack/diskimage-builder] 1007799: Handle passwd coming from shadow-utils on rhel-10 https://review.opendev.org/c/openstack/diskimage-builder/+/1007799 | ||
| - [openstack/diskimage-builder] 1007800: pip-and-virtrualenv don't depend on epel https://review.opendev.org/c/openstack/diskimage-builder/+/1007800 | ||
| - [openstack/diskimage-builder] 984486: Skip local loop device creation for no-final-image builds https://review.opendev.org/c/openstack/diskimage-builder/+/984486 | ||
| - [openstack/diskimage-builder] 1007795: Add DIB_BIND_MOUNTS option for container environments https://review.opendev.org/c/openstack/diskimage-builder/+/1007795 | ||
| - [openstack/diskimage-builder] 983815: Add tarball element for unprivileged container builds https://review.opendev.org/c/openstack/diskimage-builder/+/983815 | ||
| - [openstack/diskimage-builder] 1007796: Add dnf-assert element https://review.opendev.org/c/openstack/diskimage-builder/+/1007796 | ||
Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!