Thursday, 2026-07-09

opendevreviewIvan Anfimov proposed openstack/openstack-ansible-plugins master: Use kernel-modules-extra on RHEL for kernel headers task  https://review.opendev.org/c/openstack/openstack-ansible-plugins/+/99646002:31
opendevreviewOpenStack Proposal Bot proposed openstack/openstack-ansible master: Imported Translations from Zanata  https://review.opendev.org/c/openstack/openstack-ansible/+/99658704:17
opendevreviewMerged openstack/openstack-ansible master: Imported Translations from Zanata  https://review.opendev.org/c/openstack/openstack-ansible/+/99658708:39
kleiniI am still on 31.1.3, sorry for that. I encounter now that the periodic task in nova-compute, which should run every minute, did not run like for 20 minutes on some computes. Is there something known in those regards on Epoxy, what the possible cause is?10:05
kleiniI found nova-compute burning one full CPU core - 100% in top. After restart, that is gone and nova-compute down to 0%.10:18
noonedeadpunkI wasn't spotting anything like that tbh10:19
kleiniSo urgent time to upgrade to 31.3.0 or even to 33.0.010:25
noonedeadpunkI think we were mostly running 31.0.1 though10:42
kleinihttps://paste.opendev.org/show/bB3TjOURD65gRucyLvMS/ <- found something interesting11:41
mgariepyhmm interesting. is one of the placement api containers/server acting weird?11:54
kleiniI think, this is not the cause. That is some side effect, when getting the nova-compute process out of burning one CPU core. nova-compute opened the HTTP connection to haproxy but never sent the request for the periodic update of the compute node state.12:03
kleinisee in error message: "your browser didn't sent a request." that is haproxy response to nova-compute client not sending a request.12:03
mgariepyis nova-compute out of file descriptor?12:04
mgariepydoes it happens to all compute or only a handfull of them ?12:05
kleiniall nova-compute on all compute hosts burn one full CPU core. so, they're all stuck at the same point. will investigate, what the cause is.12:07
mgariepyi am very curious on what can causes that.12:08
mgariepydid they start all at the same time ?12:08
kleiniThis already runs now for months and mostly newly spawning VMs spread across 18 compute nodes. In this case more than 10 VMs should be spawned through server group affinity on the same compute node, revealing this issue.12:11
kleiniIt looks like it is somewhere towards connection to libvirtd going stale or something.12:14
kleinihttps://bugs.launchpad.net/nova/+bug/2147424 <- and here we are!12:14
kleinilibvirt-daemon 10.0.0-2ubuntu8.11 gotcha!12:16
mgariepyho nice.12:19
mgariepyonly 2 years.12:21
mgariepyit was a quick fix.. patch waiting in CI since june 202412:21
kleiniThe issue is, looking at it with "py-spy record --native" immediately resolves the issue as the zombie TCP connection is dropped.12:22
noonedeadpunkhm.... I can recall patching livbirt for 1012:26
noonedeadpunkI think the fix I have in mind was published in 10.0.0-2ubuntu8.12 or smth like that12:26
noonedeadpunkbut it was about https://bugs.launchpad.net/ubuntu/+source/libvirt/+bug/213318312:27
kleiniI see 10.0.0-2ubuntu8.14 available.12:27
kleinihmm, no quick fix in sight12:44
mgariepycronjob running py-spy every minutes!13:34
kleinisystemd_exporter and alert in Grafana load > 0.8 over 30m -> fire13:38
mgariepyi guess the nova patches linked should reduce the frequency tho. since it will make less calls.13:39
opendevreviewDmitriy Chubinidze proposed openstack/openstack-ansible-os_magnum master: Remove Heat-related variables and entities  https://review.opendev.org/c/openstack/openstack-ansible-os_magnum/+/99641114:20
-Guest12850- NOTICE: The Gerrit service on review.opendev.org will be offline briefly one hour from now, at 21:00 UTC, while we rename a project.20:00
-Guest12850- NOTICE: The Gerrit service on review.opendev.org is going offline momentarily at 21:00 UTC while we rename a project, but will return within a few minutes.20:58

Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!