Friday, 2026-09-18

opendevreviewGhanshyam Maan proposed openstack/nova master: Reduce the unnecessary executors shutdown logs in test jobs  https://review.opendev.org/c/openstack/nova/+/100416003:07
*** ykarel__ is now known as ykarel06:23
gibigmaan: re https://bugs.launchpad.net/nova/+bug/2166009 I think my question is about a different scenario. https://github.com/openstack/nova/blob/94de0576ca02e0668233bedfbe7e18b87316c708/nova/compute/manager.py#L6418 runs on the source, and if the virt driver raises an exception then we wont hit this cleanup code so we potentially leak a task on the *source* node of the resize07:39
gibidansmith: yeah, this wait forever on the executor is intentional (and debatable) in the functional test. I want to avoid leaking running threads across test cases as that leads to very hard to debug unexpected behaviors in later test cases. native threads are not killable so I cannot force cleanup between tests07:40
gibiyou (we) can try to add some task tracking in the functional test and log and fail instead of wait forever if the executor is not empty to help debugging.07:42
gibiI would love to work on this but swamped with important bugfixes at the moment07:42
gibiyou can try to locally flip the wait=true to wait=false in the executor shutdown and try to find a way to print the content of the excutor to help narroving which task is not finished in the test case07:43
gibiOK, maybe a way out, mock the shutdown in the func test to wait a bit but not forever for the executor and if the executor is not empty after the wait then fail the test case cleanly, that should be easy to implement and would avoid the hard to debug hang and clearly show at least which test case is effected07:47
opendevreviewBalazs Gibizer proposed openstack/nova master: Add missing resize.error notification  https://review.opendev.org/c/openstack/nova/+/100616308:13
opendevreviewBalazs Gibizer proposed openstack/nova master: Reproduce bug 2166786  https://review.opendev.org/c/openstack/nova/+/100465808:30
opendevreviewBalazs Gibizer proposed openstack/nova master: Fix PCI allocation leak on failed cold migration  https://review.opendev.org/c/openstack/nova/+/100600308:30
opendevreviewSylvain Bauza proposed openstack/nova master: Reject stale reserve_block_device_name RPC on compute  https://review.opendev.org/c/openstack/nova/+/100601108:35
opendevreviewBalazs Gibizer proposed openstack/nova master: Reproduce bug 2166786  https://review.opendev.org/c/openstack/nova/+/100465808:59
opendevreviewBalazs Gibizer proposed openstack/nova master: Fix PCI allocation leak on failed cold migration  https://review.opendev.org/c/openstack/nova/+/100600308:59
* zigo is running tempest on his packaged-based CI already, and VMs are already spawned. \o/10:23
zigo(Hibiscus, of course...)10:24
zigoI very much love that I was able to see:10:24
zigosetUpClass (tempest.api.compute.admin.test_live_migration.LiveMigrationWithVTPMTest) ... SKIPPED: LiveMigrationWithVTPMTest skipped as vTPM live migration is not enabled10:24
zigo:)10:24
zigoSuper nice feature, thanks guys !!!10:25
nicolairuckelIs anyone currently working on/thinking about this? https://bugs.launchpad.net/nova/+bug/178512312:07
gibinicolairuckel: isn't it fixed in https://review.opendev.org/c/openstack/nova/+/959682 ?12:09
gibior is that a partial fix?12:10
nicolairuckelThat was only a partial fix. That patch doesn't cover cold migration and shelve.12:10
nicolairuckelWe decided to deal with those in a separate patch (see https://review.opendev.org/c/openstack/nova/+/959682/comments/ae045d85_9e58a5a5?tab=comments)12:10
nicolairuckelUnfortunately, we can't just use libvirt for that like I did in the patch you linked.12:11
gibinicolairuckel: I'm not aware of anybody working on the cold migration part12:12
gibisean-k-mooney: ^^ maybe you have more context12:12
sean-k-mooneyam12:12
sean-k-mooneyso i tought we had started on it but i think we mainly just fixed the reboot case so far12:13
sean-k-mooneyi dont recall if there was a draft for cold migragte12:13
sean-k-mooneybut we did say we likely need an api change for rezie eventually12:13
sean-k-mooneyso sate if the nvram shold be cleared or not. sorry that was for rebuild12:14
sean-k-mooneyfor rebuild we need ot add an option to presserve/cleare the nvram12:14
sean-k-mooneyto cather for the rebuild form snapshot vs unrelated image case12:14
sean-k-mooneythere was https://review.opendev.org/c/openstack/nova/+/62164612:14
sean-k-mooneybut no https://review.opendev.org/c/openstack/nova/+/959682 did not fix cold migrate and shleve12:15
sean-k-mooneygibi: nicolairuckel  i remember there was a reason we split it that way12:16
sean-k-mooneybut i dont recall what that was12:16
gibithanks sean-k-mooney 12:17
nicolairuckelBecause of the API change we probably won't be able to backport the changes needed for cold migration. 12:17
nicolairuckelBut for the other patch we were able to backport them12:17
sean-k-mooneyfor cold migration we do not need an api change12:17
sean-k-mooneythat only for rebuild12:17
nicolairuckelah12:17
sean-k-mooneyi misspoke before12:17
sean-k-mooneyfor cold migrate we whosul always preseve the nvram12:18
nicolairuckelThe other reason was complexity. My patch turned out to be quite simple since we were able to use libvirt.12:18
nicolairuckelIs this the problem for rebuild? https://bugs.launchpad.net/nova/+bug/213907712:18
sean-k-mooneyno for rebuild we have 2 usecases12:18
sean-k-mooneysome people use rebuild for backup and restore or to upgrade the software in teh vm12:19
sean-k-mooneyso for that cohort we want to preserve the nvram12:19
sean-k-mooneybut you can also use rebuild to change the image entrily and even disable secure boot12:19
sean-k-mooneyso for that second group clearing the nvram is more desireable12:20
sean-k-mooneymy prefernce woudl be to preserve by default and have an new api option to ask for it to be cleared12:20
nicolairuckelSo I guess we should split that up even further: one patch for the cold migration and one for rebuild?12:20
sean-k-mooneyhttps://bugs.launchpad.net/nova/+bug/2139077 is a third edgecase but its a less common one12:21
sean-k-mooneyyes i would keep it split12:21
sean-k-mooneyim not sure what the curret rebuild behivor is12:21
sean-k-mooneyshelve is also the final edgbecase we did not adress12:22
sean-k-mooneywhich si it would be nice to save the nvram somewhere on shelve12:22
sean-k-mooneywe just didnt agree where12:22
sean-k-mooneyall of the edgecase for nvram also applies ot the tpm data too12:22
sean-k-mooneyfor windows in partical preserving the nvram and loosing the tpm data on cold migrate will require bitlocker recovery12:23
sean-k-mooneyso if i was to work on this i would do cold-migrate/resize next for both nvram and tpm12:24
nicolairuckelIn one patch or separate patches?13:13
bauzasgibi: hope your laptop is back, can I start to look at your patches ?13:13
nicolairuckelI'm not sure if we need TPM13:14
gibibauzas: give me 5 to push what I have (some missing tests) 13:15
bauzasack, I'll stop in ~1.5hour13:16
opendevreviewBalazs Gibizer proposed openstack/nova master: Fix PCI allocation leak on failed cold migration  https://review.opendev.org/c/openstack/nova/+/100600313:21
opendevreviewBalazs Gibizer proposed openstack/nova master: Clean PCI claim of failed resize in periodic task  https://review.opendev.org/c/openstack/nova/+/100620913:21
gibibauzas: so ^^ I have 3 patches13:22
gibi1. reproducer https://review.opendev.org/c/openstack/nova/+/100465813:22
gibi2. clean old leaks from periodics https://review.opendev.org/c/openstack/nova/+/100620913:22
gibi3. prevent new leaks https://review.opendev.org/c/openstack/nova/+/100600313:23
bauzasok13:23
gibithe 3rd is RPC sensitve :)13:23
gibiand I have to add unit test to the 2nd patch13:23
dansmithgibi: I found a deadlock in my own stuff that was definitely related (but haven't solved yet) so that may have been my only issue, but other test workers are also hanging, so not sure13:36
dansmithI guess it feels like maybe we should not wait there but fail if any tasks are left in some way instead of wait13:37
gibiyeah that is a good idea to implement13:40
gmaanyeah failing is also good, better than hanging. 15:55
gmaangibi: RE:resize_instance: on source it will not hang the task as task on source resize_instance is tracked via RPC common wrapper and it will always end the task as soon RPC requsted is ended either by success or fail16:00
gmaangibi: but yes, event cleanup will be left in failure case16:00
gmaangibi: dansmith but on dest side we have case of task leakage same as live migration so what you think of it and we can backport it to stable/2026.2 https://review.opendev.org/c/openstack/nova/+/100609516:01
*** ralonsoh is now known as ralonsoh_ooo16:02
gmaandansmith: I think I am seeing it in my patch (https://review.opendev.org/c/openstack/nova/+/1004160) also where test are hanging due to wait on executors https://zuul.opendev.org/t/openstack/build/9263ab2ac8c244f3a81f94efd50ad1c7/log/job-output.txt20:09
gmaanlet me convert the wait to failure and see 20:10
dansmithgmaan: I need to go EOD soon but are you saying we can't make the functional test fail instead of wait because of the task leakage on live migration?21:22
gmaandansmith: we can, I am working on something to propose soon22:53
gmaandansmith: live migration task leakage is when we will bump the manager_shutdown_timeout which is 0 currently 22:54

Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!