| opendevreview | Ghanshyam Maan proposed openstack/nova master: Fix task leak for live migration rollback case https://review.opendev.org/c/openstack/nova/+/1003386 | 03:25 |
|---|---|---|
| opendevreview | Ghanshyam Maan proposed openstack/nova master: [func test]Catch hanging task at graceful shutdown https://review.opendev.org/c/openstack/nova/+/1003266 | 03:25 |
| gmaan | gibi: on https://review.opendev.org/c/openstack/nova/+/1001369, i replied to your comment. basically test is failing because of RPC thread is beeing sleep and shutdown wait for that to finish si deadlock. manager_shutdown_timoeut does not solve things here | 03:27 |
| gmaan | gibi: ashigupt_ I am testing this on manager_shutdown_timeout default value change also https://review.opendev.org/c/openstack/nova/+/1003266 | 03:28 |
| gmaan | gibi: sambork: Uggla: I will not be able to join eventlet or upstream bug meeting if any. still in my kids school transition and will be online around ~10 AMish my time | 03:29 |
| opendevreview | Goutham Pacha Ravi proposed openstack/nova stable/2026.1: Use a per-instance cephx identity for CephFS shares https://review.opendev.org/c/openstack/nova/+/1004765 | 05:14 |
| opendevreview | Goutham Pacha Ravi proposed openstack/nova stable/2025.2: Use a per-instance cephx identity for CephFS shares https://review.opendev.org/c/openstack/nova/+/1004766 | 05:21 |
| opendevreview | Goutham Pacha Ravi proposed openstack/nova stable/2025.1: Use a per-instance cephx identity for CephFS shares https://review.opendev.org/c/openstack/nova/+/1004767 | 05:21 |
| sambork | gmaan, thanks for letting us know! | 06:18 |
| opendevreview | Arnaud Morin proposed openstack/os-vif master: Drop pyroute2 win32 marker https://review.opendev.org/c/openstack/os-vif/+/1004780 | 08:19 |
| opendevreview | Arnaud Morin proposed openstack/nova master: Add pyroute2 in doc dependencies https://review.opendev.org/c/openstack/nova/+/1004781 | 08:23 |
| amorin | Hey nova team, I created the two changes above ^ | 08:23 |
| amorin | to fix an issue in documentation (at least) not showing the os_vif_ovs parameters anymore | 08:24 |
| amorin | e.g. https://docs.openstack.org/nova/latest/configuration/config.html | 08:24 |
| amorin | look for ovsdb_connection | 08:24 |
| amorin | and compare with: https://docs.openstack.org/nova/2026.1/configuration/config.html#os_vif_ovs.ovsdb_connection | 08:24 |
| amorin | I believe that should be fix before 2026.2 release, or the release will miss these params | 08:25 |
| mhen | Uggla: are you around? | 08:32 |
| Uggla | Hi mhen | 08:32 |
| mhen | hi :) | 08:33 |
| mhen | Concerning https://lists.openstack.org/archives/list/openstack-discuss@lists.openstack.org/thread/AE4FL6L374WQTUJI65RDHQE3GLTI3R5B/ | 08:33 |
| mhen | eharney stated "i think the basic summary there is that we landed cinder code but never completed nova work for it, so it's pretty broken" | 08:33 |
| mhen | which seemed to confirm my suspicion that things in Nova (and os-brick) are missing | 08:34 |
| mhen | I am trying to get a cross-project session between Cinder and Nova organized for the upcoming PTG for this | 08:34 |
| mhen | Uggla: can you assist me with this? is there a PTG etherpad for Nova to register sessions? | 08:36 |
| mhen | (I will also talk to Cinder in their meeting today) | 08:36 |
| Uggla | mhen, sure I can help to discuss that, no pb. I have not created the ptg etherpad page for Indri. | 08:39 |
| Uggla | mhen, I will create the page today and ping you. So you will be able to enter the topic | 08:40 |
| mhen | Uggla: splendid, thank you! | 08:40 |
| Uggla | mhen, https://etherpad.opendev.org/p/nova-2027.1-ptg (a bit raw at the moment), but please add your topic at the bottom and I will schedule it for the vPTG. | 08:48 |
| *** ChanServ sets mode: +o Uggla | 08:55 | |
| *** ChanServ changes topic to "This channel is for Nova development | Development-planning: https://etherpad.opendev.org/p/nova-2026.2-status | PTG doc: https://etherpad.opendev.org/p/nova-2027.1-ptg | This channel is logged at https://meetings.opendev.org/irclogs/%23openstack-nova/" | 08:58 | |
| *** ChanServ sets mode: -o Uggla | 09:00 | |
| stephenfin | Uggla: a reminder that we're still waiting on your ack on https://review.opendev.org/c/openstack/project-config/+/988094/3 | 10:25 |
| mhen | Uggla: thanks, I have added the cross-project session description to the bottom | 11:15 |
| gibi | amorin: replied in your doc patches | 11:37 |
| amorin | ack thanks! | 12:05 |
| opendevreview | Arnaud Morin proposed openstack/nova master: Add pyroute2 in doc dependencies https://review.opendev.org/c/openstack/nova/+/1004781 | 12:07 |
| gibi | amorin: if we miss the RC1 deadline with this I would not sweat on it much and just fix it on master and backport as soon as the GA is out and then push a point release if needed from stable | 12:08 |
| gibi | but feel free to disagree if this somehow creates an ackward situation productifying the GA | 12:08 |
| gibi | Uggla: can arrange for an RC2 if it is warranted | 12:08 |
| gibi | s/:// | 12:08 |
| amorin | I do not care that much on my side, I was just surprised that the doc was wrong. I may not be the only one, so better trying to have this before the GA IMHO | 12:09 |
| gibi | ack | 12:17 |
| opendevreview | Balazs Gibizer proposed openstack/nova master: Reproduce bug 2166786 https://review.opendev.org/c/openstack/nova/+/1004658 | 12:32 |
| opendevreview | sean mooney proposed openstack/nova master: move nova-alt-config to debian-13 https://review.opendev.org/c/openstack/nova/+/966479 | 12:46 |
| opendevreview | Stephen Finucane proposed openstack/placement master: tox: Use constraints option https://review.opendev.org/c/openstack/placement/+/996893 | 13:15 |
| opendevreview | Stephen Finucane proposed openstack/placement master: Replace license classifier https://review.opendev.org/c/openstack/placement/+/1004422 | 13:15 |
| opendevreview | Stephen Finucane proposed openstack/placement master: Migrate requirements to pyproject.toml https://review.opendev.org/c/openstack/placement/+/1004423 | 13:15 |
| Uggla | stephenfin, I have just +1 the change above, sorry for the latency. | 13:17 |
| gibi | Uggla: finally back to https://review.opendev.org/c/openstack/nova/+/999977/ I left some small suggestion inline but it is looks good pretty good now | 13:42 |
| Uggla | gibi, thanks I will look at that. | 13:43 |
| gibi | Uggla: also left some comments in the RC1 etherpad about api microversion history and prelude I guess those are still needs to be done | 14:12 |
| Uggla | gibi, regarding prelude as it is a copy of highlight, I was waiting the HL to be settled which is the case now, so I'll create it. Thanks for the microversion history because I thought it was ok but missed it. | 14:17 |
| gibi | cool | 14:18 |
| gibi | I think the rest looks good | 14:18 |
| opendevreview | Merged openstack/nova-specs master: Move Hibiscus implemented specs https://review.opendev.org/c/openstack/nova-specs/+/1004442 | 14:20 |
| gibi | Uggla: left feedback on https://review.opendev.org/c/openstack/nova/+/999978 I think this evac bug and my evac bug discussed on the Monday call are sharing a root case. I left suggestion how to combine the two fixes that is non-controversial from the Monday discussion perspective but probably useful for both bugs. It is pretty complicated so feel free to ask me to jump on a call to explain more | 15:01 |
| bauzas | gibi: I guess we won't have eventlet meeting, right? | 15:05 |
| Uggla | gibi, yeah I had the feeling that Monday topic was more or less linked and that a 2 steps approach is required. One short term to meet customer happy and one longer term for the PTG. | 15:07 |
| opendevreview | ribaudr proposed openstack/nova master: Mark 2.104 as maximum API version for 2026.2 Hibiscus https://review.opendev.org/c/openstack/nova/+/1004872 | 15:15 |
| gibi | bauzas: it is ongoing but nothing important there | 15:22 |
| dansmith | gmaan: sorry I just realized I never replied on your task leak thing | 15:27 |
| * dansmith looks now | 15:27 | |
| Uggla | Reminder: upstream triage (https://meet.google.com/zjr-rxus-hzj) | 15:30 |
| gibi | couple of mintues lte | 15:31 |
| gibi | late | 15:31 |
| opendevreview | Grzegorz Grasza proposed openstack/nova master: Match fixed_ip literally instead of partially escaping it as a regex https://review.opendev.org/c/openstack/nova/+/1004878 | 15:39 |
| opendevreview | Clif Houck proposed openstack/nova stable/2026.1: perf(ironic): eliminate O(N²) ProviderTree deepcopy at startup https://review.opendev.org/c/openstack/nova/+/1004890 | 16:28 |
| opendevreview | Merged openstack/placement master: tox: Use constraints option https://review.opendev.org/c/openstack/placement/+/996893 | 16:37 |
| dansmith | gibi: are you still around/ | 16:43 |
| gmaan | dansmith: thanks for reply, reading | 16:52 |
| dansmith | gmaan: I'm curious what gibi thinks about this | 16:52 |
| dansmith | wondering if he was already okay with the RPC breakage or didn't realize when he +2d | 16:52 |
| gmaan | dansmith: I think he mentioned about the revert if it fails but not specific to new vs old node | 16:53 |
| dansmith | I know | 16:54 |
| opendevreview | Clif Houck proposed openstack/nova stable/2025.2: perf(ironic): eliminate O(N²) ProviderTree deepcopy at startup https://review.opendev.org/c/openstack/nova/+/1004899 | 17:03 |
| gmaan | dansmith: I think, adding a new RPC call (end_tracked_task or something) will be better here with RPC version bump. do you want me change that way or wanted to wait for gibi reply? | 17:05 |
| dansmith | I guess I'd rather not have a new call just for the very rare "cancel before started" case... why do you like that better? | 17:06 |
| gmaan | that is what i was thinking to bump the RPC for either change but was not sure if we can do the version bumo at this stage or not. if yes then no issue | 17:06 |
| gmaan | dansmith: it will not cancel before started, it can check if task is recorded otherwise log warning | 17:07 |
| dansmith | it's not great to bump for either for sure at this point | 17:07 |
| dansmith | gmaan: it's a cancel to the destination because we basically never started the migration which now doesn't need to be cleaned up right? | 17:08 |
| gmaan | I am thinking it generic way where this common RPC call can be used in other place, because I am not sure we handle all cases of error on each node for cross-node operations | 17:08 |
| dansmith | bauzas: are you still around? | 17:08 |
| bauzas | I am | 17:08 |
| * bauzas looking above | 17:08 | |
| dansmith | bauzas: can you skim this? https://review.opendev.org/c/openstack/nova/+/1003386 | 17:08 |
| dansmith | just looking for another opinion.. I don't really like either option but they both require an RPC bump to fix (or we just ignore this case and leave it as a delayed shutdown bug I guess) | 17:09 |
| bauzas | hah, ok, I can try to take a look but I don't have a lot of context for now | 17:10 |
| gmaan | dansmith: but if rollback happening on source then it means migration started in dest right? | 17:10 |
| dansmith | gmaan: started in nova but probably not libvirt yet right? | 17:11 |
| gmaan | yes not till libvirt | 17:11 |
| bauzas | wait | 17:12 |
| bauzas | now I have a bit more of knowledge | 17:12 |
| bauzas | all of the skim is for _post_live_migration, ie. once the instance is either fully migrated or rollbacked | 17:13 |
| dansmith | actually, I guess not... it's any shared storage migration, even if it has already started | 17:13 |
| bauzas | because we know we have remnants to cleanup, we have to modify the other compute to remove them | 17:13 |
| dansmith | right, that's the case where we _do_ make the call | 17:14 |
| dansmith | but the bug is in cases where we don't | 17:14 |
| dansmith | as of the graceful shutdown stuff, we _always_ need to make the (or a) call to the dest because we left a pending task if nothing else | 17:15 |
| bauzas | https://github.com/openstack/nova/blob/d3a1c95d06c2d97368207c2728fc262796683ad8/nova/compute/manager.py#L10557 | 17:15 |
| dansmith | so to me the right solution is to always make that call in the future and pass it the cleanup flag maybe, but we can only do that with a version bump | 17:15 |
| bauzas | all of that is in _post_live_migrate | 17:15 |
| dansmith | else we will ask n-1 computes to do a cleanup when there is none to do | 17:15 |
| gmaan | its during cleanup of pre live migration not post | 17:16 |
| dansmith | right | 17:16 |
| dansmith | well, during cleanup of live migration... to clean up what was started in pre_live_migration | 17:17 |
| bauzas | https://docs.openstack.org/nova/latest/reference/live-migration.html | 17:17 |
| gmaan | dansmith: yeah, that is my preference too when i proposed change, if we can delay this bug and mentioned as known bug and in part-2 we can handle by flag or new call | 17:17 |
| gmaan | dansmith: yes -> 'to clean up what was started in pre_live_migration' | 17:17 |
| dansmith | bauzas: just to be clear, I'm looking for opinions on if we can/should bump the compute RPC to add this call/param right now, just before rc1 | 17:18 |
| bauzas | we're changing the live-migration workflow per se, not exactly the RPC interface | 17:19 |
| bauzas | so that's kind of a grey area | 17:19 |
| bauzas | but in case of a rolling upgrade, we have to ensure we're not breaking for sure | 17:20 |
| dansmith | huh? | 17:20 |
| bauzas | because old compute could expect a different workflow | 17:20 |
| dansmith | it's 100% an RPC change :) | 17:20 |
| gmaan | yeah, old dest will always do the rollback at dest when not needed and most probably fail | 17:21 |
| bauzas | we don't change the RPC parameters for sure, but we changing the callers yes | 17:21 |
| dansmith | not only are we expecting the new behavior of canceling the task, we need to be passing a parameter to instruct the old side not to try to clean up something it shouldn't be cleaning up | 17:21 |
| dansmith | bauzas: we _need_ to add a parameter | 17:21 |
| bauzas | lemme explain it differently | 17:21 |
| gmaan | dansmith: flag can be avoided as as current code change where dest can calculate the cleanup is required or not | 17:21 |
| dansmith | right now this will just generate (we think) a traceback | 17:21 |
| bauzas | either we consider the existing patch good and then I have a concern due to rolling upgrades (old computes may expect the calls differently), or we assume we need to pass a new param and then I have concern witth a late RPC change | 17:22 |
| dansmith | gmaan: but the flag also gives us the indication that we're making the new call so we can avoid doing it on the older versioon | 17:23 |
| bauzas | either way, I'm not happy with such behavioural change so close to the release | 17:23 |
| dansmith | bauzas: I don't think I'm okay with no bump and either change | 17:23 |
| dansmith | it's very clearly just side-stepping RPC behavioral expectations | 17:23 |
| gmaan | dansmith: that can but we can just check the same based on pin version and cleanup flag calculation. I mean flag is better way but if we want to avoid RPC signature change that is possible | 17:24 |
| bauzas | ship has sailed for a RPC bump this release https://review.opendev.org/c/openstack/nova/+/1004424 | 17:25 |
| dansmith | gmaan: that would be compute manager inspecting the pin set in the rpcapi right? | 17:25 |
| bauzas | dansmith: that was my "either" : the existing proposal has upgrade concners | 17:25 |
| dansmith | bauzas: that hasn't merged so I don't think it has _sailed_ nor do I think landing that is terminal until we make the release :) | 17:26 |
| bauzas | can't we just defer the fix to post-rc1 ? | 17:26 |
| dansmith | meaning have this be broken in H? we can for sure, that's one of the options | 17:27 |
| gmaan | dansmith: ah yeah. i get your point now | 17:27 |
| gmaan | so what is main concern on stating it as known bug which is in scenario 1. live migration started at pre migration stage, source rolling it back , it is shared storage where cleanup not needed at dest 2. dest shutting down. Issue: it will hold dest graceful shutdown for full time as configured (manager_shutdown_time default 160 sec) | 17:27 |
| dansmith | gmaan: yep, that's one option and was also something I think is worth surveying opinions on | 17:28 |
| gmaan | only issue it will cause is to make shutdown taking full time instead of having task tracking benefits in this scenario | 17:28 |
| dansmith | yup | 17:28 |
| gmaan | which is nothing change than before this rlease | 17:28 |
| gmaan | bcz I am not not 100% ok to bump RPC version or try some workaround which we need to chaneg again | 17:28 |
| gmaan | *I am also not | 17:28 |
| bauzas | because we're leaking threads as of now ? | 17:29 |
| dansmith | not leaking threads really.. started a task that will wait for $timeout before allowing n-cpu to exit because it thinks a migration will still be coming | 17:29 |
| bauzas | I see | 17:29 |
| gmaan | bauzas: bcz we are leaking task from task tracking, live migration task started at dest willnot be ended as source do rollbacl and do not call dest | 17:29 |
| dansmith | but it's very rare that this would be hit, and if you don't enable (or do disable) graceful shutdown it won't matter anyway | 17:30 |
| bauzas | I think the stable policy still allows us to backport an RPC change provided certain constraints, but I can doublechecks | 17:30 |
| dansmith | so maybe just merge a reno with known-bug | 17:30 |
| bauzas | doublecheck* | 17:30 |
| dansmith | bauzas: we can but it is majorly less good to do that than just merge it now | 17:30 |
| dansmith | I think if we don't fix now, we won't fix in stable | 17:30 |
| bauzas | I see your point, I don't have a particular opinion about it except that I don't want the existing patch to be merged without testing rolling upgrades | 17:31 |
| gmaan | yeah, fix would not be backportable | 17:32 |
| bauzas | we don't have a job that runs live-migration on a multinode grenade env, right? | 17:33 |
| * bauzas looks at https://zuul.opendev.org/t/openstack/build/fed98842b85842aeb58b3f5a62803e39 | 17:33 | |
| dansmith | I need to take a break | 17:34 |
| dansmith | I think I'm most in favor of just #knownbug'ing it and fix it in I | 17:34 |
| dansmith | if it absolutely explodes in our face, we can backport something like the current patch (or offer it to people out of band) but I doubt it will | 17:35 |
| gmaan | dansmith: I have one more solution which does not need RPC bump or any ordering change of cleaup | 17:35 |
| bauzas | https://zuul.opendev.org/t/openstack/job/nova-grenade-multinode is testing live-migration on a rolling upgrade fashion but we don't run graceful shutdown on it | 17:36 |
| gmaan | dansmith: so this rollback at sourec happening after pre_live_migration at dest and where dest raise any error. and task is leaked beacuxse we start the live migration at dest task before dest's pre_live_migration | 17:36 |
| gmaan | what if we start the live-migration_at_dest task after dest's pre_live_migration is completed so that this rollback at source will not leak the task at dest | 17:37 |
| gmaan | I mean record lve migration at dest when we know it is actually going to start and not pre live migration failed | 17:37 |
| dansmith | I dunno, that seems like a lot more change from what we have now and I'm probably too fried to think about it | 17:38 |
| dansmith | I also thought about making the task on the dest just poll the migration object while it's waiting, but that is more work so late as well | 17:38 |
| gmaan | I mean we are starting live migration at dest too early even we do not know if dest's pre live migration things are successful or not | 17:39 |
| gmaan | poll task where and how to decide when to end it? | 17:39 |
| gmaan | dansmith: for 'late' yes, but this sems backportable fix if we do it in next release right? | 17:40 |
| dansmith | poll task in an executor task, end it when the migration goes to finished or error.. it was just a thought and something we'd have to look at | 17:40 |
| dansmith | gmaan: potentially | 17:40 |
| gmaan | dansmith: ohk so based on the migration status itself. yeah that is one possible way | 17:41 |
| bauzas | I had to move errand, but I'd indeed prefer the third solution, which is to end the task when the migration finishes, instead of overloading our RPC interface for that need | 18:11 |
| gmaan | yeah, i am finding that solution a better one. My solution of shifting the start_task is not perfect and does not solve issue when dest's pre live migration is successful but vif plugged event can cause rollback on source - | 18:27 |
| gmaan | dansmith: bauzas i commented in gerrit about going for 'listing it as a known bug' path (please check), we can wait for gibi on his opinion and i can proceed accordingly | 18:34 |
| opendevreview | ribaudr proposed openstack/nova master: Add Hibiscus prelude section https://review.opendev.org/c/openstack/nova/+/1004927 | 20:31 |
Generated by irclog2html.py 4.1.0 by Marius Gedminas - find it at https://mg.pov.lt/irclog2html/!