Skip to content

VM HA not working when there's an issue with vm restart #13510

Description

@gusmef

problem

We're trying VM HA on cloudstack. We disabled HOST HA (see issue 13371) and then simulated a crash on an host with a "fake" kernel panic. The investigators correctly found the host dead and put it in alert/disconnected state.

2026-06-29 10:01:03,855 DEBUG [o.a.c.k.h.KVMHostActivityChecker] (AgentTaskPool-3:[ctx-5facfbaf]) (logid:9ccd9518) Checking Host {"id":76,"name":"fakehost1.ha","type":"Routing","uuid":"ad118d97-03e2-4555-9d70-595864e87
1fd"} status...
2026-06-29 10:01:20,849 INFO  [c.c.a.m.ClusteredAgentManagerImpl] (AgentTaskPool-4:[ctx-dd435ced]) (logid:69b8079e) Investigating why host Host {"id":76,"name":"fakehost1.ha","type":"Routing","uuid":"ad118d97-03e2-4555
-9d70-595864e871fd"} has disconnected with event
2026-06-29 10:01:23,858 DEBUG [o.a.c.k.h.KVMHostActivityChecker] (AgentTaskPool-3:[ctx-5facfbaf]) (logid:9ccd9518) Setting Host {"id":76,"name":"fakehost1.ha","type":"Routing","uuid":"ad118d97-03e2-4555-9d70-595864e871
fd"} to "Disconnected" status.
2026-06-29 10:01:23,860 DEBUG [o.a.c.k.h.KVMHostActivityChecker] (AgentTaskPool-3:[ctx-5facfbaf]) (logid:9ccd9518) Investigating Host {"id":76,"name":"fakehost1.ha","type":"Routing","uuid":"ad118d97-03e2-4555-9d70-5958
64e871fd"} via neighboring Host {"id":75,"name":"fakehost2.ha","type":"Routing","uuid":"ffea4f8e-0d91-4a22-b600-aea1891be419"}.
2026-06-29 10:01:23,918 DEBUG [o.a.c.k.h.KVMHostActivityChecker] (AgentTaskPool-3:[ctx-5facfbaf]) (logid:9ccd9518) Host {"id":76,"name":"fakehost1.ha","type":"Routing","uuid":"ad118d97-03e2-4555-9d70-595864e871fd"} is 
not active according to neighbor Host {"id":75,"name":"fakehost2.ha","type":"Routing","uuid":"ffea4f8e-0d91-4a22-b600-aea1891be419"}, details: Heart is not beating.
2026-06-29 10:01:23,918 DEBUG [o.a.c.k.h.KVMHostActivityChecker] (AgentTaskPool-3:[ctx-5facfbaf]) (logid:9ccd9518) HA: HOST is ineligible legacy state Disconnected for host Host {"id":76,"name":"fakehost1.ha","type":"R
outing","uuid":"ad118d97-03e2-4555-9d70-595864e871fd"}
2026-06-29 10:01:23,918 INFO  [c.c.a.m.ClusteredAgentManagerImpl] (AgentTaskPool-3:[ctx-5facfbaf]) (logid:9ccd9518) The agent from host Host {"id":76,"name":"fakehost1.ha","type":"Routing","uuid":"ad118d97-03e2-4555-9d
70-595864e871fd"} state determined is Disconnected
2026-06-29 10:01:23,918 WARN  [c.c.a.m.ClusteredAgentManagerImpl] (AgentTaskPool-3:[ctx-5facfbaf]) (logid:9ccd9518) Agent is disconnected but the host is still up: Host {"id":76,"name":"fakehost1.ha","type":"Routing","
uuid":"ad118d97-03e2-4555-9d70-595864e871fd"} state: Enabled

The management servers then tried to start the vms on another host

2026-06-29 10:01:54,745 DEBUG [o.a.c.f.j.i.AsyncJobManagerImpl$5] (Work-Job-Executor-7:[ctx-fa35a62a, job-21199/job-21200]) (logid:28a7d0f4) Executing AsyncJob {"accountId":1,"cmd":"com.cloud.vm.VmWorkStart","cmdInfo":"rO0ABXNyABhjb2
0uY2xvdWQudm0uVm1Xb3JrU3RhcnR9cMGsvxz73gIAC0oABGRjSWRMAAZhdm9pZHN0ADBMY29tL2Nsb3VkL2RlcGxveS9EZXBsb3ltZW50UGxhbm5lciRFeGNsdWRlTGlzdDtMAAljbHVzdGVySWR0ABBMamF2YS9sYW5nL0xvbmc7TAAGaG9zdElkcQB-AAJMAAtqb3VybmFsTmFtZXQAEkxqYXZhL2xhbmcvU3R
yaW5nO0wAEXBoeXNpY2FsTmV0d29ya0lkcQB-AAJMAAdwbGFubmVycQB-AANMAAVwb2RJZHEAfgACTAAGcG9vbElkcQB-AAJMAAlyYXdQYXJhbXN0AA9MamF2YS91dGlsL01hcDtMAA1yZXNlcnZhdGlvbklkcQB-AAN4cgATY29tLmNsb3VkLnZtLlZtV29ya5-ZtlbwJWdrAgAESgAJYWNjb3VudElkSgAGdXNl
cklkSgAEdm1JZEwAC2hhbmRsZXJOYW1lcQB-AAN4cAAAAAAAAAABAAAAAAAAAAEAAAAAAAAFy3QAGVZpcnR1YWxNYWNoaW5lTWFuYWdlckltcGwAAAAAAAAAAHBwcHBwcHBwc3IAEWphdmEudXRpbC5IYXNoTWFwBQfawcMWYNEDAAJGAApsb2FkRmFjdG9ySQAJdGhyZXNob2xkeHA_QAAAAAAADHcIAAAAEAAAA
AF0AAtIYU9wZXJhdGlvbnQAP3JPMEFCWE55QUJGcVlYWmhMbXhoYm1jdVFtOXZiR1ZoYnMwZ2NvRFZuUHJ1QWdBQldnQUZkbUZzZFdWNGNBRXhw","cmdVersion":0,"completeMsid":null,"created":"Mon Jun 29 10:01:52 CEST 2026","id":21200,"initMsid":90520741747699,"insta
nceId":null,"instanceType":null,"lastPolled":null,"lastUpdated":null,"processStatus":0,"removed":null,"result":null,"resultCode":0,"status":"IN_PROGRESS","userId":1,"uuid":"319b5aa3-a42e-42d6-a538-c9c172529d70"}
2026-06-29 10:01:54,788 INFO  [c.c.a.m.a.i.FirstFitRoutingAllocator] (Work-Job-Executor-7:[ctx-fa35a62a, job-21199/job-21200, ctx-60b8df98, FirstFitRoutingAllocator]) (logid:28a7d0f4)  Guest VM is requested with Custom[UEFI] Boot Typ
e false
2026-06-29 10:01:54,831 DEBUG [o.a.c.e.o.NetworkOrchestrator] (Work-Job-Executor-7:[ctx-fa35a62a, job-21199/job-21200, ctx-60b8df98]) (logid:28a7d0f4) Network Network {"id": 347, "name": "DR-FE_11NET", "uuid": "48655ada-11ad-4101-abc
b-1531f6575018", "networkofferingid": 32} is already implemented
2026-06-29 10:01:54,842 DEBUG [o.a.c.e.o.NetworkOrchestrator] (Work-Job-Executor-7:[ctx-fa35a62a, job-21199/job-21200, ctx-60b8df98]) (logid:28a7d0f4) Changing active number of NICs for Network ID=Network {"id": 347, "name": "DR-FE_1
1NET", "uuid": "48655ada-11ad-4101-abcb-1531f6575018", "networkofferingid": 32} on 1
2026-06-29 10:01:54,849 DEBUG [o.a.c.e.o.VolumeOrchestrator] (Work-Job-Executor-7:[ctx-fa35a62a, job-21199/job-21200, ctx-60b8df98]) (logid:28a7d0f4) No need to recreate the volume [{"name":"ROOT-1483","uuid":"37824f75-50a3-4dea-9062
-f0ee8cde341c"}] since it already has an assigned pool: [0f61b84c-4c67-3856-8e16-eb920a49c200]. Adding disk to the VM.
.....
2026-06-29 10:01:55,976 DEBUG [o.a.c.e.o.NetworkOrchestrator] (Work-Job-Executor-7:[ctx-fa35a62a, job-21199/job-21200, ctx-60b8df98]) (logid:28a7d0f4) The nic Nic {"broadcastUri":"vlan:\/\/20","deviceId":0,"iPv4Address":null,"id":264
6,"instanceId":1483,"reservationId":"22066ce0-4ddd-4d9a-9fa3-c1d7c611ba2d","uuid":"9a4c66dc-e3f9-484f-a1af-2817ebfba998"} on NicProfile {"broadcastUri":null,"deviceId":0,"iPv4Address":null,"id":2646,"reservationId":"22066ce0-4ddd-4d9
a-9fa3-c1d7c611ba2d","uuid":"9a4c66dc-e3f9-484f-a1af-2817ebfba998","vmId":1483} was released according to VM instance {"id":1483,"instanceName":"i-2-1483-VM","state":"Starting","type":"User","uuid":"1f7afd4f-398b-4789-ae14-73d28903f0
36"} by guru com.cloud.network.guru.ExternalGuestNetworkGuru@601e3c8a, now updating record.
2026-06-29 10:01:55,977 DEBUG [o.a.c.e.o.NetworkOrchestrator] (Work-Job-Executor-7:[ctx-fa35a62a, job-21199/job-21200, ctx-60b8df98]) (logid:28a7d0f4) Changing active number of NICs for Network ID=Network {"id": 347, "name": "DR-FE_1
1NET", "uuid": "48655ada-11ad-4101-abcb-1531f6575018", "networkofferingid": 32} on -1
2026-06-29 10:01:56,038 WARN  [c.c.d.DeploymentPlanningManagerImpl] (Work-Job-Executor-7:[ctx-fa35a62a, job-21199/job-21200, ctx-60b8df98]) (logid:28a7d0f4) The last host [Host {"id":76,"name":"fakehost1.ha","type":"Ro
uting","uuid":"ad118d97-03e2-4555-9d70-595864e871fd"}] of VM [VM instance {"id":1483,"instanceName":"i-2-1483-VM","state":"Starting","type":"User","uuid":"1f7afd4f-398b-4789-ae14-73d28903f036"}] is in the avoid set. Skipping this and
 trying other available hosts.
....
2026-06-29 10:01:56,189 INFO  [o.a.c.s.a.ZoneWideStoragePoolAllocator] (Work-Job-Executor-7:[ctx-fa35a62a, job-21199/job-21200, ctx-60b8df98]) (logid:28a7d0f4) Using volume allocation algorithm firstfit to reorder pools.
2026-06-29 10:01:56,221 ERROR [c.c.v.ClusteredVirtualMachineManagerImpl] (Work-Job-Executor-7:[ctx-fa35a62a, job-21199/job-21200, ctx-60b8df98]) (logid:28a7d0f4) Unable to orchestrate start VM instance {"id":1483,"instanceName":"i-2-
1483-VM","state":"Stopped","type":"User","uuid":"1f7afd4f-398b-4789-ae14-73d28903f036"} due to [Unable to create a deployment for VM instance {"id":1483,"instanceName":"i-2-1483-VM","state":"Starting","type":"User","uuid":"1f7afd4f-3
98b-4789-ae14-73d28903f036"} after 3 attempts Last known error: Unable to start VM on Host {"id":75,"name":"fakehost2.ha","type":"Routing","uuid":"ffea4f8e-0d91-4a22-b600-aea1891be419"} due to internal error: QEMU unex
pectedly closed the monitor (vm='i-2-1483-VM'): 2026-06-29T08:01:55.363939Z qemu-kvm: -blockdev {"node-name":"libvirt-2-format","read-only":false,"discard":"unmap","cache":{"direct":true,"no-flush":false},"driver":"qcow2","file":"lib
virt-2-storage","backing":null}: Failed to get "write" lock
2026-06-29 10:01:56,249 WARN  [c.c.h.HighAvailabilityManagerExtImpl] (HA-Worker-1:[ctx-9f19d45c, work-17]) (logid:9228a43f) Encountered unhandled exception during HA process, reschedule work HAWork[17-HA-1483-Stopped-Scheduled] com.cloud.utils.exception.CloudRuntimeException: Unable to orchestrate the start of VM instance {"instanceName":"i-2-1483-VM","uuid":"1f7afd4f-398b-4789-ae14-73d28903f036"}.
        at com.cloud.vm.VirtualMachineManagerImpl.orchestrateStart(VirtualMachineManagerImpl.java:6012)
        at java.base/jdk.internal.reflect.NativeMethodAccessorImpl.invoke0(Native Method)
        at java.base/jdk.internal.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:77)
        at java.base/jdk.internal.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)
        at java.base/java.lang.reflect.Method.invoke(Method.java:569)
        at com.cloud.vm.VmWorkJobHandlerProxy.handleVmWorkJob(VmWorkJobHandlerProxy.java:102)
        at com.cloud.vm.VirtualMachineManagerImpl.handleVmWorkJob(VirtualMachineManagerImpl.java:6133)
        at com.cloud.vm.VmWorkJobDispatcher.runJob(VmWorkJobDispatcher.java:99)
        at org.apache.cloudstack.framework.jobs.impl.AsyncJobManagerImpl$5.runInContext(AsyncJobManagerImpl.java:698)
        at org.apache.cloudstack.managed.context.ManagedContextRunnable$1.run(ManagedContextRunnable.java:49)
        at org.apache.cloudstack.managed.context.impl.DefaultManagedContext$1.call(DefaultManagedContext.java:56)
        at org.apache.cloudstack.managed.context.impl.DefaultManagedContext.callWithContext(DefaultManagedContext.java:103)
        at org.apache.cloudstack.managed.context.impl.DefaultManagedContext.runWithContext(DefaultManagedContext.java:53)
        at org.apache.cloudstack.managed.context.ManagedContextRunnable.run(ManagedContextRunnable.java:46)
        at org.apache.cloudstack.framework.jobs.impl.AsyncJobManagerImpl$5.run(AsyncJobManagerImpl.java:646)
        at java.base/java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:539)
        at java.base/java.util.concurrent.FutureTask.run(FutureTask.java:264)
        at java.base/java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136)
        at java.base/java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635)
        at java.base/java.lang.Thread.run(Thread.java:840)

2026-06-29 10:01:56,250 WARN  [c.c.h.HighAvailabilityManagerExtImpl] (HA-Worker-1:[ctx-9f19d45c, work-17]) (logid:9228a43f) Rescheduling work HAWork[17-HA-1483-Stopped-Scheduled] to try again at 2026-06-29T10:02:26.432+0200. Finished attempt 1/5 times.

The job fails (expectedly i would say) because of the nfs write lock held from the dead host.
But when retrying the vm restart we get

2026-06-29 10:01:52,843 INFO  [c.c.h.HighAvailabilityManagerExtImpl] (HA-Worker-1:[ctx-9f19d45c, work-17]) (logid:9228a43f) Processing work HAWork[17-HA-1483-Stopped-Scheduled]
2026-06-29 10:01:52,845 INFO  [c.c.h.HighAvailabilityManagerExtImpl] (HA-Worker-1:[ctx-9f19d45c, work-17]) (logid:9228a43f) HA on VM instance {"id":1483,"instanceName":"i-2-1483-VM","state":"Stopped","type":"User","uuid":"1f7afd4f-398b-4789-ae14-73d28903f036"}
2026-06-29 10:01:52,860 WARN  [o.a.c.f.j.AsyncJobExecutionContext] (HA-Worker-1:[ctx-9f19d45c, work-17, ctx-b9bac272]) (logid:9228a43f) Job is executed without a context, setup psudo job for the executing thread
2026-06-29 10:01:52,871 DEBUG [o.a.c.f.j.i.AsyncJobManagerImpl] (HA-Worker-1:[ctx-9f19d45c, work-17, ctx-b9bac272]) (logid:9228a43f) Sync job-21200 execution on object VmWorkJobQueue.1483
2026-06-29 10:01:56,249 WARN  [c.c.h.HighAvailabilityManagerExtImpl] (HA-Worker-1:[ctx-9f19d45c, work-17]) (logid:9228a43f) Encountered unhandled exception during HA process, reschedule work HAWork[17-HA-1483-Stopped-Scheduled] com.cloud.utils.exception.CloudRuntimeException: Unable to orchestrate the start of VM instance {"instanceName":"i-2-1483-VM","uuid":"1f7afd4f-398b-4789-ae14-73d28903f036"}.
2026-06-29 10:01:56,250 WARN  [c.c.h.HighAvailabilityManagerExtImpl] (HA-Worker-1:[ctx-9f19d45c, work-17]) (logid:9228a43f) Rescheduling work HAWork[17-HA-1483-Stopped-Scheduled] to try again at 2026-06-29T10:02:26.432+0200. Finished attempt 1/5 times.
2026-06-29 10:02:54,848 INFO  [c.c.h.HighAvailabilityManagerExtImpl] (HA-Worker-0:[ctx-fce5735e, work-17]) (logid:f2990b6c) Processing work HAWork[17-HA-1483-Stopped-Scheduled]
2026-06-29 10:02:54,852 INFO  [c.c.h.HighAvailabilityManagerExtImpl] (HA-Worker-0:[ctx-fce5735e, work-17]) (logid:f2990b6c) HA on VM instance {"id":1483,"instanceName":"i-2-1483-VM","state":"Stopped","type":"User","uuid":"1f7afd4f-398b-4789-ae14-73d28903f036"}
2026-06-29 10:02:54,852 INFO  [c.c.h.HighAvailabilityManagerExtImpl] (HA-Worker-0:[ctx-fce5735e, work-17]) (logid:f2990b6c) VM VM instance {"id":1483,"instanceName":"i-2-1483-VM","state":"Stopped","type":"User","uuid":"1f7afd4f-398b-4789-ae14-73d28903f036"} has been changed.  Current State = Stopped Previous State = Stopped last updated = 20 previous updated = 17
2026-06-29 10:02:54,852 INFO  [c.c.h.HighAvailabilityManagerExtImpl] (HA-Worker-0:[ctx-fce5735e, work-17]) (logid:f2990b6c) Completed work HAWork[17-HA-1483-Stopped-Scheduled]. Took 2/5 attempts.

The restart process didn't even try a second time because of the mismatch in the update counter between the vm and the work (20 vs 17). The final state of the vm (without manual intervention) is Stopped.

versions

Cloudstack version: 4.22.1.0, with management server running on ubuntu22 and agent running on oracle linux 9
Shared nfsv4 (4.2) storage

The steps to reproduce the bug

  1. Disable HOST HA
  2. Enable VM HA on some vm
  3. Shutdown forcefully the host with aforementioned vm and check for nfs locks

What to do about it?

No response

Activity

  1. kiranchavala commented on Jun 29, 2026

    @kiranchavala
    Member

    @gusmef

    I had introduced the kernel panic with the following command "echo c > /proc/sysrq-trigger"

    The kvm host went to alert state and no ha entry was created for the vm under "select * from op_ha_work"

    Image
    2026-06-29 12:34:09,322 INFO  [c.c.a.m.ClusteredAgentManagerImpl] (AgentTaskPool-1:[ctx-14c00e81]) (logid:d84d661c) The agent from host Host {"id":1,"name":"Cloudstack-Kvm-Host1","type":"Routing","uuid":"c3b12744-eaa4-44db-b261-3b548fbfb49b"} state determined is Disconnected
    2026-06-29 12:34:09,322 WARN  [c.c.a.m.ClusteredAgentManagerImpl] (AgentTaskPool-1:[ctx-14c00e81]) (logid:d84d661c) Agent is disconnected but the host is still up: Host {"id":1,"name":"Cloudstack-Kvm-Host1","type":"Routing","uuid":"c3b12744-eaa4-44db-b261-3b548fbfb49b"} state: Enabled
    2026-06-29 12:34:09,333 WARN  [c.c.a.AlertManagerImpl] (AgentTaskPool-1:[ctx-14c00e81]) (logid:d84d661c) No recipients set in global setting 'alert.email.addresses', skipping sending alert with subject [Host disconnected, name: Cloudstack-Kvm-Host1 (id:c3b12744-eaa4-44db-b261-3b548fbfb49b), availability zone: kiran-home-lab-zone1, pod: pod1] and content [If the agent for host [name: Cloudstack-Kvm-Host1 (id:c3b12744-eaa4-44db-b261-3b548fbfb49b), availability zone: kiran-home-lab-zone1, pod: pod1] is not restarted within alert.wait seconds, host will go to Alert state].
    2026-06-29 12:34:09,334 DEBUG [c.c.a.m.ClusteredAgentManagerImpl] (AgentTaskPool-1:[ctx-14c00e81]) (logid:d84d661c) Deregistering link for AgentAttache {"_id":1,"_name":"Cloudstack-Kvm-Host1","_uuid":"c3b12744-eaa4-44db-b261-3b548fbfb49b"} with state Alert
    
    

    With the fix #13373

    The host goes into down state and the vm ha is triggered

    Could you please provide the command you used to trigger the kernel panic

    and also the values of the global settings

    
    "commands.timeout" = 
    
    
    
  2. gusmef commented on Jun 29, 2026

    @gusmef
    Author

    Hi @kiranchavala , the kernel panic was triggered in the exact same way you did it
    "echo c > /proc/sysrq-trigger"
    Here's the requested setting:

    Image

    To clarify, the issue is not that HA fails to trigger; the logs clearly show it attempting to restart the VM on another host. The actual problem is that after the first attempt fails, the process stops after the second try and gives up entirely, even though the VM remains offline.

    2026-06-29 10:02:54,852 INFO  [c.c.h.HighAvailabilityManagerExtImpl] (HA-Worker-0:[ctx-fce5735e, work-17]) (logid:f2990b6c) VM VM instance {"id":1483,"instanceName":"i-2-1483-VM","state":"Stopped","type":"User","uuid":"1f7afd4f-398b-4789-ae14-73d28903f036"} has been changed.  Current State = Stopped Previous State = Stopped last updated = 20 previous updated = 17
    2026-06-29 10:02:54,852 INFO  [c.c.h.HighAvailabilityManagerExtImpl] (HA-Worker-0:[ctx-fce5735e, work-17]) (logid:f2990b6c) Completed work HAWork[17-HA-1483-Stopped-Scheduled]. Took 2/5 attempts.
    

    Let me know if you need anything else or if there's something you think i'm doing wrong

  3. kiranchavala commented on Jun 29, 2026

    @kiranchavala
    Member

    @gusmef thanks for the update

    In my lab , the HA doesn't start for the vm and host stays in alert state

    What is the host status in your case ?

    Also what is value of the global setting

    force.ha
    kvm.ha.fence.on.storage.heartbeat.failure

  4. gusmef commented on Jun 29, 2026

    @gusmef
    Author

    @kiranchavala the global settings are
    force.ha: false
    kvm.ha.fence.on.storage.heartbeat.failure: false
    We do have
    force.ha: true
    in the cluster with the 2 hosts where we're trying the HA.

    The host status was "Alert". I think we somehow tricked the HA into starting by declaring the "dead" host as degraded (manually from the web gui).

  5. muthukrishnang1100 commented on Jun 30, 2026

    @muthukrishnang1100

    Hi All, I am also facing the same issue for my cloudstack environment 4.20.3.0 with ceph storage. Before, VM HA will not work after for the heartbeat detection i added the NFS storage for this, Now VM HA will work, But IF host have an 20 VMs most of the VMs are restarted to another hosts, But few VMs are stucked at stopping state, same I checked DB it shows the HA done not going to take second retry, But still those VMs are stucked for stopping state, Need to recover manually.

    Please check for this

  6. kiranchavala commented on Jun 30, 2026

    @kiranchavala
    Member

    @gusmef @muthukrishnang1100

    could you please let us know the nfs version and qemu versions from the kvm host

    
    nfsstat -m
    qemu-system-x86_64 --version
    
  7. kiranchavala commented on Jun 30, 2026

    @kiranchavala
    Member

    @gusmef @muthukrishnang1100

    According to this comment there are issues with the recent qemu and nfs version

    #10690 (comment)

  8. gusmef commented on Jun 30, 2026

    @gusmef
    Author

    Hi @kiranchavala

    /usr/libexec/qemu-kvm --version
    QEMU emulator version 10.1.0 (qemu-kvm-10.1.0-17.el9_8)
    Copyright (c) 2003-2025 Fabrice Bellard and the QEMU Project developers
    
    nfsstat -m
    Flags:
    rw,sync,nosuid,nodev,noexec,relatime,vers=4.2,rsize=65536,wsize=65536,namlen=255,acregmin=0,acregmax=0,acdirmin=0,acdirmax=0,hard,noac,proto=tcp,nconnect=8,timeo=600,retrans=2,sec=sys,clientaddr=x.x.x.x,local_lock=none,addr=y.y.y.y
    

    Regarding the comment shared, i don't think it's the exact same situation, the first attempt to migrate fails (and that's ok, there's still the nfs lock from the dead host) but then cloudstack should try again to migrate the vm, and it doesn't. It stops at attempt 2/5 because there's a mismatch between the work.updated and the vm.updated counters, but no one updated manually the vm.

  9. muthukrishnang1100 commented on Jun 30, 2026

    @muthukrishnang1100

    @gusmef @muthukrishnang1100

    could you please let us know the nfs version and qemu versions from the kvm host

    
    nfsstat -m
    qemu-system-x86_64 --version
    

    But I am using ceph for all VMs . NFS is only for an heartbeat. I have used one my management server to create NFS mount and connected to all the compute hosts.

  10. muthukrishnang1100 commented on Jun 30, 2026

    @muthukrishnang1100

    @kiranchavala
    See I share here my 3 management servers and also first 2 compute hosts all my remaining compute hosts are same.

    root@ltchl1pacsmgm01:~#
    nfsstat -m
    qemu-system-x86_64 --version
    /var/lib/cloudstack/mnt/222786140395859.3d9e9f1f from 10.29.40.12:/cloudstack-secondary
    Flags: rw,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,fatal_neterrors=none,proto=tcp,timeo=600,retrans=2,sec=sys,clientaddr=10.29.20.4,local_lock=none,addr=10.29.40.12

    Command 'qemu-system-x86_64' not found, but can be installed with:
    apt install qemu-system-x86
    root@ltchl1pacsmgm01:~#

    root@ltchl1pacsmgm02:~#
    nfsstat -m
    qemu-system-x86_64 --version
    /var/lib/cloudstack/mnt/116713698675123.11d231ec from 10.29.40.12:/cloudstack-secondary
    Flags: rw,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,proto=tcp,timeo=600,retrans=2,sec=sys,clientaddr=10.29.20.5,local_lock=none,addr=10.29.40.12

    -bash: qemu-system-x86_64: command not found
    root@ltchl1pacsmgm02:~#

    root@ltchl1pacsmgm03:~#
    nfsstat -m
    qemu-system-x86_64 --version
    /var/lib/cloudstack/mnt/244360711989334.f57ae42 from 10.29.40.12:/cloudstack-secondary
    Flags: rw,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,proto=tcp,timeo=600,retrans=2,sec=sys,clientaddr=10.29.20.6,local_lock=none,addr=10.29.40.12

    -bash: qemu-system-x86_64: command not found
    root@ltchl1pacsmgm03:~#

    root@ltchl1pacscom01:~#
    nfsstat -m
    qemu-system-x86_64 --version
    /mnt/4e5b0e96-d38a-3822-a69e-40945154553b from 10.29.20.4:/var/cloudstack-ha-heartbeat
    Flags: rw,nosuid,nodev,noexec,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,proto=tcp,timeo=600,retrans=2,sec=sys,clientaddr=10.29.30.4,local_lock=none,addr=10.29.20.4

    QEMU emulator version 8.2.2 (Debian 1:8.2.2+ds-0ubuntu1.16)
    Copyright (c) 2003-2023 Fabrice Bellard and the QEMU Project developers
    root@ltchl1pacscom01:~#

    root@ltchl1pacscom02:~# nfsstat -m
    qemu-system-x86_64 --version
    /mnt/4e5b0e96-d38a-3822-a69e-40945154553b from 10.29.20.4:/var/cloudstack-ha-heartbeat
    Flags: rw,nosuid,nodev,noexec,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,proto=tcp,timeo=600,retrans=2,sec=sys,clientaddr=10.29.30.5,local_lock=none,addr=10.29.20.4

    QEMU emulator version 8.2.2 (Debian 1:8.2.2+ds-0ubuntu1.16)
    Copyright (c) 2003-2023 Fabrice Bellard and the QEMU Project developers
    root@ltchl1pacscom02:~#

    root@ltchl1pacscom02:~# nfsstat -m
    qemu-system-x86_64 --version
    /mnt/4e5b0e96-d38a-3822-a69e-40945154553b from 10.29.20.4:/var/cloudstack-ha-heartbeat
    Flags: rw,nosuid,nodev,noexec,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,proto=tcp,timeo=600,retrans=2,sec=sys,clientaddr=10.29.30.5,local_lock=none,addr=10.29.20.4

    QEMU emulator version 8.2.2 (Debian 1:8.2.2+ds-0ubuntu1.16)
    Copyright (c) 2003-2023 Fabrice Bellard and the QEMU Project developers
    root@ltchl1pacscom02:~#

  11. muthukrishnang1100 commented on Jun 30, 2026

    @muthukrishnang1100

    Hi @kiranchavala
    Thank you for checking.
    Our environment is different from the original NFS-primary-storage case.
    Environment
    Apache CloudStack: 4.20.3.0
    Hypervisor: KVM on Ubuntu 24.04
    Management servers: 3-node management cluster
    Database: 3-node MariaDB Galera cluster behind HAProxy
    VM primary storage: Ceph RBD
    Ceph pool for VM disks: cloudstack-primary
    Ceph health during the test: HEALTH_OK
    QEMU version on all compute hosts:

    QEMU emulator version 8.2.2
    (Debian 1:8.2.2+ds-0ubuntu1.16)
    

    Important: NFS is not used as VM primary storage in our environment.
    NFS usage
    NFSv4.2 is used only for the KVM Host HA heartbeat path.
    Example heartbeat mount from compute hosts:

    /mnt/<heartbeat-mount-id> from <management-server-ip>:/var/cloudstack-ha-heartbeat
    Flags: rw,nosuid,nodev,noexec,relatime,vers=4.2,rsize=1048576,wsize=1048576,hard,proto=tcp,timeo=600,retrans=2,sec=sys,local_lock=none
    

    We also use NFSv4.2 for CloudStack secondary storage. However, all user VM root and data disks are on Ceph RBD.
    Before adding the dedicated NFS heartbeat storage, Host HA did not trigger correctly with Ceph RBD primary storage. After configuring the NFS heartbeat storage, Host HA detection and fencing are working correctly.
    Relevant HA settings

    vm.ha.enabled = true
    vm.ha.alerts.enabled = true
    ha.workers = 5
    
    enable.ha.storage.migration = true
    ha.fence.builders.exclude = RecreatableFencer
    ha.investigators.order = KVMInvestigator, PingInvestigator, SimpleInvestigator
    
    ha.max.concurrent.activity.check.operations = 25
    ha.max.concurrent.fence.operations = 25
    ha.max.concurrent.health.check.operations = 50
    ha.max.concurrent.recovery.operations = 25
    
    kvm.ha.activity.check.failure.ratio = 0.7
    kvm.ha.activity.check.interval = 60
    kvm.ha.activity.check.max.attempts = 10
    kvm.ha.activity.check.timeout = 60
    kvm.ha.degraded.max.period = 300
    
    kvm.ha.fence.on.storage.heartbeat.failure = true
    kvm.ha.fence.timeout = 60
    kvm.ha.health.check.timeout = 10
    kvm.ha.recover.failure.threshold = 1
    kvm.ha.recover.timeout = 60
    kvm.ha.recover.wait.period = 600
    
    ha.vm.restart.hostup = true
    vm.ha.migration.max.retries = 5
    

    Test performed
    We performed repeated ungraceful host-failure tests by hard-powering off a KVM compute host through OOBM/IPMI.
    Example failed host:

    Host name: ltchl1pacscom04
    CloudStack Host ID: 19
    Host IP: <redacted>
    

    At the time of one test, there were 9 user VMs running on this host.
    Expected behavior:

    Host Down
    -> fencing succeeds
    -> all VMs restart on another KVM host
    

    Actual behavior:

    Host Down
    -> KVM fencing succeeds
    -> most VMs restart successfully on other hosts
    -> a few VMs remain permanently in Stopping state
    -> HA work later becomes Done
    -> manual recovery is required
    

    For example, during one 9-VM test, 6 VMs restarted successfully and the following VMs remained in Stopping:

    i-22-248-VM
    i-22-257-VM
    i-22-258-VM
    

    Management-server log observations
    CloudStack correctly identifies the failed host as Down and KVM fencing succeeds:

    Fencer KVMFenceBuilder returned true
    

    For an affected VM, the HA stop/network cleanup workflow encountered a database transaction failure:

    Unable to release some network resources for the VM in Stopping state
    
    com.cloud.utils.exception.CloudRuntimeException:
    Unable to commit or close the connection.
    
    Caused by:
    com.mysql.cj.jdbc.exceptions.MySQLTransactionRollbackException:
    Deadlock found when trying to get lock; try restarting transaction
    

    For other affected VMs, HA logs:

    Encountered unhandled exception during HA process, reschedule work
    
    com.cloud.utils.exception.CloudRuntimeException:
    Unable to find by id on DB, due to:
    Deadlock found when trying to get lock; try restarting transaction
    

    The HA work is initially rescheduled. During the next HA attempt, CloudStack detects that the VM state has changed from Running to Stopping:

    HA on VM instance {... state="Stopping" ...}
    
    VM ... has been changed.
    Current State = Stopping
    Previous State = Running
    

    Then the HA work is completed:

    Completed work HAWork[...]
    

    However, the VM remains in Stopping and is not restarted on another host.
    Database observation
    For the affected VMs, the database shows:

    vm_instance.state = Stopping
    vm_instance.host_id = failed host ID
    
    Latest op_ha_work:
    step = Done
    

    The MariaDB Galera cluster was healthy after the test:

    wsrep_cluster_status      = Primary
    wsrep_cluster_size        = 3
    wsrep_connected           = ON
    wsrep_ready               = ON
    wsrep_local_state_comment = Synced
    

    Current counters on one database node were:

    Innodb_deadlocks          = 0
    wsrep_local_bf_aborts     = 68
    wsrep_local_cert_failures = 75
    wsrep_local_replays       = 4
    

    These Galera counters are cumulative and were not captured before and immediately after every HA test, so we cannot confirm that they are directly caused by this specific HA event. However, the CloudStack management-server logs clearly show MySQLTransactionRollbackException: Deadlock found when trying to get lock during the affected HA recovery workflow.
    Ceph RBD lock observation
    For VMs that restarted successfully, the RBD exclusive locks moved from the failed host to the destination compute hosts.
    For VMs stuck in Stopping, their root-volume RBD locks remained on the failed source host. Example:

    rbd lock ls cloudstack-primary/<rbd-image>
    
    There is 1 exclusive lock on this image.
    Locker          ID                    Address
    client.<id>     auto <id>             <failed-compute-host-ip>:0/<client-session>
    

    Since the source host was hard powered off, this stale RBD lock is expected. However, CloudStack does not complete the HA recovery because the VM remains in Stopping and the HA work becomes Done.
    Comparison with issue #13510
    I understand that the original issue uses shared NFSv4.2 primary storage and the first restart failure is caused by a QEMU/NFS write-lock error.
    Our initial HA failure appears different:

    Issue #13510:
    QEMU cannot acquire NFS write lock
    -> initial restart fails
    -> HA retry detects update counter mismatch
    -> HA work completes without retry
    
    Our environment:
    CloudStack HA stop/network cleanup encounters DB transaction deadlock
    -> VM remains Stopping
    -> HA retry detects changed VM state/update state
    -> HA work completes without successful restart
    

    The storage backend is different, but the final behavior appears similar: after an initial HA recovery failure, CloudStack stops retrying and completes the HA work even though the VM remains unavailable.

  12. kiranchavala commented on Jun 30, 2026

    @kiranchavala
    Member
  13. muthukrishnang1100 commented on Jul 6, 2026

    @muthukrishnang1100

    @kiranchavala Now, I solved that issue. I added the NFS primary storage only for KVMHA Heartbeat. After that still Host HA is not working, Then we disable the Host HA and Only worked now for VM HA.
    But, some of them VMs are stucked at the stopping state, we need to manual recover is needed, stop and start the VMs GUI then they all are running.

    I tried and many thing for this stucked VMs how to recover but we cant able to do that? Because, Many private cloud have an same issue during HA some VMs are stucked

  14. muthukrishnang1100 commented on Jul 13, 2026

    @muthukrishnang1100

    @kiranchavala Any update on this?

  15. kiranchavala commented on Jul 14, 2026

    @kiranchavala
    Member

    @muthukrishnang1100 I am not able to reproduce the issue in my lab on the recent build of Cloudstack 4.22.1 with nfs storage

    If you can reproduce the issue consistently on 4.22.1. Please let us know

    http://packages.shapeblue.com/cloudstack/upstream/
    http://packages.shapeblue.com/cloudstack/upstream/debian/4.22/

    Also I am not sure if galera mysql is offically supported with cloudsack

    https://docs.cloudstack.apache.org/en/4.22.1.0/releasenotes/compat.html#software-requirements

  16. muthukrishnang1100 commented on Jul 15, 2026

    @muthukrishnang1100

    @kiranchavala But, I am using Ceph and NFS for HA heartbeat only.

  17. locked and limited conversation to collaborators on Jul 28, 2026
  18. converted this issue into a discussion #13729 on Jul 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions