Repository navigation
VM HA not working when there's an issue with vm restart #13510
Description
Activity
I had introduced the kernel panic with the following command "echo c > /proc/sysrq-trigger"
The kvm host went to alert state and no ha entry was created for the vm under "select * from op_ha_work"
2026-06-29 12:34:09,322 INFO [c.c.a.m.ClusteredAgentManagerImpl] (AgentTaskPool-1:[ctx-14c00e81]) (logid:d84d661c) The agent from host Host {"id":1,"name":"Cloudstack-Kvm-Host1","type":"Routing","uuid":"c3b12744-eaa4-44db-b261-3b548fbfb49b"} state determined is Disconnected 2026-06-29 12:34:09,322 WARN [c.c.a.m.ClusteredAgentManagerImpl] (AgentTaskPool-1:[ctx-14c00e81]) (logid:d84d661c) Agent is disconnected but the host is still up: Host {"id":1,"name":"Cloudstack-Kvm-Host1","type":"Routing","uuid":"c3b12744-eaa4-44db-b261-3b548fbfb49b"} state: Enabled 2026-06-29 12:34:09,333 WARN [c.c.a.AlertManagerImpl] (AgentTaskPool-1:[ctx-14c00e81]) (logid:d84d661c) No recipients set in global setting 'alert.email.addresses', skipping sending alert with subject [Host disconnected, name: Cloudstack-Kvm-Host1 (id:c3b12744-eaa4-44db-b261-3b548fbfb49b), availability zone: kiran-home-lab-zone1, pod: pod1] and content [If the agent for host [name: Cloudstack-Kvm-Host1 (id:c3b12744-eaa4-44db-b261-3b548fbfb49b), availability zone: kiran-home-lab-zone1, pod: pod1] is not restarted within alert.wait seconds, host will go to Alert state]. 2026-06-29 12:34:09,334 DEBUG [c.c.a.m.ClusteredAgentManagerImpl] (AgentTaskPool-1:[ctx-14c00e81]) (logid:d84d661c) Deregistering link for AgentAttache {"_id":1,"_name":"Cloudstack-Kvm-Host1","_uuid":"c3b12744-eaa4-44db-b261-3b548fbfb49b"} with state AlertWith the fix #13373
The host goes into down state and the vm ha is triggered
Could you please provide the command you used to trigger the kernel panic
and also the values of the global settings
"commands.timeout" =Hi @kiranchavala , the kernel panic was triggered in the exact same way you did it
"echo c > /proc/sysrq-trigger"
Here's the requested setting:
To clarify, the issue is not that HA fails to trigger; the logs clearly show it attempting to restart the VM on another host. The actual problem is that after the first attempt fails, the process stops after the second try and gives up entirely, even though the VM remains offline.
2026-06-29 10:02:54,852 INFO [c.c.h.HighAvailabilityManagerExtImpl] (HA-Worker-0:[ctx-fce5735e, work-17]) (logid:f2990b6c) VM VM instance {"id":1483,"instanceName":"i-2-1483-VM","state":"Stopped","type":"User","uuid":"1f7afd4f-398b-4789-ae14-73d28903f036"} has been changed. Current State = Stopped Previous State = Stopped last updated = 20 previous updated = 17 2026-06-29 10:02:54,852 INFO [c.c.h.HighAvailabilityManagerExtImpl] (HA-Worker-0:[ctx-fce5735e, work-17]) (logid:f2990b6c) Completed work HAWork[17-HA-1483-Stopped-Scheduled]. Took 2/5 attempts.Let me know if you need anything else or if there's something you think i'm doing wrong
@gusmef thanks for the update
In my lab , the HA doesn't start for the vm and host stays in alert state
What is the host status in your case ?
Also what is value of the global setting
force.ha
kvm.ha.fence.on.storage.heartbeat.failure@kiranchavala the global settings are
force.ha: false
kvm.ha.fence.on.storage.heartbeat.failure: false
We do have
force.ha: true
in the cluster with the 2 hosts where we're trying the HA.The host status was "Alert". I think we somehow tricked the HA into starting by declaring the "dead" host as degraded (manually from the web gui).
Hi All, I am also facing the same issue for my cloudstack environment 4.20.3.0 with ceph storage. Before, VM HA will not work after for the heartbeat detection i added the NFS storage for this, Now VM HA will work, But IF host have an 20 VMs most of the VMs are restarted to another hosts, But few VMs are stucked at stopping state, same I checked DB it shows the HA done not going to take second retry, But still those VMs are stucked for stopping state, Need to recover manually.
Please check for this
could you please let us know the nfs version and qemu versions from the kvm host
nfsstat -m qemu-system-x86_64 --versionAccording to this comment there are issues with the recent qemu and nfs version
/usr/libexec/qemu-kvm --version QEMU emulator version 10.1.0 (qemu-kvm-10.1.0-17.el9_8) Copyright (c) 2003-2025 Fabrice Bellard and the QEMU Project developers nfsstat -m Flags: rw,sync,nosuid,nodev,noexec,relatime,vers=4.2,rsize=65536,wsize=65536,namlen=255,acregmin=0,acregmax=0,acdirmin=0,acdirmax=0,hard,noac,proto=tcp,nconnect=8,timeo=600,retrans=2,sec=sys,clientaddr=x.x.x.x,local_lock=none,addr=y.y.y.yRegarding the comment shared, i don't think it's the exact same situation, the first attempt to migrate fails (and that's ok, there's still the nfs lock from the dead host) but then cloudstack should try again to migrate the vm, and it doesn't. It stops at attempt 2/5 because there's a mismatch between the work.updated and the vm.updated counters, but no one updated manually the vm.
could you please let us know the nfs version and qemu versions from the kvm host
nfsstat -m qemu-system-x86_64 --versionBut I am using ceph for all VMs . NFS is only for an heartbeat. I have used one my management server to create NFS mount and connected to all the compute hosts.
@kiranchavala
See I share here my 3 management servers and also first 2 compute hosts all my remaining compute hosts are same.root@ltchl1pacsmgm01:~#
nfsstat -m
qemu-system-x86_64 --version
/var/lib/cloudstack/mnt/222786140395859.3d9e9f1f from 10.29.40.12:/cloudstack-secondary
Flags: rw,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,fatal_neterrors=none,proto=tcp,timeo=600,retrans=2,sec=sys,clientaddr=10.29.20.4,local_lock=none,addr=10.29.40.12Command 'qemu-system-x86_64' not found, but can be installed with:
apt install qemu-system-x86
root@ltchl1pacsmgm01:~#root@ltchl1pacsmgm02:~#
nfsstat -m
qemu-system-x86_64 --version
/var/lib/cloudstack/mnt/116713698675123.11d231ec from 10.29.40.12:/cloudstack-secondary
Flags: rw,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,proto=tcp,timeo=600,retrans=2,sec=sys,clientaddr=10.29.20.5,local_lock=none,addr=10.29.40.12-bash: qemu-system-x86_64: command not found
root@ltchl1pacsmgm02:~#root@ltchl1pacsmgm03:~#
nfsstat -m
qemu-system-x86_64 --version
/var/lib/cloudstack/mnt/244360711989334.f57ae42 from 10.29.40.12:/cloudstack-secondary
Flags: rw,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,proto=tcp,timeo=600,retrans=2,sec=sys,clientaddr=10.29.20.6,local_lock=none,addr=10.29.40.12-bash: qemu-system-x86_64: command not found
root@ltchl1pacsmgm03:~#root@ltchl1pacscom01:~#
nfsstat -m
qemu-system-x86_64 --version
/mnt/4e5b0e96-d38a-3822-a69e-40945154553b from 10.29.20.4:/var/cloudstack-ha-heartbeat
Flags: rw,nosuid,nodev,noexec,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,proto=tcp,timeo=600,retrans=2,sec=sys,clientaddr=10.29.30.4,local_lock=none,addr=10.29.20.4QEMU emulator version 8.2.2 (Debian 1:8.2.2+ds-0ubuntu1.16)
Copyright (c) 2003-2023 Fabrice Bellard and the QEMU Project developers
root@ltchl1pacscom01:~#root@ltchl1pacscom02:~# nfsstat -m
qemu-system-x86_64 --version
/mnt/4e5b0e96-d38a-3822-a69e-40945154553b from 10.29.20.4:/var/cloudstack-ha-heartbeat
Flags: rw,nosuid,nodev,noexec,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,proto=tcp,timeo=600,retrans=2,sec=sys,clientaddr=10.29.30.5,local_lock=none,addr=10.29.20.4QEMU emulator version 8.2.2 (Debian 1:8.2.2+ds-0ubuntu1.16)
Copyright (c) 2003-2023 Fabrice Bellard and the QEMU Project developers
root@ltchl1pacscom02:~#root@ltchl1pacscom02:~# nfsstat -m
qemu-system-x86_64 --version
/mnt/4e5b0e96-d38a-3822-a69e-40945154553b from 10.29.20.4:/var/cloudstack-ha-heartbeat
Flags: rw,nosuid,nodev,noexec,relatime,vers=4.2,rsize=1048576,wsize=1048576,namlen=255,hard,proto=tcp,timeo=600,retrans=2,sec=sys,clientaddr=10.29.30.5,local_lock=none,addr=10.29.20.4QEMU emulator version 8.2.2 (Debian 1:8.2.2+ds-0ubuntu1.16)
Copyright (c) 2003-2023 Fabrice Bellard and the QEMU Project developers
root@ltchl1pacscom02:~#Hi @kiranchavala
Thank you for checking.
Our environment is different from the original NFS-primary-storage case.
Environment
Apache CloudStack:4.20.3.0
Hypervisor: KVM on Ubuntu 24.04
Management servers: 3-node management cluster
Database: 3-node MariaDB Galera cluster behind HAProxy
VM primary storage: Ceph RBD
Ceph pool for VM disks:cloudstack-primary
Ceph health during the test:HEALTH_OK
QEMU version on all compute hosts:QEMU emulator version 8.2.2 (Debian 1:8.2.2+ds-0ubuntu1.16)Important: NFS is not used as VM primary storage in our environment.
NFS usage
NFSv4.2 is used only for the KVM Host HA heartbeat path.
Example heartbeat mount from compute hosts:/mnt/<heartbeat-mount-id> from <management-server-ip>:/var/cloudstack-ha-heartbeat Flags: rw,nosuid,nodev,noexec,relatime,vers=4.2,rsize=1048576,wsize=1048576,hard,proto=tcp,timeo=600,retrans=2,sec=sys,local_lock=noneWe also use NFSv4.2 for CloudStack secondary storage. However, all user VM root and data disks are on Ceph RBD.
Before adding the dedicated NFS heartbeat storage, Host HA did not trigger correctly with Ceph RBD primary storage. After configuring the NFS heartbeat storage, Host HA detection and fencing are working correctly.
Relevant HA settingsvm.ha.enabled = true vm.ha.alerts.enabled = true ha.workers = 5 enable.ha.storage.migration = true ha.fence.builders.exclude = RecreatableFencer ha.investigators.order = KVMInvestigator, PingInvestigator, SimpleInvestigator ha.max.concurrent.activity.check.operations = 25 ha.max.concurrent.fence.operations = 25 ha.max.concurrent.health.check.operations = 50 ha.max.concurrent.recovery.operations = 25 kvm.ha.activity.check.failure.ratio = 0.7 kvm.ha.activity.check.interval = 60 kvm.ha.activity.check.max.attempts = 10 kvm.ha.activity.check.timeout = 60 kvm.ha.degraded.max.period = 300 kvm.ha.fence.on.storage.heartbeat.failure = true kvm.ha.fence.timeout = 60 kvm.ha.health.check.timeout = 10 kvm.ha.recover.failure.threshold = 1 kvm.ha.recover.timeout = 60 kvm.ha.recover.wait.period = 600 ha.vm.restart.hostup = true vm.ha.migration.max.retries = 5Test performed
We performed repeated ungraceful host-failure tests by hard-powering off a KVM compute host through OOBM/IPMI.
Example failed host:Host name: ltchl1pacscom04 CloudStack Host ID: 19 Host IP: <redacted>At the time of one test, there were 9 user VMs running on this host.
Expected behavior:Host Down -> fencing succeeds -> all VMs restart on another KVM hostActual behavior:
Host Down -> KVM fencing succeeds -> most VMs restart successfully on other hosts -> a few VMs remain permanently in Stopping state -> HA work later becomes Done -> manual recovery is requiredFor example, during one 9-VM test, 6 VMs restarted successfully and the following VMs remained in
Stopping:i-22-248-VM i-22-257-VM i-22-258-VMManagement-server log observations
CloudStack correctly identifies the failed host as Down and KVM fencing succeeds:Fencer KVMFenceBuilder returned trueFor an affected VM, the HA stop/network cleanup workflow encountered a database transaction failure:
Unable to release some network resources for the VM in Stopping state com.cloud.utils.exception.CloudRuntimeException: Unable to commit or close the connection. Caused by: com.mysql.cj.jdbc.exceptions.MySQLTransactionRollbackException: Deadlock found when trying to get lock; try restarting transactionFor other affected VMs, HA logs:
Encountered unhandled exception during HA process, reschedule work com.cloud.utils.exception.CloudRuntimeException: Unable to find by id on DB, due to: Deadlock found when trying to get lock; try restarting transactionThe HA work is initially rescheduled. During the next HA attempt, CloudStack detects that the VM state has changed from
RunningtoStopping:HA on VM instance {... state="Stopping" ...} VM ... has been changed. Current State = Stopping Previous State = RunningThen the HA work is completed:
Completed work HAWork[...]However, the VM remains in
Stoppingand is not restarted on another host.
Database observation
For the affected VMs, the database shows:vm_instance.state = Stopping vm_instance.host_id = failed host ID Latest op_ha_work: step = DoneThe MariaDB Galera cluster was healthy after the test:
wsrep_cluster_status = Primary wsrep_cluster_size = 3 wsrep_connected = ON wsrep_ready = ON wsrep_local_state_comment = SyncedCurrent counters on one database node were:
Innodb_deadlocks = 0 wsrep_local_bf_aborts = 68 wsrep_local_cert_failures = 75 wsrep_local_replays = 4These Galera counters are cumulative and were not captured before and immediately after every HA test, so we cannot confirm that they are directly caused by this specific HA event. However, the CloudStack management-server logs clearly show
MySQLTransactionRollbackException: Deadlock found when trying to get lockduring the affected HA recovery workflow.
Ceph RBD lock observation
For VMs that restarted successfully, the RBD exclusive locks moved from the failed host to the destination compute hosts.
For VMs stuck inStopping, their root-volume RBD locks remained on the failed source host. Example:rbd lock ls cloudstack-primary/<rbd-image> There is 1 exclusive lock on this image. Locker ID Address client.<id> auto <id> <failed-compute-host-ip>:0/<client-session>Since the source host was hard powered off, this stale RBD lock is expected. However, CloudStack does not complete the HA recovery because the VM remains in
Stoppingand the HA work becomesDone.
Comparison with issue #13510
I understand that the original issue uses shared NFSv4.2 primary storage and the first restart failure is caused by a QEMU/NFS write-lock error.
Our initial HA failure appears different:Issue #13510: QEMU cannot acquire NFS write lock -> initial restart fails -> HA retry detects update counter mismatch -> HA work completes without retry Our environment: CloudStack HA stop/network cleanup encounters DB transaction deadlock -> VM remains Stopping -> HA retry detects changed VM state/update state -> HA work completes without successful restartThe storage backend is different, but the final behavior appears similar: after an initial HA recovery failure, CloudStack stops retrying and completes the HA work even though the VM remains unavailable.
Reacted by kiranchavalacc @andrijapanicsb @sureshanaparti your thoughts
@kiranchavala Now, I solved that issue. I added the NFS primary storage only for KVMHA Heartbeat. After that still Host HA is not working, Then we disable the Host HA and Only worked now for VM HA.
But, some of them VMs are stucked at the stopping state, we need to manual recover is needed, stop and start the VMs GUI then they all are running.I tried and many thing for this stucked VMs how to recover but we cant able to do that? Because, Many private cloud have an same issue during HA some VMs are stucked
@kiranchavala Any update on this?
@muthukrishnang1100 I am not able to reproduce the issue in my lab on the recent build of Cloudstack 4.22.1 with nfs storage
If you can reproduce the issue consistently on 4.22.1. Please let us know
http://packages.shapeblue.com/cloudstack/upstream/
http://packages.shapeblue.com/cloudstack/upstream/debian/4.22/Also I am not sure if galera mysql is offically supported with cloudsack
https://docs.cloudstack.apache.org/en/4.22.1.0/releasenotes/compat.html#software-requirements
@kiranchavala But, I am using Ceph and NFS for HA heartbeat only.
- locked and limited conversation to collaborators
on Jul 28, 2026
problem
We're trying VM HA on cloudstack. We disabled HOST HA (see issue 13371) and then simulated a crash on an host with a "fake" kernel panic. The investigators correctly found the host dead and put it in alert/disconnected state.
The management servers then tried to start the vms on another host
The job fails (expectedly i would say) because of the nfs write lock held from the dead host.
But when retrying the vm restart we get
The restart process didn't even try a second time because of the mismatch in the update counter between the vm and the work (20 vs 17). The final state of the vm (without manual intervention) is Stopped.
versions
Cloudstack version: 4.22.1.0, with management server running on ubuntu22 and agent running on oracle linux 9
Shared nfsv4 (4.2) storage
The steps to reproduce the bug
What to do about it?
No response