Repository navigation
Autoscale: infinite scale-up when VMs fail to START (Stopped) — #11244 guard only counts Error state #14185
Description
Activity
🎯 Triage report
Well-documented report of an infinite AutoScale scale-up loop: when a scale-up VM fails to start (landing in
State.Stoppedrather thanState.Error), neither the errored-instance guard from #11244 nor the active-member counter account forStoppedVMs, so the group scales up forever. The reporter also identifies a related defect wheredoScaleUp's cleanup only catchesServerApiException, missing theCloudRuntimeExceptionactually thrown byVirtualMachineManagerImpl.start(), so failed VMs are never destroyed. Real-world impact was VM/IP exhaustion (2,296 VMs from a max_members=2 group).📊 Assessment
Dimension Value Reasoning Type type:bug Clear functional defect with reproducible logic flaw and code-level evidence Component component:management-server AutoScale scaling logic lives in AutoScaleManagerImpl/AutoScaleVmGroupVmMapDaoImpl(management server core); no dedicatedcomponent:autoscalelabel existsSeverity Severity:Major Caused resource exhaustion (guest IP/subnet, thousands of leaked VMs) impacting unrelated deployments in the reporter's environment, though it requires a specific start-failure trigger to manifest Labels type:bug, Severity:Major, component:management-server See above Coding agent Suitable Root cause is precisely identified with exact file/method names and code snippets, and three concrete, scoped fix options are proposed (broaden guard to include Stopped, catchCloudRuntimeExceptionindoScaleUp, or countStoppedincountAvailableVmsByGroup)🔗 Similar issues
- PR Prevent infinite autoscaling #11244 (referenced in issue) — merged fix for the original infinite-scale-up bug (Autoscale Group creates Infinite Number Of VMs when unable to startup #9318), but only guards against
State.Error, notState.Stopped; this issue identifies the gap left by that fix. - PR Prevent infinite retries of autoscaling #9574 (referenced in issue) — an earlier, unmerged one-line proposal that reportedly would have addressed a related variant.
No separate open GitHub issue duplicating this exact report was found via search.
💡 Notes and suggestions
- Recommend maintainers/assignee evaluate the reporter's three suggested fixes together rather than in isolation: option 1 (include
StoppedingetErroredInstanceCount()) stops the infinite loop quickly, but combining it with option 2 (broadening thedoScaleUpcatch toCloudRuntimeException) prevents the VM leak from accumulating in the first place, which is the root cause of the exhaustion. - Worth double-checking whether other terminal/failure VM states (e.g.
Destroyed,Expunging) should also be considered incountAvailableVmsByGroup()for consistency. - A regression/unit test simulating a start failure that raises
CloudRuntimeException(notServerApiException) fromadvanceStartwould help confirm the fix and guard against future regressions.
Generated by Daily Issue Triage · sonnet50 102.2K · ◷
Add this agentic workflows to your repo
To install this agentic workflow, run
gh aw add githubnext/agentics/workflows/daily-issue-triage.md@d7c1dc4b72b00607a67caaffdcc216cb64379cf9- PR Prevent infinite autoscaling #11244 (referenced in issue) — merged fix for the original infinite-scale-up bug (Autoscale Group creates Infinite Number Of VMs when unable to startup #9318), but only guards against
- linked a pull request that will close this issueconsidder stopped VMs on autoscale #14284
on Oct 1, 2026
ISSUE TYPE
COMPONENT NAME
CLOUDSTACK VERSION
CONFIGURATION
Advanced zone, KVM, Ceph/RBD-only primary storage. One AutoScale VM group
(
min_members=1,max_members=2,interval=30) on an isolated network with a/24 guest CIDR.
OS / ENVIRONMENT
Linux (Debian 12), KVM hosts.
SUMMARY
The infinite-autoscaling guard added in #11244 (fixing #9318) only counts
instances in
State.Error. When scale-up VMs fail to start — as opposed tofailing to be created — they land in
State.Stopped, notState.Error. Theguard therefore never trips, and the group scales up on every interval
indefinitely.
In our incident this produced 2,296 VMs from a group whose
max_membersis 2,over roughly 21 hours, until the guest subnet was exhausted.
Two independent counters are involved and both exclude
Stopped:AutoScaleVmGroupVmMapDaoImpl.getErroredInstanceCount()— the Prevent infinite autoscaling #11244 guard —counts
State.Erroronly:AutoScaleVmGroupVmMapDaoImpl.countAvailableVmsByGroup()— used by everyscaling decision in
AutoScaleManagerImpl(checkConditionUp,checkConditionDown,checkAutoScaleVmGroup, and the group-state handlers) —counts only
Starting,Running,Stopping,Migrating:So with N leaked
Stoppedmembers,currentVM == 0:checkAutoScaleVmGroup:if (currentVM < minMembers)->0 < 1-> scale up, every intervalcheckAutoScaleVmGroup:if (currentVM > maxMembers)->0 > 2-> scale-down never firescheckConditionDown:if (currentVM - 1 < minVm)->-1 < 1-> scale-down additionally blockedcheckConditionUp: errored-instance guard ->0 > 10false -> guard never tripsA third defect prevents the failed VM from being cleaned up, which is what allows
the leak to accumulate in the first place.
doScaleUppersists the group map rowbefore attempting the start, and its cleanup is guarded on
ServerApiException:startNewVMdoes convertInsufficientCapacityExceptionintoServerApiException,but it never sees that exception, because
VirtualMachineManagerImpl.start()has already wrapped it into an unchecked
CloudRuntimeException:CloudRuntimeExceptionmatches none ofstartNewVM's typed catches and is not aServerApiException, so it propagates pastdoScaleUp's handler toAutoScaleManagerImpl$MonitorTask, anddestroyVm()is never called. Theobserved log line is exactly this:
Note that PR #9574 ("Prevent infinite retries of autoscaling"), which proposed a
one-line change to
AutoScaleVmGroupVmMapDaoImpl, was closed unmerged; the merged#11244 took the threshold approach instead, which is what leaves this variant
uncovered.
STEPS TO REPRODUCE
Create an AutoScale VM group (
min_members=1,max_members=2, short interval).Let it stabilise at 1 running VM.
Break VM start (not creation) in a way that returns an
InsufficientCapacityExceptionorResourceUnavailableExceptionfromadvanceStart.The trigger we actually hit was the group network's Virtual Router becoming
unreachable, so
VirtualRouterElement.applyDhcpEntriesfailed with:ResourceUnavailableException: Resource [DataCenter:1] is unreachable: Unable to apply dhcp entry on router.That was an observed failure rather than a deliberate test, so I have not
confirmed that stopping the VR is a minimal reproducer — any start-path failure
that surfaces as
CloudRuntimeExceptionout ofVirtualMachineManagerImpl.start()should exhibit the same leak.
Observe: one new VM per interval, each landing in
Stopped, each retaining itsautoscale_vmgroup_vm_maprow, indefinitely.EXPECTED RESULTS
Scale-up stops after a bounded number of consecutive failed starts, and/or failed
instances are cleaned up, and/or
Stoppedmembers count towardmax_members.ACTUAL RESULTS
Unbounded VM creation. In our case, one VM per 30s for ~21 hours:
Error(created after the guest subnet was exhausted — these failat IP allocation, before a NIC is assigned)
Stopped, each holding a NIC and therefore a guest IPincluding unrelated, non-autoscale ones — failed with
InsufficientVirtualNetworkCapacityException: Unable to acquire Guest IP addressautoscale.errored.instance.thresholdwas at its default of 10 throughout, andgetErroredInstanceCount()returned 0 the entire time, because none of the leakedinstances were in
Error— they were inStopped.SUGGESTED FIX
Any one of these would break the loop; the first two seem most direct:
State.StoppedingetErroredInstanceCount()(or add a separate"failed instance" counter covering both
ErrorandStopped).doScaleUp's catch fromServerApiExceptionto also handleCloudRuntimeException, sodestroyVm()runs and the member is not leaked.Stoppedmembers incountAvailableVmsByGroup()so that leaked memberscontribute to
max_membersand become eligible for scale-down.