Skip to content

Autoscale: infinite scale-up when VMs fail to START (Stopped) — #11244 guard only counts Error state #14185

Description

@VishnuVuggepalli
ISSUE TYPE
  • Bug Report
COMPONENT NAME
VM Autoscale Feature
CLOUDSTACK VERSION
4.22.0.0 (observed)
Also present in 4.22.1.0, 4.23.0.0 and main (verified by source inspection)
CONFIGURATION

Advanced zone, KVM, Ceph/RBD-only primary storage. One AutoScale VM group
(min_members=1, max_members=2, interval=30) on an isolated network with a
/24 guest CIDR.

OS / ENVIRONMENT

Linux (Debian 12), KVM hosts.

SUMMARY

The infinite-autoscaling guard added in #11244 (fixing #9318) only counts
instances in State.Error. When scale-up VMs fail to start — as opposed to
failing to be created — they land in State.Stopped, not State.Error. The
guard therefore never trips, and the group scales up on every interval
indefinitely.

In our incident this produced 2,296 VMs from a group whose max_members is 2,
over roughly 21 hours, until the guest subnet was exhausted.

Two independent counters are involved and both exclude Stopped:

  1. AutoScaleVmGroupVmMapDaoImpl.getErroredInstanceCount() — the Prevent infinite autoscaling #11244 guard —
    counts State.Error only:
public int getErroredInstanceCount(long vmGroupId) {
    SearchCriteria<Integer> sc = CountBy.create();
    sc.setParameters("vmGroupId", vmGroupId);
    sc.setJoinParameters("vmSearch", "states", State.Error);   // Stopped not counted
    ...
}
  1. AutoScaleVmGroupVmMapDaoImpl.countAvailableVmsByGroup() — used by every
    scaling decision in AutoScaleManagerImpl (checkConditionUp,
    checkConditionDown, checkAutoScaleVmGroup, and the group-state handlers) —
    counts only Starting, Running, Stopping, Migrating:
sc.setJoinParameters("vmSearch", "states",
        State.Starting, State.Running, State.Stopping, State.Migrating);  // Stopped not counted

So with N leaked Stopped members, currentVM == 0:

  • checkAutoScaleVmGroup: if (currentVM < minMembers) -> 0 < 1 -> scale up, every interval
  • checkAutoScaleVmGroup: if (currentVM > maxMembers) -> 0 > 2 -> scale-down never fires
  • checkConditionDown: if (currentVM - 1 < minVm) -> -1 < 1 -> scale-down additionally blocked
  • checkConditionUp: errored-instance guard -> 0 > 10 false -> guard never trips

A third defect prevents the failed VM from being cleaned up, which is what allows
the leak to accumulate in the first place. doScaleUp persists the group map row
before attempting the start, and its cleanup is guarded on ServerApiException:

// AutoScaleManagerImpl.doScaleUp
autoScaleVmGroupVmMapDao.persist(groupVmMapVO);   // persisted BEFORE the start attempt
try {
    startNewVM(vm.getId());
    ...
} catch (ServerApiException e) {
    ...
    destroyVm(vm.getId());      // never reached, see below
    break;
}

startNewVM does convert InsufficientCapacityException into ServerApiException,
but it never sees that exception, because VirtualMachineManagerImpl.start()
has already wrapped it into an unchecked CloudRuntimeException:

try {
    advanceStart(vmUuid, params, planToDeploy, planner);
} catch (ConcurrentOperationException | InsufficientCapacityException e) {
    throw new CloudRuntimeException(String.format("Unable to start a VM [%s] due to [%s].", vmUuid, e.getMessage()), e);
}

CloudRuntimeException matches none of startNewVM's typed catches and is not a
ServerApiException, so it propagates past doScaleUp's handler to
AutoScaleManagerImpl$MonitorTask, and destroyVm() is never called. The
observed log line is exactly this:

WARN  [c.c.n.a.A.MonitorTask] Caught the following exception on monitoring AutoScale Vm Group
      com.cloud.utils.exception.CloudRuntimeException: Unable to start a VM [...]

Note that PR #9574 ("Prevent infinite retries of autoscaling"), which proposed a
one-line change to AutoScaleVmGroupVmMapDaoImpl, was closed unmerged; the merged
#11244 took the threshold approach instead, which is what leaves this variant
uncovered.

STEPS TO REPRODUCE
  1. Create an AutoScale VM group (min_members=1, max_members=2, short interval).

  2. Let it stabilise at 1 running VM.

  3. Break VM start (not creation) in a way that returns an
    InsufficientCapacityException or ResourceUnavailableException from
    advanceStart.

    The trigger we actually hit was the group network's Virtual Router becoming
    unreachable, so VirtualRouterElement.applyDhcpEntries failed with:
    ResourceUnavailableException: Resource [DataCenter:1] is unreachable: Unable to apply dhcp entry on router.
    That was an observed failure rather than a deliberate test, so I have not
    confirmed that stopping the VR is a minimal reproducer — any start-path failure
    that surfaces as CloudRuntimeException out of VirtualMachineManagerImpl.start()
    should exhibit the same leak.

  4. Observe: one new VM per interval, each landing in Stopped, each retaining its
    autoscale_vmgroup_vm_map row, indefinitely.

EXPECTED RESULTS

Scale-up stops after a bounded number of consecutive failed starts, and/or failed
instances are cleaned up, and/or Stopped members count toward max_members.

ACTUAL RESULTS

Unbounded VM creation. In our case, one VM per 30s for ~21 hours:

  • 2,043 VMs in Error (created after the guest subnet was exhausted — these fail
    at IP allocation, before a NIC is assigned)
  • 253 VMs in Stopped, each holding a NIC and therefore a guest IP
  • The /24 guest network reached 254/254 NICs and all subsequent VM deployments —
    including unrelated, non-autoscale ones — failed with
    InsufficientVirtualNetworkCapacityException: Unable to acquire Guest IP address

autoscale.errored.instance.threshold was at its default of 10 throughout, and
getErroredInstanceCount() returned 0 the entire time, because none of the leaked
instances were in Error — they were in Stopped.

SUGGESTED FIX

Any one of these would break the loop; the first two seem most direct:

  1. Include State.Stopped in getErroredInstanceCount() (or add a separate
    "failed instance" counter covering both Error and Stopped).
  2. Broaden doScaleUp's catch from ServerApiException to also handle
    CloudRuntimeException, so destroyVm() runs and the member is not leaked.
  3. Count Stopped members in countAvailableVmsByGroup() so that leaked members
    contribute to max_members and become eligible for scale-down.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions