Skip to content

NetworkGarbageCollector releases a network's VLAN while it still has live NICs #14177

Description

@Pearl1594

problem

When a VM's domain is detected missing (PowerReportMissing) past the graceful period (vm.op.wait.interval), CloudStack releases its NIC (broadcast_uri/isolation_uri -> NULL, nics_count) and marks the VM Stopped. If the VM's power state later resyncs back to Running (domain restored, HA restart, etc.), the NIC is never re-reserved and nics_count is never restored.
Separately, op_networks.nics_count never counts the network's own VirtualRouter NIC, so it under-counts from network creation. Combined, a single NIC release event can drive nics_count to 0 while the network still has live NICs (including on a Running VM).
NetworkGarbageCollector trusts nics_count==0 (plus a check of CloudStack's own DB-tracked "no non-Stopped instances") without verifying against the actual nics table or hypervisor state, and proceeds to stop the VR and release the VLAN back to the dynamic allocation pool - while it may still be bridged to a live VM. In our environment, GC's cleanup step also removes the host-level bridge, making the VM unrecoverable via a normal restart.

versions

tested on 4.22 (Maybe be observed on older versions too)

The steps to reproduce the bug

  1. Deploy a VM on an isolated network - note the nics_count = 1 (instead of 2 for VM and VR)
  2. Disable HA on the VM (as HA masks the issue)
  3. Perform virsh destroy <vm-name> to simulate domain loss - backup the dumpxml virsh dumpxml <vm-name> > backup.xml
  4. After the graceful period, CloudStack releases the NIC and decrements nics_count to 0.
  5. Restore the domain via virsh create <dumped-xml> . VM resyncs to Running, but NIC stays unreserved, counter stays 0
  6. Lower network.gc.interval/network.gc.wait to accelerate the scavenger; restart management server.
  7. virsh destroy again to bring the VM's tracked state back to Stopped
  8. GC fires: stops the VR, releases the VLAN, network transitions to Allocated broadcast_uri=NULL, despite 2 live NICs still present in nics.
    ...

What to do about it?

  • On VM power-state resync to Running after a missing-VM event, reconcile the NIC (broadcast_uri/isolation_uri/state=Reserved) and restore nics_count.
  • Harden NetworkGarbageCollector to verify live NIC count from the nics table (WHERE network_id=? AND removed IS NULL) before tearing down a network, instead of trusting the cached counter alone
  • Include the VirtualRouter's own NIC in nics_count from network creation

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions