Skip to content

k8s deployer: WaitForDeploymentAvailable errors on a 0-replica Deployment (keda scale-to-zero redeploy) #4068

Description

@aliok

Summary

WaitForDeploymentAvailable (pkg/k8s/wait.go) can fail a redeploy with
could not find current pod-template-hash for deployment <name> when the live
Deployment has 0 desired replicas. This is reachable today via the keda
(HTTP-scaler) deployer, whose ScaledObject scales idle functions down to zero.

Note: this is reasoned from the code, not yet reproduced (it's a timing race).
Filing so the analysis is tracked; repro TBD.

Details

WaitForDeploymentAvailable re-fetches the live Deployment on every poll tick:

deployment, err := clientset.AppsV1().Deployments(namespace).Get(ctx, deploymentName, metav1.GetOptions{})
...
return checkIfDeploymentIsAvailable(ctx, clientset, deployment)

checkIfDeploymentIsAvailable finds the active ReplicaSet by looking for one with
*rs.Spec.Replicas > 0. A Deployment sitting at 0 desired replicas has no such
ReplicaSet, so currentPodTemplateHash stays empty and the function returns an error
(pkg/k8s/wait.go):

if currentPodTemplateHash == "" {
    return false, fmt.Errorf("could not find current pod-template-hash for deployment %s", deployment.Name)
}

The func deployer itself always applies replicas >= 1 (pkg/k8s/deployer.go), so
it never creates a 0-replica Deployment. But because the wait inspects live state,
reachability is governed by what KEDA does during the wait, not by what func wrote:

  1. An idle keda-HTTP function has been scaled to 0 replicas by KEDA.
  2. User redeploys. func applies the Deployment with replicas: 1, then starts polling.
  3. If the KEDA controller reasserts 0 replicas on the live Deployment before the poll
    observes it as ready, the live fetch returns spec.replicas == 0, and the wait
    returns the pod-template-hash error — failing the redeploy.

Suggested fix

Short-circuit when there is nothing to wait for:

desiredReplicas := *deployment.Spec.Replicas
if desiredReplicas == 0 {
    return true, nil // no replicas desired; nothing to wait for
}

This also unblocks any future spec-level scale-to-zero on the raw/keda Deployment.

Notes

  • Reasoned from code; not reproduced. A repro needs a cluster with KEDA and an idle,
    scaled-to-zero HTTP function being redeployed.
  • Fix is small and self-contained but likely wants a repro or unit test first, hence
    filing as an issue rather than an immediate PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions