Skip to content

selenosis does not delete Failed Browser CRDs; pods stuck NotReady (1/2) indefinitely #8

Description

@mideeff

Summary

Browser pods can remain stuck in 1/2 NotReady indefinitely after seleniferous exits, while the corresponding Browser CRD stays in Failed. This persists well beyond SESSION_IDLE_TIMEOUT. Cleanup requires manual intervention.
Observed when running many parallel browser sessions and under cluster resource pressure (CPU/memory/scheduling; many Pending or slow-starting pods). The product issue is lack of guaranteed cleanup when Browser is terminal Failed and selenosis does not delete the CRD.

Environment

Component Version
selenosis alcounit/selenosis:v2.0.7
browser-controller alcounit/browser-controller:v0.0.6
browser-service alcounit/browser-service:v0.0.6
browser-ui alcounit/browser-ui:v0.0.7
seleniferous (sidecar) alcounit/seleniferous:v2.0.6

Steps to reproduce

  1. Run automated tests that create many browser sessions in parallel (enough to stress CPU/memory or scheduling in the namespace).
  2. Optionally observe scheduling pressure: kubectl get pods -n <namespace> shows multiple Pending or slow-starting browser pods.
  3. Wait for tests to finish.
  4. Check pods: kubectl get pods -n <namespace>
  5. Check CRDs: kubectl get browsers -n <namespace>

Expected behavior

After SESSION_IDLE_TIMEOUT (or after a definitive failure path), selenosis should remove the Browser CRD. browser-controller should complete finalizer handling and the pod should terminate.

Actual behavior

Pods remain 1/2 NotReady. Browser CRDs remain Failed and are not removed by selenosis (stayed more than 3 days in our cluster).
Example:

kubectl get browsers -n selenosis
NAME                                   BROWSER   VERSION     PHASE    AGE
eee39701-45c9-4427-a7bd-ace1f6502377   chrome    145.0-csp   Failed   33m
f31c9a1a-0933-4607-9bff-448326efb3ca   chrome    145.0-csp   Failed   33m
kubectl get pods -n selenosis
eee39701-45c9-4427-a7bd-ace1f6502377   1/2   NotReady   0   33m
f31c9a1a-0933-4607-9bff-448326efb3ca   1/2   NotReady   0   33m

Logs

browser-controller

browser-controller — marks Failed, sees seleniferous not ready, then stops further reconciliation (waits for selenosis to delete the CRD):

{"level":"info","Browser":{"name":"eee39701-45c9-4427-a7bd-ace1f6502377","namespace":"selenosis"},"message":"Browser status set to Failed"}
{"level":"info","Browser":{"name":"eee39701-45c9-4427-a7bd-ace1f6502377","namespace":"selenosis"},"containerName":"seleniferous","containerReady":"false","restartCount":0,"message":"Browser Pod container statuses"}
{"level":"info","Browser":{"name":"eee39701-45c9-4427-a7bd-ace1f6502377","namespace":"selenosis"},"finalizers":["browserpod.selenosis.io/finalizer"],"message":"current finalizers on Browser"}
{"level":"info","Browser":{"name":"eee39701-45c9-4427-a7bd-ace1f6502377","namespace":"selenosis"},"message":"Browser is in Failed state, nothing to do"}

Suspected root cause

  1. Under resource pressure, parallel runs produce partial failures: e.g. seleniferous exits or never becomes ready while the browser container may still run → pod 1/2, Browser Failed.
  2. selenosis may not complete normal session lifecycle / idle-timeout cleanup for these objects (e.g. session never fully registered or state lost under load).
  3. browser-controller sets Failed and then "nothing to do", expecting selenosis to delete the Browser CRD.
  4. Deadlock: selenosis does not delete the CRD → finalizer / pod cleanup does not finish → pods linger NotReady.

Workaround

kubectl delete browsers -n selenosis --all

Activity

  1. alcounit commented on Apr 8, 2026

    @alcounit
    Owner

    @mideeff thanks for detailed report. After fix Browser CR and Pod will be deleted according to following cases:

    Scenario Browser CR Pod
    No matching BrowserConfig handleMissingPod → Status=Failed → next reconcile: deleteBrowser → deleted never created
    podCreationTimeout exceeded (pod stuck in Pending > 5 min) Status=Failed → next reconcile: deleteBrowser → deleted deletePod (force, grace=0) → deleted
    PodPending + container Terminated Status=Failed → next reconcile: deleteBrowser → deleted deletePod (force, grace=0) → deleted
    PodPending + container Waiting with non-transient reason (CrashLoopBackOff, ErrImagePull, etc.) Status=Failed → next reconcile: deleteBrowser → deleted deletePod (force, grace=0) → deleted
    Pod phase Failed Status=Failed → next reconcile: deleteBrowser → deleted deletePod (force, grace=0) → deleted
    Status=Failed (Failed early exit — any of the above on next reconcile) deleteBrowser → finalizer removed → Delete(CR) → deleted deletePod (force, grace=0) → deleted
    Critical container (browser or seleniferous) Terminated while Running deleteBrowser → finalizer removed → Delete(CR) → deleted deleted via OwnerReference GC after CR deleted
    Browser CR DeletionTimestamp set (external delete) finalizer removed → deleted deletePod → explicit delete in handleDeletion → deleted
    Pod DeletionTimestamp set, CR alive deleteBrowser → deleted already terminating → deleted
  2. transferred this issue fromalcounit/selenosison Apr 8, 2026
  3. self-assigned this
    on Apr 9, 2026
  4. added a commit that references this issue on Apr 9, 2026
  5. alcounit commented on Apr 9, 2026

    @alcounit
    Owner

    @mideeff please deploy browser-controller v0.0.7, let me know of any issues

  6. mideeff commented on Apr 10, 2026

    @mideeff
    Author

    Update after local testing

    I deployed browser-controller:v0.0.7 and tested locally.

    Stuck Pending pods and Browser CRs remaining after failed or timed-out sessions no longer seem to reproduce. After ~5 minutes I see pod creation timeout exceeded, Browser pod forcibly deleted, and the Browser CRs are cleaned up — consistent with PR #9.

    Please do not close the issue yet. I will re-test on our corporate cluster next week (or shortly after) and post an update.

    Questions

    1. Is podCreationTimeout intentionally fixed at 5 minutes, or is there a plan to make it configurable?

    2. Selenosis v1 had --browser-limit for a global concurrent session cap. I do not see an equivalent in Selenosis v2. Is there a supported way to enforce the same limit in v2?

    Thanks.

  7. alcounit commented on Apr 10, 2026

    @alcounit
    Owner

    Is podCreationTimeout intentionally fixed at 5 minutes, or is there a plan to make it configurable?

    it is hardcoded for now, will move it to flags in next release

    Selenosis v1 had --browser-limit for a global concurrent session cap. I do not see an equivalent in Selenosis v2. Is there a supported way to enforce the same limit in v2?

    V2 doesn't support limit, limit per namespace can be done using ResourceQuota

      apiVersion: v1                                                                                                                                                                                
      kind: ResourceQuota                                                                                                                                                                           
      metadata:                                                                                                                                                                                   
        name: browser-quota
        namespace: selenosis
      spec:
        hard:
          count/browsers.selenosis.io: "10"
  8. mideeff commented on Apr 16, 2026

    @mideeff
    Author

    Hi,

    Update: we consider #8 resolved in our environment (thank you for the fix / guidance there).
    The text below is an additional production scenario that still hurts operators in a similar way (long tail of Browser / Pod work after CI), so we wanted to document it in the same thread for context.

    We are seeing behaviour that fits the same overall theme as #8 (long‑lived Browser / browser work and cleanup after a run), but the sequence is specific and we can reproduce it reliably under one condition.

    When it reproduces
    It happens only if the Pod quota is exhausted while the test run is still going.

    While tests are running, new sessions cannot get a Pod immediately — they back up (clients keep waiting; Browser objects pile up behind the quota). So the “queue” grows during the run.

    What happens after the test run finishes
    The CI job stops asking for new sessions, but the backlog from the quota‑blocked period is still there.

    Then we see a drain phase:

    The controller creates a new browser Pod for the next queued Browser (there is finally room under the quota for a Pod object).
    Over VNC the Pod often looks wrong: no real browser (empty / stuck UI).
    After about ~5 minutes that attempt ends (the Pod / attempt goes away).
    The next queued Browser gets another Pod — repeat until the backlog is empty.
    So the “bad tail” is not random noise — it is the backlog created during the run, processed one Pod at a time after CI has already exited.

    Comparison with Selenosis v1 (old hub)
    On Selenosis v1 we did not see this pattern. When the Pod quota was full, the session request failed immediately (for example SessionNotCreatedException / pod limit). Nothing silently queued, and no Pod was created for those failed requests.

    With Selenosis v2, the same “quota full during the run” situation becomes a growing backlog while tests still execute, and after the run we watch that backlog drain as empty / broken Pods with the ~5 minute cycle above.

    Suggestions (bounded “queue” / backlog lifetime)
    We would like a configurable bound so this backlog cannot grow and drain without a practical limit from the operator’s point of view. Examples that would help us:

    Fail fast when quota blocks Pod creation (similar to v1): do not keep long‑lived Browser objects for requests that cannot succeed until quota changes — fail the session quickly and delete / expire the Browser.

    Max age / TTL for a Browser that never reached a healthy running browser (from creation or from “first Pod create attempt”), after which the CR is removed or marked terminal and cleaned up — so a CI spike cannot leave hours of post‑run work.

    Max backlog size (per namespace or per hub): when the limit is hit, reject new session requests instead of accepting work that will only be processed much later.

    Thanks.

  9. alcounit commented on Apr 19, 2026

    @alcounit
    Owner

    Thanks for the detailed write-up.

    After some digging I came up with few changes and I need you help with testing. Please use alcounit/browser-controller:develop image

    Recommended configuration for your scenario

    --browser-pending-timeout=2m
    --browser-pod-creation-timeout=5m
    --max-workers=4
    --max-retries=3

    This ensures:

    • Browsers blocked by quota fail within 2 minutes (no backlog growth)
    • Pods that do get created but are stuck (your "empty/stuck UI") are cleaned up within 5 minutes
    • Multiple workers process cleanup concurrently

    Let us know if this resolves the post-run drain behavior in your environment.

  10. alcounit commented on Apr 27, 2026

    @alcounit
    Owner

    @mideeff did you get a chance to check with develop image?

  11. mideeff commented on Apr 27, 2026

    @mideeff
    Author

    @alcounit Hi. Unfortunately, we’re currently experiencing issues in our cluster, which prevent me from verifying this image. I think I’ll be able to check it this week or next.

  12. alcounit commented on May 15, 2026

    @alcounit
    Owner

    Hi @mideeff, any updates?

  13. mideeff commented on May 18, 2026

    @mideeff
    Author

    @alcounit Hi! So far, I haven’t had time for this. I think I’ll be able to test it only at the end of the week.

  14. alcounit commented on Jun 9, 2026

    @alcounit
    Owner

    Closing issue, codebase with fix merged in release v0.0.8

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't workingv0.0.7

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions