Repository navigation
selenosis does not delete Failed Browser CRDs; pods stuck NotReady (1/2) indefinitely #8
Description
Activity
@mideeff thanks for detailed report. After fix Browser CR and Pod will be deleted according to following cases:
Scenario Browser CR Pod No matching BrowserConfig handleMissingPod→Status=Failed→ next reconcile:deleteBrowser→ deletednever created podCreationTimeoutexceeded (pod stuck inPending> 5 min)Status=Failed→ next reconcile:deleteBrowser→ deleteddeletePod(force, grace=0) → deletedPodPending+ containerTerminatedStatus=Failed→ next reconcile:deleteBrowser→ deleteddeletePod(force, grace=0) → deletedPodPending+ containerWaitingwith non-transient reason (CrashLoopBackOff,ErrImagePull, etc.)Status=Failed→ next reconcile:deleteBrowser→ deleteddeletePod(force, grace=0) → deletedPod phase FailedStatus=Failed→ next reconcile:deleteBrowser→ deleteddeletePod(force, grace=0) → deletedStatus=Failed(Failed early exit — any of the above on next reconcile)deleteBrowser→ finalizer removed →Delete(CR)→ deleteddeletePod(force, grace=0) → deletedCritical container ( browserorseleniferous)Terminatedwhile RunningdeleteBrowser→ finalizer removed →Delete(CR)→ deleteddeleted via OwnerReference GC after CR deleted Browser CR DeletionTimestampset (external delete)finalizer removed → deleted deletePod→ explicit delete inhandleDeletion→ deletedPod DeletionTimestampset, CR alivedeleteBrowser→ deletedalready terminating → deleted Reacted by mideeff- linked a pull request that will close this issueFix https://github.com/alcounit/browser-controller/issues/8 #9
on Apr 9, 2026 @mideeff please deploy browser-controller v0.0.7, let me know of any issues
Reacted by mideeffUpdate after local testing
I deployed
browser-controller:v0.0.7and tested locally.Stuck
Pendingpods andBrowserCRs remaining after failed or timed-out sessions no longer seem to reproduce. After ~5 minutes I seepod creation timeout exceeded,Browser pod forcibly deleted, and theBrowserCRs are cleaned up — consistent with PR #9.Please do not close the issue yet. I will re-test on our corporate cluster next week (or shortly after) and post an update.
Questions
-
Is
podCreationTimeoutintentionally fixed at 5 minutes, or is there a plan to make it configurable? -
Selenosis v1 had
--browser-limitfor a global concurrent session cap. I do not see an equivalent in Selenosis v2. Is there a supported way to enforce the same limit in v2?
Thanks.
-
Is podCreationTimeout intentionally fixed at 5 minutes, or is there a plan to make it configurable?
it is hardcoded for now, will move it to flags in next release
Selenosis v1 had --browser-limit for a global concurrent session cap. I do not see an equivalent in Selenosis v2. Is there a supported way to enforce the same limit in v2?
V2 doesn't support limit, limit per namespace can be done using ResourceQuota
apiVersion: v1 kind: ResourceQuota metadata: name: browser-quota namespace: selenosis spec: hard: count/browsers.selenosis.io: "10"
Hi,
Update: we consider #8 resolved in our environment (thank you for the fix / guidance there).
The text below is an additional production scenario that still hurts operators in a similar way (long tail of Browser / Pod work after CI), so we wanted to document it in the same thread for context.We are seeing behaviour that fits the same overall theme as #8 (long‑lived Browser / browser work and cleanup after a run), but the sequence is specific and we can reproduce it reliably under one condition.
When it reproduces
It happens only if the Pod quota is exhausted while the test run is still going.While tests are running, new sessions cannot get a Pod immediately — they back up (clients keep waiting; Browser objects pile up behind the quota). So the “queue” grows during the run.
What happens after the test run finishes
The CI job stops asking for new sessions, but the backlog from the quota‑blocked period is still there.Then we see a drain phase:
The controller creates a new browser Pod for the next queued Browser (there is finally room under the quota for a Pod object).
Over VNC the Pod often looks wrong: no real browser (empty / stuck UI).
After about ~5 minutes that attempt ends (the Pod / attempt goes away).
The next queued Browser gets another Pod — repeat until the backlog is empty.
So the “bad tail” is not random noise — it is the backlog created during the run, processed one Pod at a time after CI has already exited.Comparison with Selenosis v1 (old hub)
On Selenosis v1 we did not see this pattern. When the Pod quota was full, the session request failed immediately (for example SessionNotCreatedException / pod limit). Nothing silently queued, and no Pod was created for those failed requests.With Selenosis v2, the same “quota full during the run” situation becomes a growing backlog while tests still execute, and after the run we watch that backlog drain as empty / broken Pods with the ~5 minute cycle above.
Suggestions (bounded “queue” / backlog lifetime)
We would like a configurable bound so this backlog cannot grow and drain without a practical limit from the operator’s point of view. Examples that would help us:Fail fast when quota blocks Pod creation (similar to v1): do not keep long‑lived Browser objects for requests that cannot succeed until quota changes — fail the session quickly and delete / expire the Browser.
Max age / TTL for a Browser that never reached a healthy running browser (from creation or from “first Pod create attempt”), after which the CR is removed or marked terminal and cleaned up — so a CI spike cannot leave hours of post‑run work.
Max backlog size (per namespace or per hub): when the limit is hit, reject new session requests instead of accepting work that will only be processed much later.
Thanks.
Thanks for the detailed write-up.
After some digging I came up with few changes and I need you help with testing. Please use
alcounit/browser-controller:developimageRecommended configuration for your scenario
--browser-pending-timeout=2m
--browser-pod-creation-timeout=5m
--max-workers=4
--max-retries=3This ensures:
- Browsers blocked by quota fail within 2 minutes (no backlog growth)
- Pods that do get created but are stuck (your "empty/stuck UI") are cleaned up within 5 minutes
- Multiple workers process cleanup concurrently
Let us know if this resolves the post-run drain behavior in your environment.
Reacted by mideeff@mideeff did you get a chance to check with develop image?
@alcounit Hi. Unfortunately, we’re currently experiencing issues in our cluster, which prevent me from verifying this image. I think I’ll be able to check it this week or next.
Reacted by Danylo KuvshynovHi @mideeff, any updates?
@alcounit Hi! So far, I haven’t had time for this. I think I’ll be able to test it only at the end of the week.
Closing issue, codebase with fix merged in release v0.0.8
Summary
Browser pods can remain stuck in
1/2 NotReadyindefinitely afterseleniferousexits, while the correspondingBrowserCRD stays inFailed. This persists well beyondSESSION_IDLE_TIMEOUT. Cleanup requires manual intervention.Observed when running many parallel browser sessions and under cluster resource pressure (CPU/memory/scheduling; many
Pendingor slow-starting pods). The product issue is lack of guaranteed cleanup whenBrowseris terminalFailedand selenosis does not delete the CRD.Environment
alcounit/selenosis:v2.0.7alcounit/browser-controller:v0.0.6alcounit/browser-service:v0.0.6alcounit/browser-ui:v0.0.7alcounit/seleniferous:v2.0.6Steps to reproduce
kubectl get pods -n <namespace>shows multiplePendingor slow-starting browser pods.kubectl get pods -n <namespace>kubectl get browsers -n <namespace>Expected behavior
After
SESSION_IDLE_TIMEOUT(or after a definitive failure path), selenosis should remove theBrowserCRD. browser-controller should complete finalizer handling and the pod should terminate.Actual behavior
Pods remain
1/2 NotReady.BrowserCRDs remainFailedand are not removed by selenosis (stayed more than 3 days in our cluster).Example:
Logs
browser-controller
browser-controller — marks Failed, sees seleniferous not ready, then stops further reconciliation (waits for selenosis to delete the CRD):
{"level":"info","Browser":{"name":"eee39701-45c9-4427-a7bd-ace1f6502377","namespace":"selenosis"},"message":"Browser status set to Failed"} {"level":"info","Browser":{"name":"eee39701-45c9-4427-a7bd-ace1f6502377","namespace":"selenosis"},"containerName":"seleniferous","containerReady":"false","restartCount":0,"message":"Browser Pod container statuses"} {"level":"info","Browser":{"name":"eee39701-45c9-4427-a7bd-ace1f6502377","namespace":"selenosis"},"finalizers":["browserpod.selenosis.io/finalizer"],"message":"current finalizers on Browser"} {"level":"info","Browser":{"name":"eee39701-45c9-4427-a7bd-ace1f6502377","namespace":"selenosis"},"message":"Browser is in Failed state, nothing to do"}Suspected root cause
Workaround