DEVOPS-2899: Remove Trivy from job alerts, fix job count and perfdata - #53
Merged
Merged
Conversation
The deployed rollup reported the sum of event counts as the number of failed jobs, so "4 recent jobs failed" was only correct because each trivy scan job happened to have a single event. A job with a count=5 event made the number wrong. Report the job count and event total separately, and emit perfdata on both the OK and CRIT branches so the metric stays graphable instead of disappearing exactly when a job fails. Adopt the single-service rollup design that was already running on the host, and document why: job names carry a template hash, so one service per job accumulates discovered services forever and cannot alert on a new failure until the next discovery pass. Also: - Cap the detail at 10 jobs so a churning controller cannot grow the service output without bound. - Drop the now-dead service_name() and the unused svc/perf locals. - Strip '|' from event messages; CheckMK uses it to split perfdata from plugin output. - Match the hardening in statefulset_checker: list-form kubectl instead of shell=True, OSError handling, JSON shape validation, tolerate non-dict items and a null event count.
trivy-operator recreates its scan jobs constantly, and a failed scan is a scanner problem -- usually a rate-limited ghcr.io/aquasecurity/trivy-db pull -- not a problem with the workload being scanned. These were the dominant source of noise in K8s_Failed_Jobs. Exclude them by namespace and by job-name prefix. Namespace alone is not enough: trivy-operator can be configured to run scan jobs in the scanned workload's namespace instead of its own, and events carry only involvedObject, so there are no labels to match on. Suppressed failures are still counted and surfaced in the detail line and as a 'suppressed' metric, so a scanner that has stopped working is visible instead of silently ignored.
Contributor
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The focused changes correctly address the counting, perfdata, output-size, and error-handling requirements.
Review effort: Balanced
Findings: None
What changed in this PR
Corrects and hardens the Kubernetes failed-job CheckMK rollup.
Changes:
- Separates failed-job and event counts in perfdata.
- Caps failure details at 10 jobs and sanitizes messages.
- Improves kubectl execution and JSON error handling.
| File | Description |
|---|---|
lakehouse/k8s_failed_jobs |
Implements corrected rollup metrics, bounded details, and input hardening. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Contributor
Author
|
Thank you @kkellerlbl ! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes DEVOPS-2899.
Why
A CRIT from the lakehouse cluster read:
Two problems: the trivy scan jobs behind it are alert noise, and the count is wrong.
Alert fatigue (DEVOPS-2899). trivy-operator recreates its scan jobs constantly, and a failed scan is a scanner problem, not a problem with the workload being scanned. Because the check reads Events (~1h TTL) and trivy keeps retrying, these alert, self-clear, and re-alert indefinitely. Confirmed against the live cluster: at the time of writing there are zero failed Jobs and zero Job warning events in
trivy-operator, yet the operator is still logging fresh scan failures — the alert had already self-cleared.Wrong count. The check summed
event.get("count", 1)across jobs and labelled the result "recent jobs failed". It read4only because all four scan jobs happened to have one event each; two jobs where one has acount=5event would have reported "6 recent jobs failed" while listing two.Separately, the copy running on
prod-berdl-cp-01had diverged from this repo — it had been hand-edited from one-service-per-job to a single rollup service. That change was correct and is adopted here, with a comment explaining why so it does not get reverted.What changed
Trivy exclusion — filtered by namespace (
trivy-operator) and job-name prefix (scan-vulnerabilityreport-,scan-configauditreport-,scan-exposedsecretreport-,scan-sbomreport-). Namespace alone is insufficient: trivy-operator supportsscanJobsInSameNamespace, and events carry onlyinvolvedObject, so there are no labels to match on. Suppressed failures are still counted and surfaced in the detail line and as asuppressedmetric — see the root cause below for why going fully silent would be a mistake.Correctness and hygiene
failed_jobsis now the number of distinct jobs, with the event total reported separately.failed_jobs=N|events=M|suppressed=S). The CRIT branch previously emitted a literal-, so the metric vanished precisely when something was wrong.... and N more; previously every failure was concatenated into one unbounded line.service_name()and the unusedsvc/perflocals left behind by the rollup change.|stripped from event messages — CheckMK uses it to split perfdata from plugin output, and the old| {msg}separator put one into the service detail.statefulset_checker(DEVOPS-2883: Add statefulset checker #50): list-formkubectlinstead ofshell=True,OSErrorhandling, JSON shape validation, non-dict items skipped,count: nulltolerated.The docstring now records that this check reads Events, which Kubernetes expires after ~1h, so a job that failed before that window is not reported even if its Job object is still
Failed.Testing
Run against a stubbed
microk8s:failed_jobs=0|events=0|suppressed=4OK — the alert is gonesuppressed=2suppressed=2still reported in the detailkbase-prod/scanner-import(near-miss name)2 job(s) failed, 6 failure event(s); old code said "6 recent jobs failed"... and 4 more{"items": "nope"}3 ... UNKNOWN: unexpected kubectl JSON shapecount: null,|and newline in messageLC_ALL=C PYTHONIOENCODING=asciiUnicodeEncodeError(output is ASCII-only)Root cause of the scan failures (needs its own ticket)
Investigated against the live cluster. It is not a registry rate limit — the Helm values already mirror everything through
mirror.gcr.io(dbRegistry,javaDbRegistry,trivy.image.registry), and trivy runs inClientServermode against the in-clustertrivy-server. Scanning broadly works: 300+ VulnerabilityReports exist, many only hours old.The failures are a subset, hitting large images, with two distinct errors in the operator log:
Scratch space disappearing mid-scan (the dominant one, still firing):
Scan jobs run with
readOnlyRootFilesystem: trueandimageScanCacheDir: /tmp/trivy/.cache, whilescanJobCustomVolumes/scanJobCustomVolumesMountare both empty. The scan job's/tmpemptyDir is node-local and unsized, so this is consistent with ephemeral-storage exhaustion on the node running the scans. Worth checking disk pressure there, and giving scan jobs an explicitly sized/tmpvolume.HTTP/2 stream aborts on very large files, on the Spark and conda images:
Note
trivyOperator.excludeImagesis already set toghcr.io/berdatalakehouse/spark_notebook*, but failures continue onnotebook,spark-master,buildandpolaris-bootstrapcontainers, so that exclusion is not matching everything it was meant to.Also
operator.scanJobTTLis unset (""), so failed scan jobs are never reaped on a timer — that directly feeds the event churn this PR is silencing.This is why the check reports
suppressed=Nrather than dropping these silently. Prod Spark/Jupyter images are among the ones persistently failing to scan; a blanket silence would make "trivy stopped scanning your biggest prod images" look identical to a healthy cluster.