Skip to content

DEVOPS-2899: Remove Trivy from job alerts, fix job count and perfdata - #53

Merged
bio-boris merged 2 commits into
masterfrom
fix-k8s-failed-jobs-count
Oct 2, 2026
Merged

bio-boris merged 2 commits into
masterfrom
fix-k8s-failed-jobs-count

Conversation

@bio-boris

@bio-boris bio-boris commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Fixes DEVOPS-2899.

Why

A CRIT from the lakehouse cluster read:

CRIT: 4 recent jobs failed, list: trivy-operator/scan-vulnerabilityreport-65644f45f7 - BackoffLimitExceeded (events: 1) | ...

Two problems: the trivy scan jobs behind it are alert noise, and the count is wrong.

Alert fatigue (DEVOPS-2899). trivy-operator recreates its scan jobs constantly, and a failed scan is a scanner problem, not a problem with the workload being scanned. Because the check reads Events (~1h TTL) and trivy keeps retrying, these alert, self-clear, and re-alert indefinitely. Confirmed against the live cluster: at the time of writing there are zero failed Jobs and zero Job warning events in trivy-operator, yet the operator is still logging fresh scan failures — the alert had already self-cleared.

Wrong count. The check summed event.get("count", 1) across jobs and labelled the result "recent jobs failed". It read 4 only because all four scan jobs happened to have one event each; two jobs where one has a count=5 event would have reported "6 recent jobs failed" while listing two.

Separately, the copy running on prod-berdl-cp-01 had diverged from this repo — it had been hand-edited from one-service-per-job to a single rollup service. That change was correct and is adopted here, with a comment explaining why so it does not get reverted.

What changed

Trivy exclusion — filtered by namespace (trivy-operator) and job-name prefix (scan-vulnerabilityreport-, scan-configauditreport-, scan-exposedsecretreport-, scan-sbomreport-). Namespace alone is insufficient: trivy-operator supports scanJobsInSameNamespace, and events carry only involvedObject, so there are no labels to match on. Suppressed failures are still counted and surfaced in the detail line and as a suppressed metric — see the root cause below for why going fully silent would be a mistake.

Correctness and hygiene

  • failed_jobs is now the number of distinct jobs, with the event total reported separately.
  • Perfdata on both branches (failed_jobs=N|events=M|suppressed=S). The CRIT branch previously emitted a literal -, so the metric vanished precisely when something was wrong.
  • Detail capped at 10 jobs with ... and N more; previously every failure was concatenated into one unbounded line.
  • Removed dead service_name() and the unused svc/perf locals left behind by the rollup change.
  • | stripped from event messages — CheckMK uses it to split perfdata from plugin output, and the old | {msg} separator put one into the service detail.
  • Hardening to match statefulset_checker (DEVOPS-2883: Add statefulset checker #50): list-form kubectl instead of shell=True, OSError handling, JSON shape validation, non-dict items skipped, count: null tolerated.

The docstring now records that this check reads Events, which Kubernetes expires after ~1h, so a job that failed before that window is not reported even if its Job object is still Failed.

Testing

Run against a stubbed microk8s:

Case Result
The 4 trivy jobs from the original alert failed_jobs=0|events=0|suppressed=4 OK — the alert is gone
Trivy scan job in a foreign namespace + one real failure CRIT on the real job only, suppressed=2
Only trivy failures OK, suppressed=2 still reported in the detail
kbase-prod/scanner-import (near-miss name) CRIT — not over-filtered
2 jobs / 6 events (the counting bug) 2 job(s) failed, 6 failure event(s); old code said "6 recent jobs failed"
14 failures capped at 10 + ... and 4 more
{"items": "nope"} 3 ... UNKNOWN: unexpected kubectl JSON shape
count: null, | and newline in message handled, no pipe in output
LC_ALL=C PYTHONIOENCODING=ascii no UnicodeEncodeError (output is ASCII-only)

Root cause of the scan failures (needs its own ticket)

Investigated against the live cluster. It is not a registry rate limit — the Helm values already mirror everything through mirror.gcr.io (dbRegistry, javaDbRegistry, trivy.image.registry), and trivy runs in ClientServer mode against the in-cluster trivy-server. Scanning broadly works: 300+ VulnerabilityReports exist, many only hours old.

The failures are a subset, hitting large images, with two distinct errors in the operator log:

  1. Scratch space disappearing mid-scan (the dominant one, still firing):

    failed to copy file to temp: create temp error:
    open /tmp/trivy-7/analyzer-composite-.../analyzer-file-...: no such file or directory
    

    Scan jobs run with readOnlyRootFilesystem: true and imageScanCacheDir: /tmp/trivy/.cache, while scanJobCustomVolumes/scanJobCustomVolumesMount are both empty. The scan job's /tmp emptyDir is node-local and unsized, so this is consistent with ephemeral-storage exhaustion on the node running the scans. Worth checking disk pressure there, and giving scan jobs an explicitly sized /tmp volume.

  2. HTTP/2 stream aborts on very large files, on the Spark and conda images:

    failed to analyze usr/local/spark-4.1.3-bin-hadoop3/jars/bundle-2.29.52.jar:
    failed to copy: stream error: stream ID 5; PROTOCOL_ERROR; received from peer
    

Note trivyOperator.excludeImages is already set to ghcr.io/berdatalakehouse/spark_notebook*, but failures continue on notebook, spark-master, build and polaris-bootstrap containers, so that exclusion is not matching everything it was meant to.

Also operator.scanJobTTL is unset (""), so failed scan jobs are never reaped on a timer — that directly feeds the event churn this PR is silencing.

This is why the check reports suppressed=N rather than dropping these silently. Prod Spark/Jupyter images are among the ones persistently failing to scan; a blanket silence would make "trivy stopped scanning your biggest prod images" look identical to a healthy cluster.

The deployed rollup reported the sum of event counts as the number of
failed jobs, so "4 recent jobs failed" was only correct because each
trivy scan job happened to have a single event. A job with a count=5
event made the number wrong.

Report the job count and event total separately, and emit perfdata on
both the OK and CRIT branches so the metric stays graphable instead of
disappearing exactly when a job fails.

Adopt the single-service rollup design that was already running on the
host, and document why: job names carry a template hash, so one service
per job accumulates discovered services forever and cannot alert on a
new failure until the next discovery pass.

Also:
- Cap the detail at 10 jobs so a churning controller cannot grow the
  service output without bound.
- Drop the now-dead service_name() and the unused svc/perf locals.
- Strip '|' from event messages; CheckMK uses it to split perfdata from
  plugin output.
- Match the hardening in statefulset_checker: list-form kubectl instead
  of shell=True, OSError handling, JSON shape validation, tolerate
  non-dict items and a null event count.
Copilot AI balanced review requested due to automatic review settings October 2, 2026 03:12
trivy-operator recreates its scan jobs constantly, and a failed scan is a
scanner problem -- usually a rate-limited ghcr.io/aquasecurity/trivy-db
pull -- not a problem with the workload being scanned. These were the
dominant source of noise in K8s_Failed_Jobs.

Exclude them by namespace and by job-name prefix. Namespace alone is not
enough: trivy-operator can be configured to run scan jobs in the scanned
workload's namespace instead of its own, and events carry only
involvedObject, so there are no labels to match on.

Suppressed failures are still counted and surfaced in the detail line and
as a 'suppressed' metric, so a scanner that has stopped working is
visible instead of silently ignored.
@bio-boris bio-boris changed the title Fix job count and perfdata in k8s_failed_jobs DEVOPS-2899: Remove Trivy from job alerts, fix job count and perfdata Oct 2, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The focused changes correctly address the counting, perfdata, output-size, and error-handling requirements.

Review effort: Balanced
Findings: None

What changed in this PR

Corrects and hardens the Kubernetes failed-job CheckMK rollup.

Changes:

  • Separates failed-job and event counts in perfdata.
  • Caps failure details at 10 jobs and sanitizes messages.
  • Improves kubectl execution and JSON error handling.
File Description
lakehouse/​k8s_failed_jobs Implements corrected rollup metrics, bounded details, and input hardening.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@kkellerlbl kkellerlbl left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@bio-boris

Copy link
Copy Markdown
Contributor Author

Thank you @kkellerlbl !

@bio-boris
bio-boris merged commit 48bf181 into master Oct 2, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants