Skip to content

Commit fbb8eaf

Browse files
updates
1 parent dff2144 commit fbb8eaf

35 files changed

Lines changed: 2818 additions & 0 deletions

File tree

Lines changed: 15 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,15 @@
1+
import fs from 'fs';
2+
import path from 'path';
3+
import BlogArticleShell from '@/components/BlogArticleShell';
4+
5+
export default function Article() {
6+
const html = fs.readFileSync(
7+
path.join(process.cwd(), 'content/blog-articles', "cicd-pipeline-failure-troubleshooting", 'body.html'),
8+
'utf8'
9+
);
10+
return (
11+
<BlogArticleShell>
12+
<div dangerouslySetInnerHTML={{ __html: html }} />
13+
</BlogArticleShell>
14+
);
15+
}
Lines changed: 81 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,81 @@
1+
<p>A red pipeline is a symptom, not a diagnosis. The fastest engineers do not start reading logs top-to-bottom — they first place the failure in one of three layers: <strong>the application</strong> (your code / tests are genuinely broken), <strong>the pipeline infrastructure</strong> (registry, cache, credentials, dependency sources, deploy target), or <strong>the runner/agent</strong> (the machine executing the job — disk, memory, tool versions). Fixing the wrong layer is how teams end up re-running jobs for an hour. This guide is the triage.</p>
2+
3+
<hr>
4+
5+
<h3>The one question that localises 80% of failures</h3>
6+
<p><strong>"Did this exact commit pass before?"</strong></p>
7+
<ul>
8+
<li><strong>A previously-green commit now fails, no code change</strong> → the environment changed. Infrastructure or runner. Suspect a moved dependency, an expired secret, a new base image, a full disk, a registry outage.</li>
9+
<li><strong>Failure started exactly on the commit that touched the relevant code, and reproduces locally</strong> → application. Fix the code/test.</li>
10+
<li><strong>Fails intermittently on the same commit</strong> → non-determinism: flaky test, race in parallel jobs, or a resource-starved runner. Do not "fix" it with a re-run.</li>
11+
</ul>
12+
13+
<h3>Layer 1 — Application failures</h3>
14+
<p>These are legitimate and the pipeline is doing its job. Signals:</p>
15+
<ul>
16+
<li>Compile/type errors, unit or integration test assertions, lint/format gates, coverage thresholds.</li>
17+
<li>Reproduces locally with the same command the CI step runs.</li>
18+
<li>Localised to the build/test stage, tied to a specific diff.</li>
19+
</ul>
20+
<p><strong>What to check first:</strong> run the exact CI command locally (not your IDE's runner) with the same tool versions. Most "works on my machine" gaps are version drift — pin them (Step: Layer 2).</p>
21+
22+
<h3>Layer 2 — Pipeline infrastructure failures</h3>
23+
<p>The code is fine; the plumbing broke. Signals and causes:</p>
24+
<table>
25+
<thead><tr><th>Failing stage</th><th>Likely infrastructure cause</th></tr></thead>
26+
<tbody>
27+
<tr><td>Checkout / clone</td><td>SCM auth token expired, LFS quota, network to the SCM</td></tr>
28+
<tr><td>Dependency fetch (npm/pip/maven)</td><td>Registry down, a version yanked, a transitive dep moved, lockfile not honoured</td></tr>
29+
<tr><td>Docker build</td><td>Base image tag changed under you, cache poisoned, build-arg/secret missing</td></tr>
30+
<tr><td>Image push</td><td>Registry credentials expired, repository quota, immutable-tag conflict</td></tr>
31+
<tr><td>Deploy</td><td>Cluster/target credentials rotated, changed manifest, environment gate, quota</td></tr>
32+
</tbody>
33+
</table>
34+
<p>The tell for infrastructure is that it fails <em>outside your source tree</em> — before your code even runs, or when talking to an external system. The fix is usually pinning and hardening: pin base images by digest, honour lockfiles, rotate secrets with overlap, and make external fetches retry with backoff.</p>
35+
36+
<h3>Layer 3 — Runner / agent failures</h3>
37+
<p>The most under-diagnosed layer, because the logs often look like an application error. The machine itself is the problem:</p>
38+
<ul>
39+
<li><strong>Out of disk</strong> — accumulated Docker layers and build caches fill the agent; builds fail with "no space left on device" or cryptic tar/extract errors. Prune caches; size ephemeral runners.</li>
40+
<li><strong>Out of memory</strong> — parallel jobs or a heavy compile/test OOM the agent; the job is killed (137) and the log just stops mid-step. Reduce parallelism or use a bigger runner class.</li>
41+
<li><strong>Stale / poisoned cache</strong> — a corrupted dependency or Docker layer cache makes a good commit fail; a cache key that is too broad shares bad state across branches. Bust the cache key.</li>
42+
<li><strong>Version drift</strong> — the runner image updated its default Node/Java/Python/Docker and your build assumed the old one. Pin the toolchain in the pipeline, not the runner image.</li>
43+
<li><strong>Clock skew / TLS</strong> — a badly-timed agent breaks certificate validation on every HTTPS call.</li>
44+
</ul>
45+
<p><strong>The tell:</strong> multiple unrelated pipelines start failing at once, or the failure follows a specific runner/agent and clears on a different one. That is never your code.</p>
46+
47+
<h3>Flaky pipelines: the "fixed by re-run" trap</h3>
48+
<p>A job that passes on re-run without any change is <strong>not fixed</strong> — it is flaky, and flakiness erodes trust in the whole pipeline until people reflexively re-run real failures too. Root-cause the flake:</p>
49+
<ul>
50+
<li><strong>Test flakiness</strong> — order dependence, shared fixtures, real time/network in tests, unmocked external calls. Isolate and quarantine, then fix; don't let it mask regressions.</li>
51+
<li><strong>Concurrency races</strong> — parallel jobs sharing a database, a port, a fixed resource name, or the same Terraform state lock (see the state-locking guide).</li>
52+
<li><strong>Resource starvation</strong> — the runner was momentarily out of memory/disk; intermittent by nature.</li>
53+
</ul>
54+
55+
<h3>A triage checklist</h3>
56+
<ol>
57+
<li>Read the <strong>failing stage</strong>, not the whole log — it names the layer.</li>
58+
<li>Ask <strong>"did this commit pass before?"</strong> to split code-vs-environment.</li>
59+
<li>Check whether <strong>other pipelines</strong> are failing too (points to runner/infra).</li>
60+
<li>Reproduce the <strong>exact command</strong> locally with pinned versions.</li>
61+
<li>If intermittent, treat as <strong>flaky</strong> and root-cause — never accept a re-run as the fix.</li>
62+
</ol>
63+
64+
<h3>Common wrong approaches</h3>
65+
<ul>
66+
<li><strong>Re-running until green.</strong> Hides flakiness and burns CI minutes; the failure returns.</li>
67+
<li><strong>Editing application code to satisfy an infrastructure failure.</strong> e.g. loosening a test because the registry was down.</li>
68+
<li><strong>Widening cache keys to "speed things up."</strong> Broad keys share poisoned caches across branches.</li>
69+
<li><strong>Unpinned base images and toolchains.</strong> Guarantees future "green commit suddenly fails" incidents.</li>
70+
</ul>
71+
72+
<h3>Related resources</h3>
73+
<ul>
74+
<li><a href="/blog/terraform-state-locking-drift-troubleshooting/">Terraform state locking and drift</a> — a frequent source of pipeline concurrency failures.</li>
75+
<li><a href="/blog/kubernetes-pod-pending-troubleshooting/">Kubernetes pod Pending decision tree</a> — when the "runner" is a Kubernetes-based build pod that won't schedule.</li>
76+
<li><a href="/devops-job-support-guide/">DevOps job support guide</a> and <a href="/sre-job-support-guide/">SRE job support guide</a>.</li>
77+
</ul>
78+
79+
<p>If a release pipeline is red and blocking a deploy right now, <a href="/proxy-job-support/">real-time proxy job support</a> can help you localise the layer and unblock the release. And explaining a CI/CD failure clearly — which layer, why, how you'd prevent it — is exactly the kind of scenario a <a href="/devops-proxy-interview-support/">DevOps proxy interview support</a> session prepares you to handle.</p>
80+
81+
<p><em>Last reviewed: September 2026.</em></p>
Lines changed: 15 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,15 @@
1+
export const meta = {
2+
slug: "cicd-pipeline-failure-troubleshooting",
3+
permalink: "/blog/cicd-pipeline-failure-troubleshooting/",
4+
title: "CI/CD Pipeline Failures: Application vs Infrastructure vs Runner",
5+
description: "A triage method for red pipelines — how to tell whether a CI/CD failure is your application, the pipeline infrastructure, or the runner/agent, so you fix the real layer instead of blindly re-running the job.",
6+
date: "2026-09-07",
7+
keywords: "ci cd pipeline failure, pipeline debugging, github actions failure, gitlab ci, jenkins agent, flaky pipeline, runner out of disk, docker build cache, deployment failure vs build failure",
8+
about: "CI/CD pipeline failure triage",
9+
faqs: [
10+
{ q: "How do I tell if a pipeline failure is my code or the infrastructure?", a: "Ask whether the same commit failed before. If a previously-green commit now fails with no code change, suspect the environment — runner, cache, registry, credentials, or a moved dependency. If the failure appeared exactly on the commit that changed the relevant code and reproduces locally, it is the application. The failing stage also localises it: compile/test failures are usually application; checkout/dependency-fetch/push/deploy failures are usually infrastructure or runner." },
11+
{ q: "Why does re-running the job sometimes fix it?", a: "Because the failure was non-deterministic — a flaky test, a network blip fetching a dependency, a race in parallel jobs, or a runner that was momentarily out of disk or memory. A re-run that passes does not mean the problem is gone; it means you have a flaky pipeline that will fail again. Treat a 'fixed by re-run' as a bug to root-cause, not a resolution." },
12+
{ q: "What are the most common runner/agent causes of failure?", a: "Out of disk (Docker layers and build caches fill the agent), out of memory (parallel jobs or a heavy build OOM the agent), expired or missing credentials/secrets, a stale or poisoned build cache, clock skew breaking TLS, and version drift between the runner image and what the build expects. These fail regardless of your code and often affect many pipelines at once." },
13+
{ q: "How should I structure a pipeline to make failures debuggable?", a: "Fail fast with clear stage boundaries (lint, build, test, package, deploy) so the failing stage names the layer; make steps idempotent and re-runnable; pin tool and base-image versions; separate application tests from infrastructure/deploy steps; surface logs and artifacts on failure; and quarantine known-flaky tests instead of letting them mask real regressions." },
14+
],
15+
} as const;
Lines changed: 15 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,15 @@
1+
import fs from 'fs';
2+
import path from 'path';
3+
import BlogArticleShell from '@/components/BlogArticleShell';
4+
5+
export default function Article() {
6+
const html = fs.readFileSync(
7+
path.join(process.cwd(), 'content/blog-articles', "how-to-explain-your-project-in-a-technical-interview", 'body.html'),
8+
'utf8'
9+
);
10+
return (
11+
<BlogArticleShell>
12+
<div dangerouslySetInnerHTML={{ __html: html }} />
13+
</BlogArticleShell>
14+
);
15+
}
Lines changed: 67 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,67 @@
1+
<p>"Tell me about a project you're proud of" sounds like an easy question and is one of the hardest rounds in a senior loop. It is where the interviewer decides whether you are an <strong>owner</strong> who understands what they built and why, or an <strong>observer</strong> who was nearby while it happened. The difference isn't the size of the project — it's the depth you can defend under follow-up pressure. This is a structure for the walkthrough, and a survival guide for the questions that come after.</p>
2+
3+
<hr>
4+
5+
<h3>Pick the right project first</h3>
6+
<p>Before structure, selection. Choose the project you <strong>know best</strong>, not the most impressive-sounding one. The deep-dive round rewards depth of ownership, and an interviewer probing a flagship system you only touched at the edges will find the gap in minutes. A modest service you designed, debugged and operated end to end gives you a defensible answer to every follow-up — which reads as far stronger than name-dropping something huge you can't explain at the component level.</p>
7+
<p><strong>Good candidate projects:</strong> something you made real decisions on, that had a genuine problem or failure, and that produced a measurable outcome. Avoid pure maintenance work and projects where your role was peripheral.</p>
8+
9+
<h3>The layered structure</h3>
10+
<p>Do not narrate the project chronologically from kickoff to launch — that buries the signal in history. Lead with the <em>shape</em>, then go deep where pulled:</p>
11+
<ol>
12+
<li><strong>Context in one sentence.</strong> What the system did, who used it, and the scale that made it non-trivial. "A real-time fraud-scoring service handling ~5k requests/sec for our payments flow."</li>
13+
<li><strong>Your role and scope.</strong> Be precise and honest — what <em>you</em> owned versus the team. Interviewers score your contribution, not the org's.</li>
14+
<li><strong>Architecture at a high level.</strong> The main components and how data flows. A clean 60-second overview that gives the interviewer hooks to probe.</li>
15+
<li><strong>The two or three decisions that mattered.</strong> This is the core. Each decision, the alternative you rejected, and the trade-off.</li>
16+
<li><strong>A real problem or failure.</strong> Something that broke, how you diagnosed it, how you fixed it, what you changed afterward.</li>
17+
<li><strong>The result.</strong> A measurable outcome — latency, cost, reliability, delivery — tied back to the work.</li>
18+
</ol>
19+
<p>Steps 4 and 5 are where senior candidates win or lose. Everything else is setup.</p>
20+
21+
<h3>Decisions and trade-offs are the signal</h3>
22+
<p>Features are forgettable; decisions are memorable. Compare:</p>
23+
<table>
24+
<thead><tr><th>Forgettable (feature)</th><th>Memorable (decision + trade-off)</th></tr></thead>
25+
<tbody>
26+
<tr><td>"We used Kafka."</td><td>"We chose Kafka over a simple queue because we needed event replay and multiple independent consumers, accepting the operational overhead of running it."</td></tr>
27+
<tr><td>"It was a microservices architecture."</td><td>"We split two services out of the monolith along the boundaries that scaled differently, and deliberately kept the rest together to avoid distributed-transaction pain."</td></tr>
28+
<tr><td>"We cached results."</td><td>"We added a read-through cache with a 30s TTL because the data tolerated brief staleness, trading consistency for a big drop in DB load."</td></tr>
29+
</tbody>
30+
</table>
31+
<p>The pattern: <strong>name the alternative you rejected and why.</strong> That single move demonstrates the judgment the round exists to assess.</p>
32+
33+
<h3>Surviving the follow-up questions</h3>
34+
<p>The interviewer will probe every claim. Common lines of attack and how to handle them:</p>
35+
<ul>
36+
<li><strong>"Why that database/queue/language?"</strong> — Have the trade-off ready. If the honest answer is "it was already in the stack," say that <em>and</em> whether it turned out to be a good fit.</li>
37+
<li><strong>"What broke in production?"</strong> — Have a real incident: symptom, how you localised it, the fix, the follow-up. This is often the highest-signal answer in the whole round.</li>
38+
<li><strong>"What would you change now?"</strong> — Shows growth. A candidate who thinks their design was flawless reads as junior. Name a specific thing and why.</li>
39+
<li><strong>"How would you scale this 10×?"</strong> — Reason about the bottleneck that hits first (usually the write path or a hot resource), not a generic "add more servers."</li>
40+
<li><strong>"Walk me through the exact request path."</strong> — The ownership test. If you can trace a request from entry to storage and back, you built it; if you hand-wave, you didn't.</li>
41+
</ul>
42+
43+
<h3>When you hit the edge of what you know</h3>
44+
<p>You will get a question past your knowledge. The move that keeps you strong is honesty plus reasoning: <em>"I didn't own that part, but here's how I'd reason about it…"</em> or <em>"I don't know for certain — my mental model is X, let me reason from there."</em> Interviewers detect fabricated depth almost instantly, and a confident wrong bluff costs far more than a graceful boundary. Reasoning from first principles at the edge of your knowledge is itself a positive signal.</p>
45+
46+
<h3>Common wrong approaches</h3>
47+
<ul>
48+
<li><strong>Chronological narration.</strong> "First we gathered requirements, then we…" — the interviewer stops listening before you reach the interesting part.</li>
49+
<li><strong>Feature tours.</strong> Listing what the system did without any decision or trade-off behind it.</li>
50+
<li><strong>Overclaiming scope.</strong> "We" everywhere, so the interviewer can't tell what you did — then can't credit you for it.</li>
51+
<li><strong>Choosing the impressive project over the owned one.</strong> Depth you can't defend is a liability.</li>
52+
<li><strong>No numbers.</strong> "It worked well" — versus "cut p99 from 800ms to 120ms." One is a claim, the other is evidence.</li>
53+
</ul>
54+
55+
<h3>Prepare it in advance</h3>
56+
<p>This round is rehearsable. For your two or three best projects, pre-write: the one-sentence context, the three key decisions with rejected alternatives, one real failure story, the measurable result, and answers to the five follow-ups above. Say them out loud until the structure is automatic and you can go deep on any branch the interviewer chooses.</p>
57+
58+
<h3>Related resources</h3>
59+
<ul>
60+
<li><a href="/blog/what-happens-in-senior-technical-interview/">What actually happens in a senior technical interview</a> — where this round fits in the loop.</li>
61+
<li>Depth to draw on for real decisions and failures: <a href="/blog/why-rag-retrieval-quality-drops-in-production/">production RAG debugging</a>, the <a href="/blog/kubernetes-oomkilled-troubleshooting/">OOMKilled root-cause guide</a>, and <a href="/blog/spark-job-slow-data-skew-troubleshooting/">Spark data-skew troubleshooting</a>.</li>
62+
<li><a href="/final-round-interview-support-guide/">Final round interview support guide</a> and <a href="/how-live-technical-interview-support-works/">how live technical interview support works</a>.</li>
63+
</ul>
64+
65+
<p>If you have a project deep-dive round coming up and want a senior engineer to help you structure and pressure-test your walkthrough — and to support you live during the interview itself — that is exactly what <a href="/proxy-interview-support/">proxy interview support</a> provides, always framed around your own participation.</p>
66+
67+
<p><em>Last reviewed: September 2026.</em></p>

0 commit comments

Comments
 (0)