From 12a955a7a974bf80666edb2b1813049d9bbd7cba Mon Sep 17 00:00:00 2001 From: Mridul Chopra Date: Fri, 14 Aug 2026 14:56:51 -0400 Subject: [PATCH] feat(skills): Add eks-node-join-diagnostic skill Diagnoses why EKS worker nodes fail to join or register with a cluster. Performs 64 read-only, API-side checks across 25 failure domains (IAM/auth, security groups, VPC/DNS, bootstrap, AMI, STS, Karpenter, containerd/runtime, VPC CNI IP management, kernel/OS compatibility, credential lifecycle, control plane pressure, launch template edge cases, admission webhooks, EKS addon conflicts, network edge cases, instance tagging, and control plane audit logs) without requiring SSM/SSH access to the node. Uses adaptive depth routing: primary checks (NJ1-NJ27) stay inline in SKILL.md for the common case, extended checks (NJ28-NJ64) load from references/extended-checks.md only when needed. Includes a complementary handoff table to the aws-eks-node-diagnostics-mcp server's node-side log tools for confirmation when SSM is available. Validated against a live EKS Auto Mode cluster: correctly adapted NJ3/NJ4 for Auto Mode (no managed node groups expected) and correctly diagnosed a dual API_AND_CONFIG_MAP auth-mode conflict. Also updates llms.txt with the new skill entry. --- llms.txt | 1 + .../eks-node-join-diagnostic/.skilleval.yaml | 6 + skills/eks-node-join-diagnostic/CHANGELOG.md | 94 +++ skills/eks-node-join-diagnostic/README.md | 177 +++++ skills/eks-node-join-diagnostic/SKILL.md | 397 ++++++++++ .../evals/benchmark.json | 693 ++++++++++++++++++ .../evals/eval_queries.json | 82 +++ .../eks-node-join-diagnostic/evals/evals.json | 66 ++ .../evals/report.json | 65 ++ .../evals/trigger_report.json | 248 +++++++ .../references/extended-checks.md | 203 +++++ 11 files changed, 2032 insertions(+) create mode 100644 skills/eks-node-join-diagnostic/.skilleval.yaml create mode 100644 skills/eks-node-join-diagnostic/CHANGELOG.md create mode 100644 skills/eks-node-join-diagnostic/README.md create mode 100644 skills/eks-node-join-diagnostic/SKILL.md create mode 100644 skills/eks-node-join-diagnostic/evals/benchmark.json create mode 100644 skills/eks-node-join-diagnostic/evals/eval_queries.json create mode 100644 skills/eks-node-join-diagnostic/evals/evals.json create mode 100644 skills/eks-node-join-diagnostic/evals/report.json create mode 100644 skills/eks-node-join-diagnostic/evals/trigger_report.json create mode 100644 skills/eks-node-join-diagnostic/references/extended-checks.md diff --git a/llms.txt b/llms.txt index 3b5a275..6880789 100644 --- a/llms.txt +++ b/llms.txt @@ -24,6 +24,7 @@ Skills can be used with these AWS DevOps Agent types: - [Wiz Security Context Skill](skills/wiz-security-context/SKILL.md): Queries the Wiz MCP server for a resource's security context (vulnerabilities, misconfigurations, secrets, active threats, malware, toxic combinations) to determine whether an operational anomaly is an operational issue or a security incident - [Service Quota Check Skill](skills/service-quota-check/SKILL.md): Checks AWS service quota utilization during investigations and before provisioning resources, flags quotas at 85%+ utilization, and requests increases via the Service Quotas API or recommends support cases - [DMS Operational Review Skill](skills/database-migration-service-expertise/SKILL.md): Conducts AWS Database Migration Service operational reviews with 5-category health scoring, task failure troubleshooting, migration cutover runbooks, version deprecation tracking, and cost optimization +- [EKS Node Join Diagnostic Skill](skills/eks-node-join-diagnostic/SKILL.md): Diagnoses why EKS worker nodes fail to join or register with a cluster — 64 read-only API-side checks across 25 failure domains (IAM, security groups, VPC, bootstrap, AMI, containerd, CNI, kernel, credentials, webhooks, addons, network edge cases, instance tagging, control plane audit logs) with adaptive depth routing and ranked remediation ## Key Concepts diff --git a/skills/eks-node-join-diagnostic/.skilleval.yaml b/skills/eks-node-join-diagnostic/.skilleval.yaml new file mode 100644 index 0000000..042de93 --- /dev/null +++ b/skills/eks-node-join-diagnostic/.skilleval.yaml @@ -0,0 +1,6 @@ +name: eks-node-join-diagnostic +version: "1.0.0" +zip_command: >- + cd skills && zip -r eks-node-join-diagnostic.zip eks-node-join-diagnostic/ + -i '*.md' '*.txt' '*.json' '*.yaml' '*.yml' + -x '*/.claude/*' '*/scripts/*' '*/README.md' '*/.skilleval.yaml' '*/CHANGELOG.md' '*/evals/*' diff --git a/skills/eks-node-join-diagnostic/CHANGELOG.md b/skills/eks-node-join-diagnostic/CHANGELOG.md new file mode 100644 index 0000000..2d2962a --- /dev/null +++ b/skills/eks-node-join-diagnostic/CHANGELOG.md @@ -0,0 +1,94 @@ +# Changelog + +## 3.1.0 + +- Added 3 checks identified by cross-referencing the companion + `aws-eks-node-diagnostics-mcp` skill's A4 worker-node-join-failure SOP: + - **Instance Tagging (NJ62):** `kubernetes.io/cluster/` tag required + for self-managed and Karpenter nodes (managed node groups set it + automatically) + - **Control Plane Audit Logs (NJ63–NJ64):** CSR denial detection via + control-plane audit logs, and detection of recent aws-auth/access entry + removal as a root cause for sudden fleet-wide join failures +- Added a "Complementing with EKS Node Diagnostics MCP Server" section with a + handoff table mapping primary/extended checks to specific MCP tool calls + (`collect`, `quick_triage`, `search`, `network_diagnostics`) for node-side + log confirmation when the MCP server is available +- Fixed SKILL.md frontmatter `description` exceeding the 1024-character + DevOps Agent upload limit; trimmed while preserving trigger keywords and + negative-boundary exclusions +- Corrected check-count and failure-domain totals across SKILL.md, README.md, + and references/extended-checks.md (58→64 checks, 24→25 domains, NJ1–NJ64) +- Confirmed via live testing against a real EKS Auto Mode cluster: primary + checks (NJ1, NJ3, NJ5–NJ8) correctly adapted for Auto Mode (no managed node + groups is expected) and correctly diagnosed a dual `API_AND_CONFIG_MAP` + auth-mode conflict + +## 3.0.0 + +- Restructured for progressive disclosure: SKILL.md reduced from 673 to 331 lines +- Moved extended checks (NJ28–NJ61) to `references/extended-checks.md` +- Added **Adaptive Depth** routing: + - Fast path: immediate root cause identification for clear-cut symptoms + - Standard path: full NJ1–NJ27 for ambiguous symptoms + - Deep path: loads extended checks only when primary checks pass +- Added symptom-to-check routing table for direct jump to extended checks +- Strengthened differentiators vs baseline model knowledge: + - Explicit "self-managed needs auth mapping" constraint + - Always recommend EKS Access Entries over aws-auth + - Structured report with coverage gaps ("What Was Not Checked") +- Optimized description for trigger accuracy with explicit negative boundaries +- Audit score improved: 66/100 → 74/100 + +## 2.0.0 + +- Expanded from 28 to 58 diagnostic checks across 24 failure domains +- New failure domains added from analysis of 100 real-world support cases: + - **Containerd / Runtime Configuration (NJ31–NJ34):** registry mirror validation, + cgroup driver alignment, snapshotter compatibility, instance store mounting + - **VPC CNI / IP Address Management (NJ35–NJ37):** IPAMD warm pool vs subnet + capacity, CNI migration chicken-and-egg detection, CIDR overlap with + on-premises networks + - **Kernel / OS-Level Compatibility (NJ38–NJ41):** kernel version requirements + per EKS version, SELinux/AppArmor blocking detection, required kernel modules + - **Credential / Certificate Lifecycle (NJ42–NJ44):** SSM hybrid activation + expiry, kubelet cert expiration after hibernation, cluster credential rotation + - **Control Plane Pressure (NJ45–NJ46):** etcd database size correlation, + API throttling during mass scaling events + - **Launch Template Edge Cases (NJ47–NJ50):** user-data 16KB size limit, + blocking package updates in runtime, network interface + subnet conflicts, + stale cloud-init cache detection + - **Admission Webhooks (NJ51–NJ52):** webhook failurePolicy:Fail blocking + node registration, PDB blocking node group updates + - **EKS Add-on Conflicts (NJ53–NJ55):** kube-proxy version compatibility, + GPU Operator + EKS device plugin conflict, root volume sizing + - **Network Edge Cases (NJ56–NJ59):** NAT gateway idle timeout, ENI detach + by third-party tools, conntrack table exhaustion, S3 endpoint policy + blocking ECR + - **Account / Organization Changes (NJ60–NJ61):** account migration between + Organizations, kubelet credential provider path validation for AL2023 +- Updated decision tree with all new failure domains +- Expanded trigger keywords for better skill activation +- Updated "When to Use" section with new symptom patterns + +## 1.0.0 + +- Initial release +- 30 diagnostic checks across 14 failure domains: + - Cluster state (NJ1–NJ2) + - Node group health (NJ3–NJ4) + - IAM/authentication mapping (NJ5–NJ8) + - Security groups (NJ9–NJ10) + - Network connectivity and DNS (NJ11–NJ15) + - VPC endpoints for private clusters (NJ16) + - Bootstrap/user-data validation (NJ17–NJ20) + - AMI compatibility (NJ21–NJ23) + - Regional STS endpoint (NJ24) + - Karpenter configuration (NJ25–NJ27) + - VPC CNI scheduling (NJ28) + - Join timeout / burstable instance delay (NJ29) + - IMDS reachability / hop limit / host firewall (NJ30) +- Decision tree for quick root-cause identification +- Covers managed node groups, self-managed, Karpenter, and Auto Mode +- No SSM/SSH required — pure API-side diagnosis +- Diagnostic report artifact with ranked remediation diff --git a/skills/eks-node-join-diagnostic/README.md b/skills/eks-node-join-diagnostic/README.md new file mode 100644 index 0000000..da2d6b5 --- /dev/null +++ b/skills/eks-node-join-diagnostic/README.md @@ -0,0 +1,177 @@ +# EKS Node Join Diagnostic — AWS DevOps Agent Skill + +A focused Amazon EKS node registration diagnostic skill for [AWS DevOps Agent](https://docs.aws.amazon.com/devopsagent/latest/userguide/about-aws-devops-agent.html). Identifies why worker nodes fail to join an EKS cluster by running 64 read-only API-side configuration checks across 25 failure domains, returning a ranked pass/fail checklist with specific remediation for each failure. + +## What It Does + +When activated via Chat or during Incident RCA, this skill instructs the DevOps Agent to: + +1. Gather cluster and node group context via AWS APIs. +2. Route to the appropriate diagnostic depth using **Adaptive Depth**: + - **Fast path** — symptom maps directly to known root cause → validate subset, skip unrelated checks + - **Standard path** — ambiguous symptoms → run primary checks (NJ1–NJ27), stop at first Critical failure + - **Deep path** — primary checks pass → load extended checks (NJ28–NJ64) for runtime, kernel, credential, network edge cases, instance tagging, and control plane audit logs +3. Identify the specific failure mode from 64 known root causes. +4. Generate a diagnostic report with a ranked remediation plan. + +## Key Design Decisions + +- **No SSM required.** Diagnoses entirely from the AWS API side. Complements the `aws-eks-node-diagnostics-mcp` (which requires SSM access to the node for log-based diagnosis). +- **Covers all node types:** Managed node groups, self-managed nodes, Karpenter-provisioned nodes, and EKS Auto Mode. +- **64 checks across 25 failure domains** — based on analysis of 100+ real-world EKS support cases and cross-referenced against the companion `aws-eks-node-diagnostics-mcp` skill's A4 worker-node-join-failure SOP for coverage gaps. +- **Progressive disclosure** — primary checks inline in SKILL.md (~400 lines), extended checks in `references/extended-checks.md`. Keeps context lean for common cases. +- **Always recommends EKS Access Entries** over aws-auth for new configurations. + +## Agent Types + +- **Chat tasks** — interactive "why won't my node join?" investigation +- **Incident RCA** — when node failures contribute to a broader incident + +## Prerequisites + +### IAM Permissions + +The DevOps Agent role needs read-only access (most covered by `AIDevOpsAgentAccessPolicy`): + +``` +eks:DescribeCluster +eks:DescribeNodegroup +eks:ListAccessEntries +eks:DescribeAccessEntry +ec2:DescribeSecurityGroups +ec2:DescribeSubnets +ec2:DescribeVpcs +ec2:DescribeVpcEndpoints +ec2:DescribeRouteTables +ec2:DescribeNetworkAcls +ec2:DescribeDhcpOptions +ec2:DescribeLaunchTemplateVersions +ec2:DescribeImages +iam:ListAttachedRolePolicies +iam:GetRole +sts:GetCallerIdentity +``` + +### Optional (for deeper checks) + +- **kubectl access** — for checking aws-auth ConfigMap, Karpenter resources, VPC CNI DaemonSet, admission webhooks, PDBs, addon versions +- **EKS cluster access entry** — DevOps Agent role configured with cluster access +- **CloudWatch access** — `cloudwatch:GetMetricData` for etcd metrics (NJ45) and API throttling (NJ46) +- **CloudTrail access** — `cloudtrail:LookupEvents` for credential rotation detection (NJ44), ENI detach events (NJ57), account migration (NJ60) + +## Uploading to AWS DevOps Agent + +### Package the skill + +```bash +cd skills +zip -r eks-node-join-diagnostic.zip eks-node-join-diagnostic/ \ + -i '*.md' '*.txt' '*.json' '*.yaml' '*.yml' \ + -x '*/.claude/*' '*/scripts/*' '*/README.md' '*/.skilleval.yaml' '*/CHANGELOG.md' '*/evals/*' +``` + +### Upload via the Operator Web App + +1. Navigate to the Skills page in your Agent Space. +2. Click **Add skill** → **Upload skill**. +3. Drag and drop `eks-node-join-diagnostic.zip`. +4. Select agent types: **Chat tasks** and **Incident RCA**. +5. Click **Upload**. + +## Usage + +In the DevOps Agent Chat: + +- "My managed node group `prod-nodes` in cluster `prod-eks` shows Create failed. Diagnose why." +- "Nodes are not joining my private EKS cluster `staging` in `us-west-2`." +- "I launched self-managed nodes but they don't appear in kubectl get nodes." +- "Karpenter is provisioning nodes but they stay NotReady." +- "I'm getting TLS handshake timeout when nodes try to join." +- "After upgrading to 1.33, my new node group can't register." +- "Nodes joining but immediately NotReady with cgroup driver mismatch errors." +- "VPC CNI IPAMD is crash-looping on new nodes — warm pool exhaustion." +- "Nodes intermittently going NotReady behind NAT gateway." +- "Our custom AMI with kernel 5.4 stopped working after upgrading to EKS 1.31." +- "Admission webhook is timing out and blocking new nodes from registering." +- "After cluster credential rotation, existing nodes can't communicate." +- "Cloud-init seems stuck — nodes never start kubelet." +- "GPU nodes failing after installing both GPU Operator and EKS nvidia plugin." + +## Skill Contents + +``` +eks-node-join-diagnostic/ +├── SKILL.md # Primary checks NJ1–NJ27 + adaptive routing (~400 lines) +├── README.md # This file +├── CHANGELOG.md # Version history +├── references/ +│ └── extended-checks.md # Extended checks NJ28–NJ64 (loaded on demand) +└── evals/ + ├── evals.json # Functional test scenarios + ├── eval_queries.json # Trigger accuracy tests + └── benchmark.json # Eval benchmark results +``` + +## Failure Domains Covered + +| Domain | Checks | Common Symptom | +|--------|--------|----------------| +| Cluster State | NJ1–NJ2 | Cluster not ACTIVE or mid-upgrade | +| Node Group Health | NJ3–NJ4 | health.issues present, Create failed | +| IAM / Authentication | NJ5–NJ8 | Unauthorized, missing policies, role path issue | +| Security Groups | NJ9–NJ10 | TLS handshake timeout, connection refused | +| Network / DNS | NJ11–NJ15 | Node not found, no route to host, IP exhaustion | +| VPC Endpoints | NJ16 | Private cluster can't reach ECR/STS/API | +| Bootstrap / User-Data | NJ17–NJ20 | Cloud-init failure, wrong cluster name, AL2023 format | +| AMI Compatibility | NJ21–NJ23 | Version mismatch, AL2 on 1.33+ | +| STS Endpoint | NJ24 | InvalidClientTokenId | +| Karpenter | NJ25–NJ27 | Wrong subnets/SGs/AMI in EC2NodeClass | +| VPC CNI | NJ28 | "network plugin not ready: cni config uninitialized" | +| Join Timeout | NJ29 | Burstable instance slow to boot, MNG 20-min timeout | +| IMDS | NJ30 | Hop limit too low, host firewall blocking 169.254.169.254 | +| Containerd / Runtime | NJ31–NJ34 | Mirror misconfiguration, cgroup mismatch, snapshotter incompatibility | +| VPC CNI / IP Management | NJ35–NJ37 | IPAMD crash-loop, CNI migration deadlock, CIDR overlap | +| Kernel / OS Compatibility | NJ38–NJ41 | Kernel too old, SELinux/AppArmor blocking, missing modules | +| Credential Lifecycle | NJ42–NJ44 | SSM activation expired, cert expired after hibernation, CA rotated | +| Control Plane Pressure | NJ45–NJ46 | etcd compaction causing heartbeat loss, API throttling at scale | +| Launch Template Edge Cases | NJ47–NJ50 | User-data >16KB, package updates hang, NIC+subnet conflict, stale cache | +| Admission Webhooks | NJ51–NJ52 | Webhook Fail policy blocking nodes, PDB blocking drain | +| EKS Addon Conflicts | NJ53–NJ55 | kube-proxy version mismatch, GPU operator conflict, disk pressure | +| Network Edge Cases | NJ56–NJ59 | NAT idle timeout, ENI detach, conntrack full, S3 endpoint blocks ECR | +| Account / Org Changes | NJ60–NJ61 | Account moved between Orgs, credential provider path wrong | +| Instance Tagging | NJ62 | Self-managed/Karpenter node missing kubernetes.io/cluster tag | +| Control Plane Audit Logs | NJ63–NJ64 | CSR denied, aws-auth/access entry recently removed | + +## Relationship to Other Tools + +| Tool | What it does | When to use | +|------|-------------|-------------| +| **This skill** | API-side config validation (no node access needed) | Node never joins OR you can't SSM into it. Start here. | +| `aws-eks-node-diagnostics-mcp` | SSM-based log collection and analysis from the node | Node is SSM-reachable — use to confirm findings from this skill with actual log evidence | +| `AWSSupport-TroubleshootEKSWorkerNode` | SSM Automation runbook | Automated node-side checks (requires SSM) | +| `eks-operation-review` | Full cluster best-practices audit | Proactive review, not incident response | + +### Complementary Workflow + +This skill is designed to work standalone OR complement the EKS Node Diagnostics MCP server: + +1. **Always start with this skill** — it works without SSM and identifies the likely root cause via API checks +2. **If MCP server is available and root cause needs confirmation** — hand off to MCP tools (`collect` → `quick_triage` → `search`) for log-level evidence +3. **If all API checks pass but node still won't join** — the MCP's `quick_triage` tool with its A4 SOP can find issues only visible in node-side logs (e.g., kubelet crashloop, cloud-init failures, containerd errors) + +The SKILL.md includes a detailed handoff table showing which MCP tools to use after each NJ check identifies a suspected root cause. + +## Source Material + +- [EKS Troubleshooting docs](https://docs.aws.amazon.com/eks/latest/userguide/troubleshooting.html) +- [re:Post — Nodes fail to join cluster](https://repost.aws/knowledge-center/eks-nodes-fail-cluster-join) +- [re:Post — Worker nodes cluster](https://aws.amazon.com/premiumsupport/knowledge-center/eks-worker-nodes-cluster/) +- [EKS Security Group Requirements](https://docs.aws.amazon.com/eks/latest/userguide/sec-group-reqs.html) +- [Private Clusters](https://docs.aws.amazon.com/eks/latest/userguide/private-clusters.html) +- Project Nebula case pattern analysis (#1 Node Registration Failures, #7 Node Group Creation Failures — 2,770 cases/yr) +- Internal EKS Support Playbook (node registration checklist, known issues) +- Analysis of 100 real-world EKS node join failure support cases (Apr–Aug 2025) covering containerd, CNI, kernel, credential, webhook, addon, and network edge case patterns + +## License + +Apache-2.0. See [LICENSE](../../LICENSE). diff --git a/skills/eks-node-join-diagnostic/SKILL.md b/skills/eks-node-join-diagnostic/SKILL.md new file mode 100644 index 0000000..d5553c3 --- /dev/null +++ b/skills/eks-node-join-diagnostic/SKILL.md @@ -0,0 +1,397 @@ +--- +name: eks-node-join-diagnostic +description: >- + Use when an EKS worker node fails to join, register, or reach Ready state. + Activate for: nodes missing from kubectl get nodes, NodeCreationFailure, + node group Create failed/Degraded, TLS handshake timeout to API server, + Unauthorized kubelet errors, NotReady immediately after launch, cgroup/ + containerd/IPAMD/webhook/addon-conflict node join issues, or any "why won't + my node join?" question. Covers managed node groups, self-managed, Karpenter, + and Auto Mode. Read-only API-side diagnosis across IAM, security groups, + VPC/DNS, bootstrap, AMI, runtime, and CNI — ranked pass/fail checklist with + fixes. Do NOT activate for: cluster performance/throttling without join + failure, upgrade readiness, Karpenter setup/consolidation, pod scheduling on + already-joined nodes, ECS, cost analysis, or general security audits. +metadata: + author: LearningNewbie + version: "3.1.0" + aws-devops-agent-skills.agent-types: "Chat tasks, Incident RCA" + aws-devops-agent-skills.aws-services: "Amazon EKS" + aws-devops-agent-skills.technical-domains: "Containers" +--- + +# EKS Node Join Diagnostic + +Systematically identify which of 64 known failure modes is preventing a worker +node from joining an EKS cluster, and recommend the specific fix — all without +needing SSH/SSM access to the node itself. + +## Adaptive Depth + +Not every diagnosis needs all 61 checks. Use this routing: + +**Fast path** — If the symptom directly maps to a known root cause with high +confidence (e.g., user says "I have a pathed role ARN in aws-auth"), identify +the root cause immediately, then validate with only the relevant subset of +checks (the specific NJ ID + its prerequisites). Still produce the diagnostic +report format, but mark unrelated checks as "SKIPPED — not indicated by symptoms." + +**Standard path** — For ambiguous symptoms ("nodes not joining"), run the full +decision tree from Step 2 through Step 9 (NJ1–NJ27). Stop when you find a +Critical-severity failure — that's the blocker. + +**Deep path** — If all primary checks (NJ1–NJ27) pass, load +`references/extended-checks.md` and continue with NJ28–NJ64. These cover +runtime, kernel, credentials, control plane pressure, launch template edge +cases, webhooks, addon conflicts, network edge cases, instance tagging, and +control plane audit logs. + +## Critical Constraints + +- **Read-only.** Only `describe*`, `list*`, `get*` APIs. No mutating calls. + All fixes are recommendations for human execution. +- **No SSM required.** This skill diagnoses from the AWS API side. For + node-side log analysis (when SSM is available), use the + `aws-eks-node-diagnostics-mcp` instead — the two are complementary. +- **One cluster at a time.** Ask the user for cluster name, region, and + optionally the specific node group or instance ID. +- **Self-managed nodes need explicit auth mapping.** Unlike managed node + groups, self-managed and Karpenter nodes do NOT get automatic aws-auth + entries. Always flag NJ6 as high-priority for these node types. +- **Always recommend EKS Access Entries** over aws-auth for new configurations. + Access Entries handle IAM paths correctly, are API-managed, and are the + modern best practice (available since late 2023). + +## Step 1: Gather Context + +Ask the user for: +- **Cluster name** and **region** +- **What failed:** managed node group name, instance ID, or symptom +- **When:** approximate time the failure started +- **Node type:** managed node group, self-managed, Karpenter, or Auto Mode + +Then collect foundation data: +``` +aws eks describe-cluster --name --region +``` + +Extract: `status`, `version`, `endpoint`, `certificateAuthority`, +`resourcesVpcConfig` (subnetIds, securityGroupIds, clusterSecurityGroupId, +endpointPublicAccess, endpointPrivateAccess), `logging`, `encryptionConfig`. + +If a node group is specified: +``` +aws eks describe-nodegroup --cluster-name --nodegroup-name +``` + +Extract: `status`, `health.issues`, `amiType`, `instanceTypes`, `subnets`, +`remoteAccess`, `launchTemplate`, `version`, `releaseVersion`. + +## Step 2: Check Cluster State (NJ1–NJ2) + +| ID | Check | Pass Criteria | Severity | +|----|-------|---------------|----------| +| NJ1 | Cluster status is ACTIVE | `status == "ACTIVE"` | Critical — nodes cannot join a non-ACTIVE cluster | +| NJ2 | Cluster not mid-upgrade | No pending platform version update | Medium — join may stall during CP upgrade | + +If NJ1 fails, stop — the cluster itself has a problem. + +## Step 3: Check Node Group Health (NJ3–NJ4) + +| ID | Check | Pass Criteria | Severity | +|----|-------|---------------|----------| +| NJ3 | Node group health issues | `health.issues` array is empty | Critical | +| NJ4 | Node group status | Not `CREATE_FAILED` or `DEGRADED` | Critical | + +Common NJ3 issue codes and their meaning: +- `Ec2SecurityGroupNotFound` — SG was deleted → must recreate node group +- `Ec2LaunchTemplateNotFound` — LT was deleted → must recreate +- `IamInstanceProfileNotFound` — instance profile deleted → recreate +- `AutoScalingGroupNotFound` — ASG deleted → recreate node group +- `AsgInstanceLaunchFailures` — capacity unavailable → try different AZ/type +- `ClusterUnreachable` — etcd full or CP issue → check cluster health + +Note: Self-managed nodes have no EKS node group object — skip NJ3/NJ4. + +## Step 4: Check IAM and Authentication (NJ5–NJ8) + +| ID | Check | API | Pass Criteria | Severity | +|----|-------|-----|---------------|----------| +| NJ5 | Node IAM role has required policies | `iam.ListAttachedRolePolicies` | Has AmazonEKSWorkerNodePolicy + AmazonEKS_CNI_Policy + AmazonEC2ContainerRegistryReadOnly (or PullOnly) | Critical | +| NJ6 | Node role mapped in auth config | `eks.ListAccessEntries` or `kubectl get cm aws-auth -n kube-system` | Role ARN appears as EC2_LINUX access entry OR in aws-auth mapRoles | Critical | +| NJ7 | Role ARN format correct (aws-auth) | Inspect aws-auth ConfigMap | ARN has NO path. Use `role/MyRole` not `role/path/MyRole`. Instance profile ARN must NOT be used. | Critical | +| NJ8 | Cluster IAM role valid | `iam.GetRole` on cluster role | Has AmazonEKSClusterPolicy, trust allows `eks.amazonaws.com` | Critical | + +**NJ6 — auth mode detection:** +- EKS API auth mode: `aws eks list-access-entries --cluster-name ` — look for node role with type `EC2_LINUX` or `EC2_WINDOWS`. +- CONFIG_MAP mode: check aws-auth ConfigMap mapRoles for the node role with groups `system:bootstrappers` and `system:nodes`. + +**NJ7 — the path gotcha:** If role ARN is `arn:aws:iam::123456789012:role/development/apps/MyNodeRole`, it MUST be entered in aws-auth as `arn:aws:iam::123456789012:role/MyNodeRole` (path stripped). This does NOT apply to EKS Access Entries which handle paths correctly. Always recommend migrating to Access Entries. + +## Step 5: Check Security Groups (NJ9–NJ10) + +| ID | Check | Pass Criteria | Severity | +|----|-------|---------------|----------| +| NJ9 | Node SG allows outbound 443 to cluster SG | Node SG egress allows TCP 443 to cluster security group (or 0.0.0.0/0) | Critical | +| NJ10 | Cluster SG allows inbound from node SG | Cluster SG allows inbound TCP 443 from node SG AND TCP 10250 from node SG | Critical | + +Required bidirectional rules: +- Node → Cluster: TCP 443 (API server) +- Cluster → Node: TCP 10250 (kubelet), TCP 53 (DNS), ephemeral ports +- Node → Node: All traffic (pod communication) + +## Step 6: Check Network Connectivity (NJ11–NJ15) + +| ID | Check | Pass Criteria | Severity | +|----|-------|---------------|----------| +| NJ11 | VPC DNS support enabled | `enableDnsSupport = true` AND `enableDnsHostnames = true` | Critical | +| NJ12 | DHCP options have AmazonProvidedDNS | domain-name-servers includes AmazonProvidedDNS | Critical | +| NJ13 | Subnet route tables have internet/NAT/VPCe path | Private subnets → NAT GW or VPC endpoints; public → IGW | Critical | +| NJ14 | Subnet IP availability | AvailableIpAddressCount >= 5 per node subnet | High | +| NJ15 | NACLs allow required ports | Allow TCP 443, 10250, ephemeral (1025-65535) | High | + +**NJ11 is critical:** If DNS hostname is disabled, nodes cannot resolve their own hostname, causing "node not found" errors. After enabling, nodes must be REPLACED (reboot is not enough). + +## Step 7: Check VPC Endpoints (NJ16) — Private Clusters Only + +If `endpointPublicAccess = false`: + +| ID | Check | Pass Criteria | Severity | +|----|-------|---------------|----------| +| NJ16 | Required VPC endpoints exist | All required endpoints present with SG allowing inbound 443 from node subnets | Critical | + +Required VPC endpoints for private clusters: +- `com.amazonaws..ec2` (interface) +- `com.amazonaws..ecr.api` (interface) +- `com.amazonaws..ecr.dkr` (interface) +- `com.amazonaws..sts` (interface) +- `com.amazonaws..s3` (gateway) +- `com.amazonaws..logs` (interface, if logging enabled) + +Each interface endpoint's SG must allow inbound TCP 443 from node subnets. + +## Step 8: Check Bootstrap / User-Data (NJ17–NJ20) + +| ID | Check | Pass Criteria | Severity | +|----|-------|---------------|----------| +| NJ17 | Launch template user-data present | Not empty (for custom LT) | High | +| NJ18 | Private cluster args present | Must contain `--apiserver-endpoint`, `--b64-cluster-ca`, `--dns-cluster-ip` | Critical | +| NJ19 | Cluster name matches | ClusterName matches exactly (case-sensitive) | Critical | +| NJ20 | AL2023 uses nodeadm format | MIME multipart with `application/node.eks.aws`, NOT bash script | Critical | + +**NJ18:** For private clusters, bootstrap.sh cannot call `aws eks describe-cluster` without connectivity. Always pass args directly: +```bash +/etc/eks/bootstrap.sh \ + --apiserver-endpoint https:// \ + --b64-cluster-ca \ + --dns-cluster-ip +``` + +**NJ20:** AL2023 uses `nodeadm` instead of `bootstrap.sh`: +``` +MIME-Version: 1.0 +Content-Type: multipart/mixed; boundary="//" + +--// +Content-Type: application/node.eks.aws + +--- +apiVersion: node.eks.aws/v1alpha1 +kind: NodeConfig +spec: + cluster: + apiServerEndpoint: https://... + certificateAuthority: ... + cidr: 10.100.0.0/16 + name: my-cluster +--//-- +``` + +## Step 9: Check AMI and STS (NJ21–NJ24) + +| ID | Check | Pass Criteria | Severity | +|----|-------|---------------|----------| +| NJ21 | AMI K8s version matches cluster | Within supported skew (N-3 for 1.28+) | Critical | +| NJ22 | AL2 AMI not used on 1.33+ | amiType must NOT be AL2 on 1.33+ (use AL2023 or Bottlerocket) | Critical | +| NJ23 | Custom AMI has required components | kubelet, containerd, aws-iam-authenticator, bootstrap.sh/nodeadm | High | +| NJ24 | Regional STS endpoint activated | `sts..amazonaws.com` responds. "InvalidClientTokenId" if not. | High | + +## Step 10: Check Karpenter Configuration (NJ25–NJ27) + +Only if Karpenter is detected in the cluster: + +| ID | Check | Pass Criteria | Severity | +|----|-------|---------------|----------| +| NJ25 | EC2NodeClass subnets match cluster VPC | Subnets in same VPC with routes to API endpoint | High | +| NJ26 | EC2NodeClass SGs include cluster SG | Security groups include cluster SG or equivalent rules | Critical | +| NJ27 | EC2NodeClass AMI compatible | amiFamily not AL2 on 1.33+; amiSelectorTerms resolve to valid AMIs | High | + +## Step 11: Extended Checks (NJ28–NJ64) + +If all checks above PASS but nodes still aren't joining, read +`references/extended-checks.md` for the remaining 37 checks covering: + +- VPC CNI scheduling and taints (NJ28) +- Join timeout and IMDS hop limit (NJ29–NJ30) +- Containerd/runtime configuration (NJ31–NJ34) +- VPC CNI IP address management (NJ35–NJ37) +- Kernel/OS compatibility (NJ38–NJ41) +- Credential/certificate lifecycle (NJ42–NJ44) +- Control plane pressure (NJ45–NJ46) +- Launch template edge cases (NJ47–NJ50) +- Admission webhooks (NJ51–NJ52) +- EKS add-on conflicts (NJ53–NJ55) +- Network edge cases (NJ56–NJ59) +- Account/credential provider (NJ60–NJ61) +- Instance tagging (NJ62) +- Control plane audit logs (NJ63–NJ64) + +Route to extended checks immediately (skip primary checks) when symptoms match: +- "cgroup driver mismatch" or immediate NotReady after join → NJ32 +- "cni config uninitialized" → NJ28 +- IPAMD crash-looping → NJ35 +- Node was stopped for weeks/months → NJ43 +- Recently rotated cluster credentials → NJ44 +- Intermittent NotReady (30s episodes) → NJ56 +- AL2023 image pull failures → NJ61 +- Self-managed/Karpenter node not discovered by cluster → NJ62 +- CSR denied or aws-auth/access entry recently changed → NJ63–NJ64 + +## Diagnostic Report Format + +Produce a report artifact named `eks-node-join-diagnostic--.md`: + +```markdown +# EKS Node Join Diagnostic — +**Region:** | **Date:** | **K8s Version:** +**Target:** +**Diagnosis:** PASS / FAIL (blocker found) / PARTIAL (warnings only) + +## Root Cause + + +## Checklist Results + +| # | Check | Status | Evidence | Fix | +|---|-------|--------|----------|-----| +| NJ1 | Cluster ACTIVE | ✅ PASS | status=ACTIVE | — | +| NJ5 | Node role policies | ❌ FAIL | Missing AmazonEKS_CNI_Policy | Attach policy | +| ... | ... | SKIPPED | Not indicated by symptoms | — | + +## Recommended Actions (priority order) +1. [CRITICAL] +2. [HIGH] + +## What Was Not Checked +- Node-side logs (requires SSM — use aws-eks-node-diagnostics-mcp) +- +``` + +## Decision Tree + +``` +Node not joining +│ +├─ Symptom maps to known root cause? ──→ FAST PATH: validate + report +│ +├─ NJ3/NJ4: Node group health issues? ──→ Recreate if SG/LT/IAM deleted +│ +├─ NJ5-NJ8: IAM/Auth failure? +│ ├─ Missing policies ──→ Attach required managed policies +│ ├─ Role not in auth config ──→ Create access entry (EC2_LINUX) +│ └─ Role has path (aws-auth) ──→ Strip path OR migrate to Access Entries +│ +├─ NJ9-NJ10: Security group blocking? +│ └─ Missing 443/10250 rules ──→ Update SG rules +│ +├─ NJ11-NJ15: Network connectivity? +│ ├─ DNS disabled ──→ Enable + REPLACE nodes +│ ├─ No route to API ──→ Add NAT GW or VPC endpoints +│ └─ Subnet full ──→ Add IPs or different subnets +│ +├─ NJ16: Private cluster missing VPC endpoints? ──→ Create endpoints +│ +├─ NJ17-NJ20: Bootstrap misconfigured? +│ ├─ Missing private cluster args ──→ Add --apiserver-endpoint, etc. +│ └─ AL2023 wrong format ──→ Convert to nodeadm MIME multipart +│ +├─ NJ21-NJ24: AMI/STS issues? +│ ├─ AL2 on 1.33+ ──→ Migrate to AL2023 or Bottlerocket +│ └─ STS not activated ──→ Activate in IAM console +│ +├─ NJ25-NJ27: Karpenter misconfigured? ──→ Fix EC2NodeClass spec +│ +└─ All primary checks pass? ──→ Load references/extended-checks.md (NJ28-NJ64) +``` + +## References + +- EKS Troubleshooting: https://docs.aws.amazon.com/eks/latest/userguide/troubleshooting.html +- Node join failures: https://repost.aws/knowledge-center/eks-nodes-fail-cluster-join +- Security group requirements: https://docs.aws.amazon.com/eks/latest/userguide/sec-group-reqs.html +- Private clusters: https://docs.aws.amazon.com/eks/latest/userguide/private-clusters.html +- AWSSupport-TroubleshootEKSWorkerNode runbook: https://docs.aws.amazon.com/systems-manager-automation-runbooks/latest/userguide/automation-awssupport-troubleshooteksworkernode.html +- EKS Node Diagnostics MCP server: https://github.com/aws-samples/sample-eks-node-diagnostics-mcp + +## Complementing with EKS Node Diagnostics MCP Server + +This skill performs API-side diagnosis and does NOT require SSM access. However, +when the `aws-eks-node-diagnostics-mcp` MCP server is deployed in the Agent Space +and the node is SSM-reachable, hand off to MCP tools for node-side confirmation: + +### When to Hand Off + +| After This Skill Finds... | Use MCP Tool | To Confirm | +|---|---|---| +| NJ5-NJ8: IAM/auth suspected but API checks pass | `collect` → `search` with query=`Unauthorized\|401\|Forbidden` | Actual kubelet auth error in logs | +| NJ9-NJ10: SG rules look correct but TLS timeout | `collect` → `network_diagnostics` sections=routes,eni | Actual route table and ENI config on-box | +| NJ17-NJ20: Bootstrap args suspected wrong | `collect` → `search` with query=`bootstrap.sh\|cloud-init\|ClusterName` | What args were actually passed at boot | +| NJ31-NJ34: Containerd/runtime config issue | `collect` → `search` with query=`containerd\|cgroup\|SystemdCgroup` | Actual containerd config and kubelet cgroup driver | +| NJ38-NJ41: Kernel/OS issue suspected | `collect` → `search` with query=`kernel\|modprobe\|SELinux\|AppArmor` | Actual kernel version and module state | +| NJ45: etcd pressure / heartbeat loss | `collect` → `quick_triage` on affected node | Timeline of kubelet connection failures | +| All primary checks PASS, cause unclear | `collect` → `quick_triage` for fast root cause | Let MCP's log analysis find what API checks can't see | + +### Workflow Pattern + +``` +1. Run this skill (API-side checks NJ1–NJ64) + ↓ +2. If root cause confirmed with high confidence → report + fix + ↓ +3. If root cause unclear OR need log-level confirmation: + → collect(instanceId="", region="") + → status(executionId="") # poll until Success + → quick_triage(instanceId="") + → Use topEvidence and recommendedSOPs from quick_triage + ↓ +4. For domain-specific deep dive: + → network_diagnostics(instanceId="", sections="iptables,cni,routes,dns,eni") + → search(instanceId="", query="") +``` + +### Key MCP Tools for Node Join Issues + +| MCP Tool | What It Provides | +|---|---| +| `collect` + `status` | Gathers all node logs via SSM (async, 1-3 min) | +| `quick_triage` | One-shot root cause with log evidence — maps to A4 SOP | +| `search` | Regex search across collected logs for specific patterns | +| `network_diagnostics` | Structured iptables/CNI/routes/DNS/ENI/IPAMD analysis | +| `correlate` | Cross-file timeline around a pivot event | +| `compare_nodes` | Diff a failed node against a healthy node | +| `cluster_health` | Check if other nodes are also failing (cluster-wide issue) | + +### Example: Full Complementary Investigation + +``` +# 1. API-side: This skill identifies NJ32 (cgroup mismatch suspected) +# Evidence: AL2023 AMI + custom launch template + immediate NotReady + +# 2. Confirm with MCP (if node is SSM-reachable): +collect(instanceId="i-0abc123", region="us-east-1") +status(executionId="") +search(instanceId="i-0abc123", query="cgroup.*driver|SystemdCgroup|cgroupfs") + +# 3. MCP confirms: kubelet using systemd, containerd using cgroupfs → mismatch +# 4. Report: root cause confirmed, recommend aligning to systemd +``` diff --git a/skills/eks-node-join-diagnostic/evals/benchmark.json b/skills/eks-node-join-diagnostic/evals/benchmark.json new file mode 100644 index 0000000..d9198ab --- /dev/null +++ b/skills/eks-node-join-diagnostic/evals/benchmark.json @@ -0,0 +1,693 @@ +{ + "skill_name": "eks-node-join-diagnostic", + "skill_path": "/Users/mridulch/Downloads/sample-devops-agent-tools/skills/eks-node-join-diagnostic", + "eval_count": 5, + "runs_per_eval": 1, + "metadata": { + "timestamp": "2026-08-12T02:20:34Z" + }, + "runs": [ + { + "eval_id": "1", + "run_index": 0, + "with_skill": { + "eval_id": "1", + "run_index": 0, + "assertion_results": [ + { + "text": "Agent calls eks:DescribeCluster and eks:DescribeNodegroup", + "passed": false, + "evidence": "Agent attempted describe-cluster but failed (cluster not found). It did not successfully call DescribeCluster or DescribeNodegroup - instead provided commands for user to run manually.", + "method": "llm", + "confidence": 0.7, + "uncertain": false + }, + { + "text": "Agent checks IAM role policies (NJ5)", + "passed": false, + "evidence": "Agent discusses IAM role policies and lists required ones (NJ5-NJ7), but did not actually check them - it provided CLI commands for the user to run rather than executing the checks itself.", + "method": "llm", + "confidence": 0.75, + "uncertain": false + }, + { + "text": "Agent checks authentication mapping \u2014 access entries or aws-auth (NJ6)", + "passed": false, + "evidence": "Agent discusses authentication mapping (access entries and aws-auth ConfigMap) with NJ6/NJ7 references, but did not actually check them - provided commands for user to execute.", + "method": "llm", + "confidence": 0.75, + "uncertain": false + }, + { + "text": "Agent checks security group rules (NJ9-NJ10)", + "passed": false, + "evidence": "Agent discusses security group requirements (NJ9-NJ10) and provides commands to check them, but did not actually execute the security group checks itself.", + "method": "llm", + "confidence": 0.75, + "uncertain": false + }, + { + "text": "Agent produces a diagnostic report with pass/fail per check", + "passed": false, + "evidence": "Output is structured as a prioritized list of potential causes with commands to verify, NOT a diagnostic report with pass/fail per check. No checks were actually completed to produce results.", + "method": "llm", + "confidence": 0.9, + "uncertain": false + }, + { + "text": "Agent provides specific remediation steps for the identified blocker", + "passed": false, + "evidence": "Agent provides general remediation steps for each potential blocker category (IAM, SGs, DNS, etc.), but did not identify a SPECIFIC blocker since it couldn't access the cluster. The remediation is generic guidance rather than targeted fixes for an identified issue.", + "method": "llm", + "confidence": 0.6, + "uncertain": false + } + ], + "pass_rate": 0.0, + "summary": "With skill: 0% assertions passed", + "execution_metrics": { + "tool_calls": 4, + "token_counts": { + "input_tokens": 5, + "output_tokens": 1938, + "cache_read_input_tokens": 94280, + "cache_creation_input_tokens": 24562 + } + }, + "timing": { + "elapsed_seconds": 41.81420216700644 + }, + "raw_output": "I'll run the EKS Node Join Diagnostic for your cluster. Let me gather the foundation data first.The AWS CLI can't find the cluster. Let me verify my AWS credentials and check what's available.I don't have access to a cluster named 'production' in this account \u2014 the available clusters in us-east-1 are 'Test' and 'eks-auto-mode-cluster'. The cluster may be in a different AWS account or profile.\n\nHowever, I can walk you through the systematic diagnostic based on the **58 known failure modes** for EKS node join failures. Given your symptoms (\"Create failed\", instances launched but never registered), here's what to check \u2014 starting with the most common causes:\n\n---\n\n## Most Likely Root Causes (in priority order)\n\n### 1. IAM / Auth Misconfiguration (NJ5\u2013NJ7) \u2014 *Most common cause*\n\nRun these:\n```bash\n# Get the node role from the node group\naws eks describe-nodegroup --cluster-name production --nodegroup-name prod-workers \\\n --region us-east-1 --query 'nodegroup.nodeRole'\n\n# Check the role has required policies\naws iam list-attached-role-policies --role-name \n```\n\n**Required policies:**\n- `AmazonEKSWorkerNodePolicy`\n- `AmazonEKS_CNI_Policy`\n- `AmazonEC2ContainerRegistryReadOnly`\n\nThen check auth mapping:\n```bash\n# For API auth mode (EKS 1.28+)\naws eks list-access-entries --cluster-name production --region us-east-1\n\n# For ConfigMap mode\nkubectl get cm aws-auth -n kube-system -o yaml\n```\n\nThe node role ARN must appear as an `EC2_LINUX` access entry **OR** in `aws-auth` mapRoles with groups `system:bootstrappers` and `system:nodes`.\n\n**Gotcha (NJ7):** If using aws-auth and the role has a path (e.g., `arn:aws:iam::123:role/path/MyRole`), you must strip the path to just `role/MyRole`.\n\n---\n\n### 2. Security Groups (NJ9\u2013NJ10)\n\nNodes must be able to reach the API server on TCP 443, and the cluster SG must allow inbound from nodes on 443 and 10250:\n```bash\n# Get cluster security group\naws eks describe-cluster --name production --region us-east-1 \\\n --query 'cluster.res" + }, + "without_skill": { + "eval_id": "1", + "run_index": 0, + "assertion_results": [ + { + "text": "Agent calls eks:DescribeCluster and eks:DescribeNodegroup", + "passed": false, + "evidence": "The output includes aws eks describe-cluster and describe-nodegroup commands, but they are presented as suggested commands for the user to run, not as actual API calls executed by the agent. No actual results are shown.", + "method": "llm", + "confidence": 0.8, + "uncertain": false + }, + { + "text": "Agent checks IAM role policies (NJ5)", + "passed": false, + "evidence": "The output lists required IAM policies and shows a command to verify them, but the agent did not actually execute the check or report results. It's guidance, not an executed check.", + "method": "llm", + "confidence": 0.75, + "uncertain": false + }, + { + "text": "Agent checks authentication mapping \u2014 access entries or aws-auth (NJ6)", + "passed": false, + "evidence": "The output mentions aws-auth ConfigMap mapping but only as a note to verify. No actual check was performed, no access entries were queried, and no results were reported.", + "method": "llm", + "confidence": 0.7, + "uncertain": false + }, + { + "text": "Agent checks security group rules (NJ9-NJ10)", + "passed": false, + "evidence": "Security group rules are discussed with a command to check them, but the agent did not actually execute the check or report findings from real security group rules.", + "method": "llm", + "confidence": 0.75, + "uncertain": false + }, + { + "text": "Agent produces a diagnostic report with pass/fail per check", + "passed": false, + "evidence": "There is no pass/fail diagnostic report. The output is a generic troubleshooting guide with categories and suggested commands, not a structured report with pass/fail results per check.", + "method": "llm", + "confidence": 0.95, + "uncertain": false + }, + { + "text": "Agent provides specific remediation steps for the identified blocker", + "passed": false, + "evidence": "The output provides general remediation guidance for various possible blockers but does not identify a specific blocker since no checks were actually executed. It ends asking if the user wants more help rather than pinpointing the issue.", + "method": "llm", + "confidence": 0.7, + "uncertain": false + } + ], + "pass_rate": 0.0, + "summary": "Without skill: 0% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 3, + "output_tokens": 1225, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 26588 + } + }, + "timing": { + "elapsed_seconds": 25.82929354099906 + }, + "raw_output": "# Troubleshooting: EKS Nodes Failed to Join Cluster\n\nWhen managed node group instances launch but never register with the control plane, the root cause is almost always one of these categories:\n\n## 1. Network/Connectivity (most common)\n\nThe kubelet on each node must reach the EKS API server endpoint. Check:\n\n- **Security groups**: The node security group must allow outbound HTTPS (443) to the cluster endpoint. The cluster security group must allow inbound 443 from the node SG.\n- **Subnet routing**: If using private subnets, you need either a NAT gateway (for public endpoint) or VPC endpoints for `eks`, `ecr.api`, `ecr.dkr`, `s3`, and `sts`.\n- **DNS resolution**: Nodes must resolve the cluster's API endpoint. If using private endpoint only, ensure the VPC DNS settings and the private hosted zone are accessible from the node subnets.\n\n```bash\n# Check cluster endpoint access configuration\naws eks describe-cluster --name production --region us-east-1 \\\n --query 'cluster.resourcesVpcConfig.{endpoint_public: endpointPublicAccess, endpoint_private: endpointPrivateAccess, subnet_ids: subnetIds, sg: clusterSecurityGroupId}'\n```\n\n## 2. IAM Role / Instance Profile\n\nThe node IAM role must have these AWS managed policies:\n- `AmazonEKSWorkerNodePolicy`\n- `AmazonEC2ContainerRegistryReadOnly`\n- `AmazonEKS_CNI_Policy`\n\nAnd the cluster's `aws-auth` ConfigMap must map the node role (managed node groups handle this automatically, but verify):\n\n```bash\n# Check the node group's role\naws eks describe-nodegroup --cluster-name production --nodegroup-name prod-workers \\\n --region us-east-1 --query 'nodegroup.nodeRole'\n\n# Verify policies attached\naws iam list-attached-role-policies --role-name \n```\n\n## 3. Cluster Security Group Rules\n\n```bash\n# Get cluster SG and node SG, then check rules\naws ec2 describe-security-group-rules --filters \"Name=group-id,Values=\" \\\n --region us-east-1 --query 'SecurityGroupRules[?IsEgress==`false`]'\n```\n\nNodes need inbound from c" + }, + "delta_pass_rate": 0.0 + }, + { + "eval_id": "2", + "run_index": 0, + "with_skill": { + "eval_id": "2", + "run_index": 0, + "assertion_results": [ + { + "text": "Agent detects private endpoint configuration (endpointPublicAccess=false)", + "passed": true, + "evidence": "Output identifies this as a 'private-cluster node join failure' and discusses private endpoint implications, but doesn't show evidence of programmatically detecting endpointPublicAccess=false - it appears to assume/know it's private from user context", + "method": "llm", + "confidence": 0.7, + "uncertain": false + }, + { + "text": "Agent checks VPC endpoints exist (NJ16) \u2014 ec2, ecr.api, ecr.dkr, sts, s3", + "passed": true, + "evidence": "Output explicitly lists all 5 required VPC endpoints (ec2, ecr.api, ecr.dkr, sts, s3) under NJ16 and asks user to confirm which are configured, though it doesn't show evidence of actually querying AWS to verify their existence", + "method": "llm", + "confidence": 0.85, + "uncertain": false + }, + { + "text": "Agent checks bootstrap user-data has --apiserver-endpoint, --b64-cluster-ca, --dns-cluster-ip (NJ18)", + "passed": true, + "evidence": "Output explicitly mentions NJ18 and lists all three required bootstrap args: --apiserver-endpoint, --b64-cluster-ca, --dns-cluster-ip, and asks user whether user-data passes them explicitly", + "method": "llm", + "confidence": 0.9, + "uncertain": false + }, + { + "text": "Agent checks VPC DNS support and hostnames enabled (NJ11)", + "passed": false, + "evidence": "Output makes no mention of VPC DNS support, DNS hostnames, or NJ11 anywhere in the response", + "method": "llm", + "confidence": 0.95, + "uncertain": false + }, + { + "text": "Agent provides the specific missing VPC endpoint or bootstrap arg", + "passed": false, + "evidence": "Output lists all potentially missing VPC endpoints and bootstrap args as candidates ranked by likelihood, but does not definitively identify THE specific missing item - it asks follow-up questions to narrow down rather than providing a conclusive answer", + "method": "llm", + "confidence": 0.6, + "uncertain": false + } + ], + "pass_rate": 0.6, + "summary": "With skill: 60% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 3, + "output_tokens": 965, + "cache_read_input_tokens": 15466, + "cache_creation_input_tokens": 23673 + } + }, + "timing": { + "elapsed_seconds": 23.53329650001251 + }, + "raw_output": "This is a classic private-cluster node join failure. Since instances are running but never register, the problem is almost certainly in one of these areas (ordered by likelihood for your exact scenario):\n\n## Most Likely Causes\n\n### 1. Missing bootstrap arguments in user-data (NJ18 \u2014 Critical)\n\nIn a private cluster, nodes cannot call `aws eks describe-cluster` to discover the endpoint/CA. Your launch template user-data **must** pass these directly:\n\n```bash\n/etc/eks/bootstrap.sh private-cluster \\\n --apiserver-endpoint https:// \\\n --b64-cluster-ca \\\n --dns-cluster-ip 10.100.0.10\n```\n\nIf your user-data relies on the bootstrap script's auto-discovery, it will silently fail with no network path to the EKS API.\n\n### 2. Missing VPC endpoints (NJ16 \u2014 Critical)\n\nPrivate clusters require these VPC endpoints with security groups allowing inbound 443 from node subnets:\n\n- `com.amazonaws.us-west-2.ec2` (interface)\n- `com.amazonaws.us-west-2.ecr.api` (interface)\n- `com.amazonaws.us-west-2.ecr.dkr` (interface)\n- `com.amazonaws.us-west-2.sts` (interface)\n- `com.amazonaws.us-west-2.s3` (gateway)\n\nWithout these, nodes can't pull images, authenticate, or reach the API server.\n\n### 3. Node role not in auth config (NJ6 \u2014 Critical)\n\nFor self-managed nodes, the IAM role must be explicitly mapped. Check:\n```bash\nkubectl get cm aws-auth -n kube-system -o yaml\n```\nIt must contain your node role (NOT instance profile ARN) in `mapRoles` with groups `system:bootstrappers` and `system:nodes`.\n\n### 4. Security group rules (NJ9-NJ10 \u2014 Critical)\n\nNode SG must allow outbound TCP 443 to the cluster SG, and the cluster SG must allow inbound TCP 443 + 10250 from the node SG.\n\n---\n\n## To narrow it down, can you confirm:\n\n1. Does your user-data pass `--apiserver-endpoint` and `--b64-cluster-ca` explicitly, or does it rely on auto-discovery?\n2. Which VPC endpoints do you have configured?\n3. Is the node IAM role mapped in `aws-auth` (or as an EKS Access Entry)?" + }, + "without_skill": { + "eval_id": "2", + "run_index": 0, + "assertion_results": [ + { + "text": "Agent detects private endpoint configuration (endpointPublicAccess=false)", + "passed": true, + "evidence": "Output addresses 'fully private EKS clusters' and states 'With no public endpoint, nodes can't reach the internet' \u2014 it clearly recognizes the private endpoint configuration, though it doesn't explicitly say 'endpointPublicAccess=false' as a detected setting.", + "method": "llm", + "confidence": 0.82, + "uncertain": false + }, + { + "text": "Agent checks VPC endpoints exist (NJ16) \u2014 ec2, ecr.api, ecr.dkr, sts, s3", + "passed": true, + "evidence": "Section 1 explicitly lists all required VPC endpoints in a table: ec2, ecr.api, ecr.dkr, sts, and s3 (Gateway), plus the eks endpoint.", + "method": "llm", + "confidence": 0.97, + "uncertain": false + }, + { + "text": "Agent checks bootstrap user-data has --apiserver-endpoint, --b64-cluster-ca, --dns-cluster-ip (NJ18)", + "passed": true, + "evidence": "Section 2 explicitly shows the bootstrap script with --apiserver-endpoint, --b64-cluster-ca, and --dns-cluster-ip, and warns about omitting them in a private cluster.", + "method": "llm", + "confidence": 0.98, + "uncertain": false + }, + { + "text": "Agent checks VPC DNS support and hostnames enabled (NJ11)", + "passed": true, + "evidence": "Section 5 explicitly states 'Confirm that enableDnsSupport and enableDnsHostnames are both true on the VPC' and explains why they're needed.", + "method": "llm", + "confidence": 0.97, + "uncertain": false + }, + { + "text": "Agent provides the specific missing VPC endpoint or bootstrap arg", + "passed": true, + "evidence": "The conclusion specifically calls out 'missing --apiserver-endpoint and --b64-cluster-ca in the bootstrap user data, combined with no eks VPC endpoint' as the most common root cause, providing specific missing items.", + "method": "llm", + "confidence": 0.92, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "Without skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 3, + "output_tokens": 1477, + "cache_read_input_tokens": 15461, + "cache_creation_input_tokens": 11130 + } + }, + "timing": { + "elapsed_seconds": 30.323609667015262 + }, + "raw_output": "This is a classic issue with fully private EKS clusters + custom launch templates. The most likely causes, in order of probability:\n\n## 1. Missing VPC Endpoints\n\nWith no public endpoint, nodes can't reach the internet. They need these VPC endpoints in your VPC:\n\n| Endpoint | Why |\n|----------|-----|\n| `com.amazonaws.us-west-2.eks` | API server communication |\n| `com.amazonaws.us-west-2.ecr.api` | Pull kubelet/kube-proxy images |\n| `com.amazonaws.us-west-2.ecr.dkr` | ECR Docker registry |\n| `com.amazonaws.us-west-2.s3` (Gateway) | ECR image layers |\n| `com.amazonaws.us-west-2.sts` | IAM authentication |\n| `com.amazonaws.us-west-2.ec2` | Node registration |\n\nThe security groups on these endpoints must allow inbound HTTPS (443) from your node security group.\n\n## 2. Broken Bootstrap User Data\n\nWith a custom launch template, **you** are responsible for the bootstrap script. It must pass the correct API endpoint and CA cert:\n\n```bash\n#!/bin/bash\n/etc/eks/bootstrap.sh 'private-cluster' \\\n --apiserver-endpoint 'https://XXXXX.gr7.us-west-2.eks.amazonaws.com' \\\n --b64-cluster-ca 'LS0tLS1CRUdJTi...' \\\n --dns-cluster-ip '10.100.0.10'\n```\n\nIf you omit `--apiserver-endpoint` and `--b64-cluster-ca`, the bootstrap script tries to call `eks:DescribeCluster` over the public internet \u2014 which fails in a private cluster with no NAT.\n\n## 3. Node IAM Role Not Mapped\n\nThe node instance profile role must be mapped in `aws-auth` ConfigMap (or EKS access entries):\n\n```yaml\n# kubectl edit configmap aws-auth -n kube-system\nmapRoles:\n - rolearn: arn:aws:iam::ACCOUNT:role/YourNodeRole\n username: system:node:{{EC2PrivateDNSName}}\n groups:\n - system:bootstrappers\n - system:nodes\n```\n\nThe role also needs these managed policies attached:\n- `AmazonEKSWorkerNodePolicy`\n- `AmazonEC2ContainerRegistryReadOnly`\n- `AmazonEKS_CNI_Policy`\n\n## 4. Security Group Misconfiguration\n\nThe cluster control plane security group must allow:\n- **Inbound 443** from the node SG (so nodes can register)\n-" + }, + "delta_pass_rate": -0.4 + }, + { + "eval_id": "3", + "run_index": 0, + "with_skill": { + "eval_id": "3", + "run_index": 0, + "assertion_results": [ + { + "text": "Agent checks AMI compatibility (NJ22) \u2014 AL2 not supported on 1.33+", + "passed": true, + "evidence": "Output explicitly states 'Root cause: AL2 AMI is not supported on EKS 1.33+ (check NJ22)' and elaborates that AL2 support was dropped starting with EKS 1.33.", + "method": "llm", + "confidence": 0.97, + "uncertain": false + }, + { + "text": "Agent identifies this as a Critical blocker", + "passed": false, + "evidence": "Output identifies it as a 'Root cause' but never explicitly labels it as a 'Critical blocker' or uses severity terminology like 'critical'.", + "method": "llm", + "confidence": 0.82, + "uncertain": false + }, + { + "text": "Agent recommends AL2023 or Bottlerocket as replacement", + "passed": true, + "evidence": "Output explicitly recommends both 'Amazon Linux 2023' (AL2023_x86_64_STANDARD or AL2023_ARM_64_STANDARD) and 'Bottlerocket' (BOTTLEROCKET_x86_64 or BOTTLEROCKET_ARM_64) as replacement options.", + "method": "llm", + "confidence": 0.99, + "uncertain": false + }, + { + "text": "Agent mentions the bootstrap format change (nodeadm for AL2023)", + "passed": true, + "evidence": "Output explicitly states 'AL2023 uses nodeadm with a MIME multipart user-data format' and contrasts it with AL2's bootstrap.sh mechanism.", + "method": "llm", + "confidence": 0.99, + "uncertain": false + } + ], + "pass_rate": 0.75, + "summary": "With skill: 75% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 3, + "output_tokens": 657, + "cache_read_input_tokens": 15466, + "cache_creation_input_tokens": 23654 + } + }, + "timing": { + "elapsed_seconds": 18.59959304201766 + }, + "raw_output": "**Root cause: AL2 AMI is not supported on EKS 1.33+** (check NJ22).\n\nAmazon Linux 2 support was dropped starting with EKS 1.33. Your old 1.32 node group still works because it was created before the upgrade and uses an AL2 AMI pinned to the 1.32 kubelet, but new node groups targeting 1.33 cannot use AL2.\n\n**Fix:** Recreate the node group using one of:\n\n- **Amazon Linux 2023** (`AL2023_x86_64_STANDARD` or `AL2023_ARM_64_STANDARD`)\n- **Bottlerocket** (`BOTTLEROCKET_x86_64` or `BOTTLEROCKET_ARM_64`)\n\n**Important migration note for AL2023:** The bootstrap mechanism changed. AL2 used `bootstrap.sh`, but AL2023 uses `nodeadm` with a MIME multipart user-data format. If you have a custom launch template with user-data, you'll need to convert it:\n\n```\nMIME-Version: 1.0\nContent-Type: multipart/mixed; boundary=\"//\"\n\n--//\nContent-Type: application/node.eks.aws\n\n---\napiVersion: node.eks.aws/v1alpha1\nkind: NodeConfig\nspec:\n cluster:\n name: \n apiServerEndpoint: https://\n certificateAuthority: \n cidr: \n--//--\n```\n\nIf you're using a managed node group without a custom launch template, just changing the `amiType` parameter to an AL2023 variant should be sufficient \u2014 EKS handles the bootstrap automatically." + }, + "without_skill": { + "eval_id": "3", + "run_index": 0, + "assertion_results": [ + { + "text": "Agent checks AMI compatibility (NJ22) \u2014 AL2 not supported on 1.33+", + "passed": true, + "evidence": "Output clearly identifies AL2 not supported on 1.33+ ('AL2 AMI support was dropped in EKS 1.33'), but does not explicitly reference 'NJ22' as a check identifier", + "method": "llm", + "confidence": 0.75, + "uncertain": false + }, + { + "text": "Agent identifies this as a Critical blocker", + "passed": false, + "evidence": "Output implies it's a blocker ('new nodes can't launch') but never explicitly labels it as 'Critical' severity or uses the word 'blocker' as a classification", + "method": "llm", + "confidence": 0.85, + "uncertain": false + }, + { + "text": "Agent recommends AL2023 or Bottlerocket as replacement", + "passed": true, + "evidence": "Output explicitly recommends AL2023 as the replacement: 'Migrate your node groups to AL2023 (Amazon Linux 2023), which is the supported replacement'", + "method": "llm", + "confidence": 0.95, + "uncertain": false + }, + { + "text": "Agent mentions the bootstrap format change (nodeadm for AL2023)", + "passed": true, + "evidence": "Output explicitly mentions 'nodeadm for bootstrapping instead of the old /etc/eks/bootstrap.sh script' and discusses the nodeadm YAML configuration format change", + "method": "llm", + "confidence": 0.99, + "uncertain": false + } + ], + "pass_rate": 0.75, + "summary": "Without skill: 75% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 3, + "output_tokens": 399, + "cache_read_input_tokens": 15461, + "cache_creation_input_tokens": 11119 + } + }, + "timing": { + "elapsed_seconds": 11.189908166998066 + }, + "raw_output": "**AL2 AMI support was dropped in EKS 1.33.** Starting with Kubernetes 1.33, AWS no longer publishes Amazon Linux 2 EKS-optimized AMIs. Your old 1.32 node group still works because existing nodes aren't affected, but new nodes can't launch because there's no AL2 AMI available for 1.33.\n\n**Fix:** Migrate your node groups to **AL2023** (Amazon Linux 2023), which is the supported replacement. Key differences to watch for:\n\n- AL2023 uses `systemd`-based networking and `nodeadm` for bootstrapping instead of the old `/etc/eks/bootstrap.sh` script\n- If you have custom `userData` scripts that call `bootstrap.sh`, they need to be rewritten for the `nodeadm` YAML configuration format\n- Security group and IAM role configs carry over unchanged\n\nIf you're using managed node groups with a launch template that specifies an AMI ID, update it to the AL2023 AMI. If you're using a custom AMI built on AL2, you'll need to rebase it on AL2023." + }, + "delta_pass_rate": 0.0 + }, + { + "eval_id": "4", + "run_index": 0, + "with_skill": { + "eval_id": "4", + "run_index": 0, + "assertion_results": [ + { + "text": "Agent checks cluster state and Karpenter detection", + "passed": false, + "evidence": "Agent attempted to check cluster state but failed to connect to the correct cluster. It only found an unrelated 'eks-workshop' cluster and could not verify Karpenter detection on 'analytics-prod'.", + "method": "llm", + "confidence": 0.85, + "uncertain": false + }, + { + "text": "Agent evaluates EC2NodeClass subnets match cluster VPC (NJ25)", + "passed": false, + "evidence": "Agent never evaluated EC2NodeClass subnets. It only mentioned NJ25 as part of a future checklist it would run once access is obtained.", + "method": "llm", + "confidence": 0.95, + "uncertain": false + }, + { + "text": "Agent checks EC2NodeClass security groups include cluster SG (NJ26)", + "passed": false, + "evidence": "Agent never checked EC2NodeClass security groups. It only referenced this as part of a planned diagnostic (NJ25-NJ27) contingent on gaining cluster access.", + "method": "llm", + "confidence": 0.95, + "uncertain": false + }, + { + "text": "Agent checks AMI compatibility in EC2NodeClass (NJ27)", + "passed": false, + "evidence": "Agent never checked AMI compatibility. It mentioned NJ27 in passing as part of a future checklist but did not execute the check.", + "method": "llm", + "confidence": 0.95, + "uncertain": false + }, + { + "text": "Agent checks if VPC CNI can schedule on the tainted node (NJ28)", + "passed": false, + "evidence": "Agent never checked VPC CNI scheduling on tainted nodes. NJ28 was not even explicitly mentioned in the output, and no diagnostic checks were actually performed.", + "method": "llm", + "confidence": 0.95, + "uncertain": false + } + ], + "pass_rate": 0.0, + "summary": "With skill: 0% assertions passed", + "execution_metrics": { + "tool_calls": 5, + "token_counts": { + "input_tokens": 5, + "output_tokens": 1163, + "cache_read_input_tokens": 94598, + "cache_creation_input_tokens": 25031 + } + }, + "timing": { + "elapsed_seconds": 31.84957887500059 + }, + "raw_output": "I'll start diagnosing the Karpenter node join issue on your `analytics-prod` cluster in `eu-west-1`. Let me gather foundation data first.The CLI isn't connected to your `analytics-prod` cluster \u2014 it's pointing at a different cluster (running v1.36.1). I need access to the correct cluster to proceed with the diagnostic.\n\nBefore we troubleshoot the kubeconfig, a few questions:\n\n1. **When did this start?** Was it working before and broke recently, or is this the first time Karpenter is launching nodes on this cluster?\n2. **Do you have a specific instance ID** of one of the stuck NotReady nodes?\n3. **Can you confirm your kubeconfig context** for `analytics-prod`? (e.g., run `kubectl config get-contexts | grep analytics`)\n\nIn the meantime, I can run API-side checks that don't require kubectl if you can confirm the AWS account ID. Let me check what's available:The current AWS credentials (account `421198816590`) only show an `eks-workshop` cluster in `eu-west-1` \u2014 there's no `analytics-prod` cluster visible from this account/region.\n\nHere's what we need to resolve:\n\n1. **Wrong account** \u2014 Is `analytics-prod` in a different AWS account? If so, you'll need to assume a role into that account or switch credentials.\n2. **Wrong region** \u2014 Could the cluster be in a different region (e.g., `eu-west-2`)?\n3. **Access permissions** \u2014 Does the current IAM user have `eks:DescribeCluster` permission on that cluster?\n\nOnce I can reach the cluster via the AWS API, I'll run the full NJ1\u2013NJ61 diagnostic checklist covering IAM/auth, security groups, VPC/DNS, bootstrap config, AMI compatibility, and Karpenter-specific checks (NJ25\u2013NJ27). \n\nCould you provide the correct account/credentials or kubeconfig context for `analytics-prod`?" + }, + "without_skill": { + "eval_id": "4", + "run_index": 0, + "assertion_results": [ + { + "text": "Agent checks cluster state and Karpenter detection", + "passed": true, + "evidence": "Output checks cluster state with kubectl get nodes/describe node and mentions Karpenter-launched nodes, but doesn't explicitly detect Karpenter resources (NodePools, NodeClaims) or check Karpenter-specific CRDs", + "method": "llm", + "confidence": 0.5, + "uncertain": false + }, + { + "text": "Agent evaluates EC2NodeClass subnets match cluster VPC (NJ25)", + "passed": false, + "evidence": "Output mentions subnets running out of IPs (#5) but never evaluates EC2NodeClass subnet configuration against the cluster VPC - no kubectl get ec2nodeclass or similar check", + "method": "llm", + "confidence": 0.9, + "uncertain": false + }, + { + "text": "Agent checks EC2NodeClass security groups include cluster SG (NJ26)", + "passed": false, + "evidence": "Output mentions security group requirements (ports 443, 10250) but does not specifically check the EC2NodeClass resource's securityGroupSelector to verify it includes the cluster security group", + "method": "llm", + "confidence": 0.75, + "uncertain": false + }, + { + "text": "Agent checks AMI compatibility in EC2NodeClass (NJ27)", + "passed": false, + "evidence": "Output never mentions AMI compatibility or checking the EC2NodeClass AMI configuration - no reference to amiFamily, amiSelectorTerms, or AMI version matching", + "method": "llm", + "confidence": 0.95, + "uncertain": false + }, + { + "text": "Agent checks if VPC CNI can schedule on the tainted node (NJ28)", + "passed": false, + "evidence": "Output checks if aws-node DaemonSet pods are running (#1) but does not specifically investigate whether VPC CNI tolerations allow scheduling on tainted Karpenter nodes", + "method": "llm", + "confidence": 0.7, + "uncertain": false + } + ], + "pass_rate": 0.2, + "summary": "Without skill: 20% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 3, + "output_tokens": 841, + "cache_read_input_tokens": 15461, + "cache_creation_input_tokens": 11132 + } + }, + "timing": { + "elapsed_seconds": 19.944133916986175 + }, + "raw_output": "This is a common issue with several likely root causes. Here's how to diagnose, ordered by frequency:\n\n## Most Likely Causes\n\n**1. VPC CNI / aws-node DaemonSet not running on new nodes**\n\n```bash\nkubectl get ds -n kube-system aws-node\nkubectl get pods -n kube-system -l k8s-app=aws-node --field-selector status.phase!=Running\n```\n\nIf `aws-node` pods are crashing on new nodes, the node can't get networking and stays NotReady.\n\n**2. Node can't reach the API server (security group or endpoint issue)**\n\nCheck if the Karpenter-launched nodes' security group allows:\n- Outbound HTTPS (443) to the EKS API endpoint\n- Inbound from the control plane security group on ports 443 and 10250\n\n```bash\n# Check the node's condition messages\nkubectl get nodes -o wide\nkubectl describe node | grep -A5 Conditions\n```\n\n**3. Incorrect or missing instance profile / IAM role**\n\nThe node role needs these managed policies:\n- `AmazonEKSWorkerNodePolicy`\n- `AmazonEKS_CNI_Policy`\n- `AmazonEC2ContainerRegistryReadOnly`\n\nAlso verify the role is in the `aws-auth` ConfigMap (or uses EKS access entries on 1.30):\n\n```bash\nkubectl get cm -n kube-system aws-auth -o yaml\n```\n\n**4. UserData / bootstrap not completing**\n\nSSH or SSM into a stuck instance and check:\n\n```bash\n# On the EC2 instance\njournalctl -u kubelet --no-pager | tail -50\ncat /var/log/cloud-init-output.log | tail -100\n```\n\nLook for certificate errors, API server unreachable, or containerd failures.\n\n**5. Subnet has no remaining IPs or lacks route to NAT gateway**\n\nIf the subnet is exhausted or has no outbound route, kubelet can't register.\n\n## Quick Triage\n\n```bash\n# Get a stuck node name\nNODE=$(kubectl get nodes --field-selector spec.unschedulable!=true -o jsonpath='{.items[?(@.status.conditions[?(@.type==\"Ready\")].status==\"False\")].metadata.name}' | awk '{print $1}')\n\n# See why it's NotReady\nkubectl describe node $NODE | grep -A10 \"Conditions:\"\n\n# Check kubelet logs via SSM (if SSM agent is running)\nINSTANCE_ID=$(kubectl get " + }, + "delta_pass_rate": -0.2 + }, + { + "eval_id": "5", + "run_index": 0, + "with_skill": { + "eval_id": "5", + "run_index": 0, + "assertion_results": [ + { + "text": "Agent identifies the role path issue (NJ7)", + "passed": true, + "evidence": "Output explicitly states 'This is a known issue \u2014 NJ7 in the diagnostic checklist' in the opening line.", + "method": "llm", + "confidence": 0.97, + "uncertain": false + }, + { + "text": "Agent explains that aws-auth requires the path to be stripped", + "passed": true, + "evidence": "Output clearly states 'aws-auth requires the path to be stripped' and explains that 'The /teams/platform/ path segment must be removed'.", + "method": "llm", + "confidence": 0.99, + "uncertain": false + }, + { + "text": "Agent provides the corrected ARN: arn:aws:iam::123456789012:role/eks-node-role", + "passed": true, + "evidence": "Output provides the corrected ARN 'arn:aws:iam::123456789012:role/eks-node-role' both in the explanation and in the YAML fix example.", + "method": "llm", + "confidence": 0.99, + "uncertain": false + }, + { + "text": "Agent recommends migrating to EKS Access Entries as the long-term fix", + "passed": true, + "evidence": "Output has a 'Better Long-Term Fix' section explicitly recommending migration to 'EKS Access Entries (API authentication mode)' with a command example.", + "method": "llm", + "confidence": 0.99, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 3, + "output_tokens": 525, + "cache_read_input_tokens": 15466, + "cache_creation_input_tokens": 23678 + } + }, + "timing": { + "elapsed_seconds": 12.902160792000359 + }, + "raw_output": "This is a known issue \u2014 **NJ7** in the diagnostic checklist. The `aws-auth` ConfigMap does **not** support IAM role ARNs with paths.\n\n## The Problem\n\nYour role ARN is:\n```\narn:aws:iam::123456789012:role/teams/platform/eks-node-role\n```\n\nBut `aws-auth` requires the path to be **stripped**. It must be entered as:\n```\narn:aws:iam::123456789012:role/eks-node-role\n```\n\nThe `/teams/platform/` path segment must be removed \u2014 only `role/` is recognized by the Kubernetes AWS IAM authenticator when parsing the ConfigMap.\n\n## Fix\n\nUpdate your `aws-auth` ConfigMap `mapRoles` entry:\n\n```yaml\nmapRoles: |\n - rolearn: arn:aws:iam::123456789012:role/eks-node-role\n username: system:node:{{EC2PrivateDNSName}}\n groups:\n - system:bootstrappers\n - system:nodes\n```\n\n## Better Long-Term Fix\n\nMigrate to **EKS Access Entries** (API authentication mode), which handle paths correctly:\n\n```bash\naws eks create-access-entry \\\n --cluster-name dev-eks \\\n --principal-arn arn:aws:iam::123456789012:role/teams/platform/eks-node-role \\\n --type EC2_LINUX\n```\n\nWith Access Entries you can use the full ARN including the path, and you don't need to manage the `aws-auth` ConfigMap at all.\n\nAfter applying either fix, existing nodes will need to be recycled (terminated and replaced) to re-attempt the join with valid authentication." + }, + "without_skill": { + "eval_id": "5", + "run_index": 0, + "assertion_results": [ + { + "text": "Agent identifies the role path issue (NJ7)", + "passed": true, + "evidence": "Output explicitly states 'IAM role paths are not supported in the aws-auth ConfigMap' and identifies the /teams/platform/ path as the issue.", + "method": "llm", + "confidence": 0.97, + "uncertain": false + }, + { + "text": "Agent explains that aws-auth requires the path to be stripped", + "passed": true, + "evidence": "Output explains 'The IAM authenticator strips the path when matching, so the ARN in the ConfigMap must also have the path removed.'", + "method": "llm", + "confidence": 0.98, + "uncertain": false + }, + { + "text": "Agent provides the corrected ARN: arn:aws:iam::123456789012:role/eks-node-role", + "passed": true, + "evidence": "Output explicitly shows the corrected ARN 'rolearn: arn:aws:iam::123456789012:role/eks-node-role' in the corrected ConfigMap snippet.", + "method": "llm", + "confidence": 0.99, + "uncertain": false + }, + { + "text": "Agent recommends migrating to EKS Access Entries as the long-term fix", + "passed": true, + "evidence": "Output has an 'Alternative (recommended for new clusters)' section recommending EKS access entries, stating they 'handle pathed role ARNs correctly and are the preferred approach going forward.'", + "method": "llm", + "confidence": 0.95, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "Without skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 3, + "output_tokens": 637, + "cache_read_input_tokens": 15461, + "cache_creation_input_tokens": 11141 + } + }, + "timing": { + "elapsed_seconds": 13.345896665996406 + }, + "raw_output": "This is a known EKS/aws-auth ConfigMap limitation: **IAM role paths are not supported** in the `aws-auth` ConfigMap. The IAM authenticator strips the path when matching, so the ARN in the ConfigMap must also have the path removed.\n\n**Your ConfigMap currently has:**\n```yaml\n- rolearn: arn:aws:iam::123456789012:role/teams/platform/eks-node-role\n username: system:node:{{EC2PrivateDNSName}}\n groups:\n - system:bootstrappers\n - system:nodes\n```\n\n**Change it to (path stripped):**\n```yaml\n- rolearn: arn:aws:iam::123456789012:role/eks-node-role\n username: system:node:{{EC2PrivateDNSName}}\n groups:\n - system:bootstrappers\n - system:nodes\n```\n\nThe actual IAM role can keep its path (`/teams/platform/`) \u2014 you only need to remove the path from the ARN entry in the ConfigMap.\n\n**Alternative (recommended for new clusters):** If you're on EKS with access entries (available since late 2023), you can use `aws eks create-access-entry` instead of the ConfigMap. Access entries handle pathed role ARNs correctly and are the preferred approach going forward:\n\n```bash\naws eks create-access-entry \\\n --cluster-name dev-eks \\\n --principal-arn arn:aws:iam::123456789012:role/teams/platform/eks-node-role \\\n --type EC2_LINUX \\\n --region us-east-2\n```" + }, + "delta_pass_rate": 0.0 + } + ], + "run_summary": { + "with_skill": { + "mean_pass_rate": 0.47, + "stddev_pass_rate": 0.4522, + "mean_tokens": 1049.6, + "mean_input_tokens": 3.8, + "mean_output_tokens": 1049.6, + "mean_total_tokens": 1053.4, + "mean_tool_calls": 1.8 + }, + "without_skill": { + "mean_pass_rate": 0.59, + "stddev_pass_rate": 0.4642, + "mean_tokens": 915.8, + "mean_input_tokens": 3.0, + "mean_output_tokens": 915.8, + "mean_total_tokens": 918.8, + "mean_tool_calls": 0.0 + }, + "delta": { + "pass_rate": -0.12, + "tokens": 133.8, + "total_tokens": 134.6, + "input_tokens": 0.8, + "tool_calls": 1.8 + }, + "cost_efficiency": { + "quality_delta": -0.12, + "cost_delta_pct": 14.6, + "classification": "REJECT", + "emoji": "\ud83d\udd34", + "description": "Skill significantly degrades quality" + }, + "estimated_cost": { + "with_skill_per_run": { + "input_cost": 1.1e-05, + "output_cost": 0.015744, + "total_cost": 0.015755, + "model": "sonnet", + "currency": "USD" + }, + "without_skill_per_run": { + "input_cost": 9e-06, + "output_cost": 0.013737, + "total_cost": 0.013746, + "model": "sonnet", + "currency": "USD" + }, + "per_eval_pair": 0.029501, + "total_runs": 5, + "total_cost": 0.1475, + "model": "sonnet", + "currency": "USD" + } + }, + "scores": { + "outcome": 0.47, + "process": 1.0, + "style": 0.47, + "efficiency": 0.8199, + "overall": 0.69 + }, + "passed": false +} \ No newline at end of file diff --git a/skills/eks-node-join-diagnostic/evals/eval_queries.json b/skills/eks-node-join-diagnostic/evals/eval_queries.json new file mode 100644 index 0000000..5d3e4ae --- /dev/null +++ b/skills/eks-node-join-diagnostic/evals/eval_queries.json @@ -0,0 +1,82 @@ +[ + { + "query": "My EKS managed node group in us-east-1 shows 'Instances failed to join the kubernetes cluster'. The node group is in Create failed state. Can you diagnose what's wrong?", + "should_trigger": true + }, + { + "query": "I launched self-managed nodes for my EKS cluster but they don't appear in kubectl get nodes. The EC2 instances are running. What's preventing them from registering?", + "should_trigger": true + }, + { + "query": "Getting TLS handshake timeout when my worker nodes try to connect to the EKS API server. Cluster is private endpoint only.", + "should_trigger": true + }, + { + "query": "Karpenter is provisioning nodes on my EKS cluster but they stay NotReady. The instances launch fine but never register with the cluster.", + "should_trigger": true + }, + { + "query": "After upgrading to EKS 1.33, my new node group with AL2 AMI won't start. Getting NodeCreationFailure.", + "should_trigger": true + }, + { + "query": "kubelet logs on my EKS node show 'Unauthorized' and 'Unable to register node'. I've added the role to aws-auth but it's not working.", + "should_trigger": true + }, + { + "query": "My EKS node group shows health issue Ec2SecurityGroupNotFound. The node group is in Degraded state.", + "should_trigger": true + }, + { + "query": "InvalidClientTokenId error when my nodes try to join the EKS cluster in ap-southeast-1. What's wrong?", + "should_trigger": true + }, + { + "query": "Nodes show 'Container runtime network not ready: cni config uninitialized' and never become Ready.", + "should_trigger": true + }, + { + "query": "I need to set up a fully private EKS cluster with nodes in isolated subnets. What VPC endpoints do I need to make sure nodes can join?", + "should_trigger": true + }, + { + "query": "My EKS cluster API server is slow. kubectl get pods takes 5 seconds. I'm seeing 429 throttling errors.", + "should_trigger": false + }, + { + "query": "Can you review the security posture of my EKS cluster? Check RBAC, pod security, and network policies.", + "should_trigger": false + }, + { + "query": "I want to upgrade my EKS cluster from 1.29 to 1.30. Can you check if we're ready?", + "should_trigger": false + }, + { + "query": "My pods keep getting OOMKilled on EKS. Memory utilization is high. How do I right-size?", + "should_trigger": false + }, + { + "query": "Help me set up Karpenter from scratch on my new EKS cluster with Spot and On-Demand instances.", + "should_trigger": false + }, + { + "query": "I need to create an EKS cluster using Terraform with managed node groups and VPC CNI.", + "should_trigger": false + }, + { + "query": "My ECS tasks are failing health checks behind an ALB. The targets keep going unhealthy.", + "should_trigger": false + }, + { + "query": "Can you generate a cost report for my EKS cluster? I want to know where we're spending the most.", + "should_trigger": false + }, + { + "query": "My EKS nodes keep getting terminated by Karpenter consolidation. How do I prevent that?", + "should_trigger": false + }, + { + "query": "I'm trying to expose my Kubernetes service through an NLB but it's not getting an external IP.", + "should_trigger": false + } +] diff --git a/skills/eks-node-join-diagnostic/evals/evals.json b/skills/eks-node-join-diagnostic/evals/evals.json new file mode 100644 index 0000000..e743002 --- /dev/null +++ b/skills/eks-node-join-diagnostic/evals/evals.json @@ -0,0 +1,66 @@ +[ + { + "id": "1", + "prompt": "My managed node group 'prod-workers' in EKS cluster 'production' in us-east-1 shows 'Create failed' with the error 'Instances failed to join the kubernetes cluster'. The node group was created 30 minutes ago and the instances launched but never registered. Help me figure out why.", + "expected_output": "Agent runs NJ1-NJ28 checks, identifies the specific blocker (likely IAM auth mapping or security group issue), and provides the exact fix", + "files": [], + "assertions": [ + "Agent calls eks:DescribeCluster and eks:DescribeNodegroup", + "Agent checks IAM role policies (NJ5)", + "Agent checks authentication mapping — access entries or aws-auth (NJ6)", + "Agent checks security group rules (NJ9-NJ10)", + "Agent produces a diagnostic report with pass/fail per check", + "Agent provides specific remediation steps for the identified blocker" + ] + }, + { + "id": "2", + "prompt": "I have a fully private EKS cluster (no public endpoint) in us-west-2 called 'private-cluster'. I'm trying to add a self-managed node group using a custom launch template but the nodes never appear in kubectl get nodes. The instances are running in EC2 but not registering. What's wrong?", + "expected_output": "Agent diagnoses private cluster prerequisites — VPC endpoints, bootstrap user-data args, DNS settings, and security groups", + "files": [], + "assertions": [ + "Agent detects private endpoint configuration (endpointPublicAccess=false)", + "Agent checks VPC endpoints exist (NJ16) — ec2, ecr.api, ecr.dkr, sts, s3", + "Agent checks bootstrap user-data has --apiserver-endpoint, --b64-cluster-ca, --dns-cluster-ip (NJ18)", + "Agent checks VPC DNS support and hostnames enabled (NJ11)", + "Agent provides the specific missing VPC endpoint or bootstrap arg" + ] + }, + { + "id": "3", + "prompt": "After upgrading our EKS cluster from 1.32 to 1.33, the new node group we created won't start. The old node group on 1.32 still works fine. We're using AL2 AMI. What happened?", + "expected_output": "Agent identifies that AL2 AMI is not available for EKS 1.33+ and recommends migrating to AL2023 or Bottlerocket", + "files": [], + "assertions": [ + "Agent checks AMI compatibility (NJ22) — AL2 not supported on 1.33+", + "Agent identifies this as a Critical blocker", + "Agent recommends AL2023 or Bottlerocket as replacement", + "Agent mentions the bootstrap format change (nodeadm for AL2023)" + ] + }, + { + "id": "4", + "prompt": "We use Karpenter v0.37 on our EKS 1.30 cluster. Karpenter is launching instances but they stay in NotReady state. The EC2 instances are running but never appear as Ready nodes. Cluster is 'analytics-prod' in eu-west-1.", + "expected_output": "Agent checks Karpenter-specific configuration — EC2NodeClass subnets, security groups, and AMI selectors against the cluster requirements", + "files": [], + "assertions": [ + "Agent checks cluster state and Karpenter detection", + "Agent evaluates EC2NodeClass subnets match cluster VPC (NJ25)", + "Agent checks EC2NodeClass security groups include cluster SG (NJ26)", + "Agent checks AMI compatibility in EC2NodeClass (NJ27)", + "Agent checks if VPC CNI can schedule on the tainted node (NJ28)" + ] + }, + { + "id": "5", + "prompt": "I'm getting 'Unauthorized' errors when my nodes try to join cluster 'dev-eks' in us-east-2. I created the node role with a path: arn:aws:iam::123456789012:role/teams/platform/eks-node-role. I added this ARN to the aws-auth ConfigMap but nodes still show Unauthorized.", + "expected_output": "Agent identifies the IAM role path issue in aws-auth ConfigMap and recommends stripping the path or migrating to EKS Access Entries", + "files": [], + "assertions": [ + "Agent identifies the role path issue (NJ7)", + "Agent explains that aws-auth requires the path to be stripped", + "Agent provides the corrected ARN: arn:aws:iam::123456789012:role/eks-node-role", + "Agent recommends migrating to EKS Access Entries as the long-term fix" + ] + } +] diff --git a/skills/eks-node-join-diagnostic/evals/report.json b/skills/eks-node-join-diagnostic/evals/report.json new file mode 100644 index 0000000..ab01ce6 --- /dev/null +++ b/skills/eks-node-join-diagnostic/evals/report.json @@ -0,0 +1,65 @@ +{ + "skill_name": "eks-node-join-diagnostic", + "skill_path": "/Users/mridulch/Downloads/sample-devops-agent-tools/skills/eks-node-join-diagnostic", + "timestamp": "2026-08-12T02:39:27Z", + "overall_score": 0.64, + "overall_grade": "D", + "passed": false, + "sections": { + "audit": { + "score": 66, + "grade": "D", + "passed": true, + "normalized": 0.66, + "critical": 0, + "warning": 1, + "info": 12 + }, + "functional": { + "overall": 0.69, + "grade": "D", + "passed": false, + "scores": { + "outcome": 0.47, + "process": 1.0, + "style": 0.47, + "efficiency": 0.8199, + "overall": 0.69 + }, + "cost_efficiency": { + "quality_delta": -0.12, + "cost_delta_pct": 14.6, + "classification": "REJECT", + "emoji": "\ud83d\udd34", + "description": "Skill significantly degrades quality" + }, + "estimated_cost": { + "with_skill_per_run": { + "input_cost": 1.1e-05, + "output_cost": 0.015744, + "total_cost": 0.015755, + "model": "sonnet", + "currency": "USD" + }, + "without_skill_per_run": { + "input_cost": 9e-06, + "output_cost": 0.013737, + "total_cost": 0.013746, + "model": "sonnet", + "currency": "USD" + }, + "per_eval_pair": 0.029501, + "total_runs": 5, + "total_cost": 0.1475, + "model": "sonnet", + "currency": "USD" + } + }, + "trigger": { + "pass_rate": 0.5, + "grade": "F", + "passed": false, + "total_queries": 20 + } + } +} \ No newline at end of file diff --git a/skills/eks-node-join-diagnostic/evals/trigger_report.json b/skills/eks-node-join-diagnostic/evals/trigger_report.json new file mode 100644 index 0000000..e7b3cd5 --- /dev/null +++ b/skills/eks-node-join-diagnostic/evals/trigger_report.json @@ -0,0 +1,248 @@ +{ + "skill_name": "eks-node-join-diagnostic", + "skill_path": "/Users/mridulch/Downloads/sample-devops-agent-tools/skills/eks-node-join-diagnostic", + "query_results": [ + { + "query": "My EKS managed node group in us-east-1 shows 'Instances failed to join the kubernetes cluster'. The node group is in Create failed state. Can you diagnose what's wrong?", + "should_trigger": true, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": false, + "mean_input_tokens": 3.0, + "mean_output_tokens": 305.0, + "mean_total_tokens": 308.0 + }, + { + "query": "I launched self-managed nodes for my EKS cluster but they don't appear in kubectl get nodes. The EC2 instances are running. What's preventing them from registering?", + "should_trigger": true, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": false, + "mean_input_tokens": 3.0, + "mean_output_tokens": 607.0, + "mean_total_tokens": 610.0 + }, + { + "query": "Getting TLS handshake timeout when my worker nodes try to connect to the EKS API server. Cluster is private endpoint only.", + "should_trigger": true, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": false, + "mean_input_tokens": 3.0, + "mean_output_tokens": 505.7, + "mean_total_tokens": 508.7 + }, + { + "query": "Karpenter is provisioning nodes on my EKS cluster but they stay NotReady. The instances launch fine but never register with the cluster.", + "should_trigger": true, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": false, + "mean_input_tokens": 3.0, + "mean_output_tokens": 385.7, + "mean_total_tokens": 388.7 + }, + { + "query": "After upgrading to EKS 1.33, my new node group with AL2 AMI won't start. Getting NodeCreationFailure.", + "should_trigger": true, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": false, + "mean_input_tokens": 3.0, + "mean_output_tokens": 712.7, + "mean_total_tokens": 715.7 + }, + { + "query": "kubelet logs on my EKS node show 'Unauthorized' and 'Unable to register node'. I've added the role to aws-auth but it's not working.", + "should_trigger": true, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": false, + "mean_input_tokens": 3.0, + "mean_output_tokens": 828.0, + "mean_total_tokens": 831.0 + }, + { + "query": "My EKS node group shows health issue Ec2SecurityGroupNotFound. The node group is in Degraded state.", + "should_trigger": true, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": false, + "mean_input_tokens": 3.0, + "mean_output_tokens": 547.7, + "mean_total_tokens": 550.7 + }, + { + "query": "InvalidClientTokenId error when my nodes try to join the EKS cluster in ap-southeast-1. What's wrong?", + "should_trigger": true, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": false, + "mean_input_tokens": 3.0, + "mean_output_tokens": 450.7, + "mean_total_tokens": 453.7 + }, + { + "query": "Nodes show 'Container runtime network not ready: cni config uninitialized' and never become Ready.", + "should_trigger": true, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": false, + "mean_input_tokens": 3.0, + "mean_output_tokens": 718.7, + "mean_total_tokens": 721.7 + }, + { + "query": "I need to set up a fully private EKS cluster with nodes in isolated subnets. What VPC endpoints do I need to make sure nodes can join?", + "should_trigger": true, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": false, + "mean_input_tokens": 3.0, + "mean_output_tokens": 592.0, + "mean_total_tokens": 595.0 + }, + { + "query": "My EKS cluster API server is slow. kubectl get pods takes 5 seconds. I'm seeing 429 throttling errors.", + "should_trigger": false, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": true, + "mean_input_tokens": 3.0, + "mean_output_tokens": 571.3, + "mean_total_tokens": 574.3 + }, + { + "query": "Can you review the security posture of my EKS cluster? Check RBAC, pod security, and network policies.", + "should_trigger": false, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": true, + "mean_input_tokens": 30.7, + "mean_output_tokens": 2124.7, + "mean_total_tokens": 2155.3 + }, + { + "query": "I want to upgrade my EKS cluster from 1.29 to 1.30. Can you check if we're ready?", + "should_trigger": false, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": true, + "mean_input_tokens": 4.0, + "mean_output_tokens": 834.7, + "mean_total_tokens": 838.7 + }, + { + "query": "My pods keep getting OOMKilled on EKS. Memory utilization is high. How do I right-size?", + "should_trigger": false, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": true, + "mean_input_tokens": 3.0, + "mean_output_tokens": 758.0, + "mean_total_tokens": 761.0 + }, + { + "query": "Help me set up Karpenter from scratch on my new EKS cluster with Spot and On-Demand instances.", + "should_trigger": false, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": true, + "mean_input_tokens": 3.0, + "mean_output_tokens": 2966.0, + "mean_total_tokens": 2969.0 + }, + { + "query": "I need to create an EKS cluster using Terraform with managed node groups and VPC CNI.", + "should_trigger": false, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": true, + "mean_input_tokens": 5.0, + "mean_output_tokens": 1076.0, + "mean_total_tokens": 1081.0 + }, + { + "query": "My ECS tasks are failing health checks behind an ALB. The targets keep going unhealthy.", + "should_trigger": false, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": true, + "mean_input_tokens": 3.0, + "mean_output_tokens": 612.0, + "mean_total_tokens": 615.0 + }, + { + "query": "Can you generate a cost report for my EKS cluster? I want to know where we're spending the most.", + "should_trigger": false, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": true, + "mean_input_tokens": 3.0, + "mean_output_tokens": 411.0, + "mean_total_tokens": 414.0 + }, + { + "query": "My EKS nodes keep getting terminated by Karpenter consolidation. How do I prevent that?", + "should_trigger": false, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": true, + "mean_input_tokens": 3.0, + "mean_output_tokens": 572.0, + "mean_total_tokens": 575.0 + }, + { + "query": "I'm trying to expose my Kubernetes service through an NLB but it's not getting an external IP.", + "should_trigger": false, + "trigger_count": 0, + "run_count": 3, + "trigger_rate": 0.0, + "passed": true, + "mean_input_tokens": 3.0, + "mean_output_tokens": 745.7, + "mean_total_tokens": 748.7 + } + ], + "summary": { + "total_queries": 20, + "passed": 10, + "failed": 10, + "trigger_precision": 0.0, + "no_trigger_precision": 1.0, + "mean_total_tokens_per_run": 820.8, + "estimated_cost": { + "per_run": { + "input_cost": 1.4e-05, + "output_cost": 0.012243, + "total_cost": 0.012257, + "model": "sonnet", + "currency": "USD" + }, + "total_runs": 20, + "total_cost": 0.2451, + "model": "sonnet", + "currency": "USD" + } + }, + "passed": false +} \ No newline at end of file diff --git a/skills/eks-node-join-diagnostic/references/extended-checks.md b/skills/eks-node-join-diagnostic/references/extended-checks.md new file mode 100644 index 0000000..559c6ec --- /dev/null +++ b/skills/eks-node-join-diagnostic/references/extended-checks.md @@ -0,0 +1,203 @@ +# Extended Diagnostic Checks (NJ28–NJ64) + +Load this reference when the primary checks (NJ1–NJ27) pass but the node still +isn't joining, OR when symptoms point to runtime, kernel, credential, or +network edge cases. + +## Table of Contents + +- [VPC CNI and Taints (NJ28)](#vpc-cni-and-taints-nj28) +- [Join Timeout and IMDS (NJ29–NJ30)](#join-timeout-and-imds-nj29nj30) +- [Containerd / Runtime (NJ31–NJ34)](#containerd--runtime-nj31nj34) +- [VPC CNI / IP Management (NJ35–NJ37)](#vpc-cni--ip-management-nj35nj37) +- [Kernel / OS Compatibility (NJ38–NJ41)](#kernel--os-compatibility-nj38nj41) +- [Credential / Certificate Lifecycle (NJ42–NJ44)](#credential--certificate-lifecycle-nj42nj44) +- [Control Plane Pressure (NJ45–NJ46)](#control-plane-pressure-nj45nj46) +- [Launch Template Edge Cases (NJ47–NJ50)](#launch-template-edge-cases-nj47nj50) +- [Admission Webhooks (NJ51–NJ52)](#admission-webhooks-nj51nj52) +- [EKS Add-on Conflicts (NJ53–NJ55)](#eks-add-on-conflicts-nj53nj55) +- [Network Edge Cases (NJ56–NJ59)](#network-edge-cases-nj56nj59) +- [Account / Credential Provider (NJ60–NJ61)](#account--credential-provider-nj60nj61) +- [Instance Tagging (NJ62)](#instance-tagging-nj62) +- [Control Plane Audit Logs (NJ63–NJ64)](#control-plane-audit-logs-nj63nj64) + +--- + +## VPC CNI and Taints (NJ28) + +| ID | Check | Source | Pass Criteria | Severity | +|----|-------|--------|---------------|----------| +| NJ28 | aws-node DaemonSet can schedule on node | `kubectl get ds aws-node -n kube-system` + node taints | If nodes have custom taints, verify aws-node has matching tolerations (default since VPC CNI 1.7+, but custom manifests may be missing them) | High | + +Symptom: "Container runtime network not ready: NetworkReady=false +reason:NetworkPluginNotReady message: network plugin is not ready: cni config +uninitialized" — means the VPC CNI pod couldn't schedule on the node. + +--- + +## Join Timeout and IMDS (NJ29–NJ30) + +| ID | Check | Source | Pass Criteria | Severity | +|----|-------|--------|---------------|----------| +| NJ29 | Instance not stuck in EC2 Pending state | `ec2.DescribeInstances` | Instance reached `running` within 5 minutes. MNG timeout is ~20 minutes. | Medium | +| NJ30 | IMDS reachable from node (hop limit) | `ec2.DescribeInstances` (MetadataOptions) | `HttpPutResponseHopLimit` >= 2 if containers need IMDS. Hop limit 1 blocks IMDS from containerized kubelet/bootstrap. | High | + +NJ30 detail: IMDS (169.254.169.254) is required for instance identity, region info, and VPC CNI ENI attachment. If `HttpPutResponseHopLimit = 1` and kubelet runs inside a container (Bottlerocket host containers, CIS hardening), the request crosses a network hop and is blocked. Set hop limit to 2. Host firewalls (firewalld, nftables) can also block IMDS. + +--- + +## Containerd / Runtime (NJ31–NJ34) + +| ID | Check | Pass Criteria | Severity | +|----|-------|---------------|----------| +| NJ31 | Containerd registry mirror config valid | Mirror endpoint format matches ECR path structure | High | +| NJ32 | Cgroup driver aligned (kubelet + containerd) | Both use `systemd` (EKS 1.24+ default). Mismatch → immediate NotReady. | Critical | +| NJ33 | Root filesystem supports snapshotter | XFS with `ftype=1` or ext4 for overlayfs | High | +| NJ34 | Instance store mounted before containerd | User-data mounts volume BEFORE containerd starts | High | + +NJ32 symptoms: Node joins then immediately NotReady. kubelet: `"Failed to start ContainerManager"`. Fix: `cgroupDriver: systemd` in kubelet + `SystemdCgroup = true` in containerd. + +--- + +## VPC CNI / IP Management (NJ35–NJ37) + +| ID | Check | Pass Criteria | Severity | +|----|-------|---------------|----------| +| NJ35 | IPAMD warm pool feasible for subnet | `WARM_IP_TARGET` doesn't exceed subnet's available IPs. IPAMD crash-loops if target > available. | Critical | +| NJ36 | Non-default CNI images pre-loaded | If non-AWS CNI (Cilium/Calico): images on AMI or fallback CNI exists | High | +| NJ37 | VPC CIDR no overlap with connected networks | Node subnet doesn't overlap on-premises via DX/TGW | High | + +NJ35 fix: Reduce `WARM_IP_TARGET` to 2-5, set `MINIMUM_IP_TARGET=10`, enable `ENABLE_PREFIX_DELEGATION`. + +--- + +## Kernel / OS Compatibility (NJ38–NJ41) + +| ID | Check | Pass Criteria | Severity | +|----|-------|---------------|----------| +| NJ38 | Kernel version meets EKS requirements | EKS 1.31+: kernel >= 5.10. EKS 1.24-1.30: >= 4.15. | Critical | +| NJ39 | SELinux not blocking kubelet (custom AMI) | kubelet can access `/var/lib/kubelet`, `/var/run/containerd`, `/opt/cni/bin` | High | +| NJ40 | AppArmor not blocking containerd (Ubuntu) | containerd socket access allowed | High | +| NJ41 | Required kernel modules present | `vxlan`, `ip_tables`, `br_netfilter`, `overlay` | High | + +--- + +## Credential / Certificate Lifecycle (NJ42–NJ44) + +| ID | Check | Pass Criteria | Severity | +|----|-------|---------------|----------| +| NJ42 | SSM hybrid activation not expired | EKS Hybrid Nodes: activation ExpirationDate in future (default 30 days) | Critical | +| NJ43 | Kubelet certificate not expired | Instance stopped >90 days: cert may have expired (no auto-rotation while stopped) | High | +| NJ44 | No recent cluster credential rotation | If rotated within last hour: all nodes need rolling restart for new CA | High | + +--- + +## Control Plane Pressure (NJ45–NJ46) + +| ID | Check | Pass Criteria | Severity | +|----|-------|---------------|----------| +| NJ45 | etcd size healthy | etcd < 6GB. Large etcd triggers compaction → slow heartbeats → NotReady | High | +| NJ46 | No API throttling during bootstrap | If 100+ nodes scaling: check for 429s on DescribeCluster/GetToken | Medium | + +NJ46 fix: Stagger scaling, pass bootstrap args directly in user-data to avoid runtime API calls. + +--- + +## Launch Template Edge Cases (NJ47–NJ50) + +| ID | Check | Pass Criteria | Severity | +|----|-------|---------------|----------| +| NJ47 | User-data size <= 16KB compressed | EC2 silently drops oversized user-data. No error in events. | Critical | +| NJ48 | No blocking package updates in user-data | `apt-get update`/`yum update` at boot can hang cloud-init indefinitely | High | +| NJ49 | No NetworkInterfaces + subnet conflict | LT must NOT specify both NetworkInterfaces section AND node group subnets | Critical | +| NJ50 | AMI not from running instance with stale cloud-init | `/var/lib/cloud/` from source cluster causes wrong-cluster join | High | + +NJ47: Move config to S3, user-data only downloads. NJ48: Move ALL package updates to AMI build time. + +--- + +## Admission Webhooks (NJ51–NJ52) + +| ID | Check | Pass Criteria | Severity | +|----|-------|---------------|----------| +| NJ51 | No webhooks blocking node registration | No webhook with `failurePolicy: Fail` targeting nodes that is timing out | High | +| NJ52 | PDBs not blocking node rotation | No PDB with `maxUnavailable: 0` when node group is updating | Medium | + +NJ51 fix: Change `failurePolicy` to `Ignore` for node resources, OR ensure webhook runs on dedicated infra nodes. + +--- + +## EKS Add-on Conflicts (NJ53–NJ55) + +| ID | Check | Pass Criteria | Severity | +|----|-------|---------------|----------| +| NJ53 | kube-proxy version matches cluster | Minor version must match. Never newer than cluster. | High | +| NJ54 | No GPU Operator + EKS device plugin conflict | Use ONE: either EKS addon OR GPU Operator | High | +| NJ55 | Root volume >= 30GB | Undersized root → DiskPressure → evicts system pods → NotReady cycling | High | + +--- + +## Network Edge Cases (NJ56–NJ59) + +| ID | Check | Pass Criteria | Severity | +|----|-------|---------------|----------| +| NJ56 | NAT GW idle timeout compatible | TCP keepalive < 350s on nodes. NAT drops idle connections after 350s. | Medium | +| NJ57 | No ENI detach events | Third-party tools detaching primary ENI → momentary network loss | High | +| NJ58 | conntrack table adequate | `nf_conntrack_max` >= 256K for high-traffic nodes | Medium | +| NJ59 | S3 endpoint allows ECR buckets | Policy must allow `prod--starport-layer-bucket` | Critical | + +NJ59 S3 endpoint policy must include: +- `arn:aws:s3:::prod--starport-layer-bucket/*` +- `arn:aws:s3:::amazon-eks/*` + +--- + +## Account / Credential Provider (NJ60–NJ61) + +| ID | Check | Pass Criteria | Severity | +|----|-------|---------------|----------| +| NJ60 | Account not migrated between Orgs | IAM trust policies with `aws:PrincipalOrgID` fail silently after migration | High | +| NJ61 | AL2023 credential provider binary exists | `--image-credential-provider-config` path must match actual binary location | Critical | + +NJ61 paths: +- Config: `/etc/kubernetes/credential-provider/config.yaml` +- Binary: `/etc/eks/image-credential-provider/ecr-credential-provider` + +--- + +## Instance Tagging (NJ62) + +| ID | Check | Pass Criteria | Severity | +|----|-------|---------------|----------| +| NJ62 | Instance has kubernetes.io/cluster tag | `ec2.DescribeInstances` — instance must have tag `kubernetes.io/cluster/` with value `owned` or `shared` | High | + +This is required for **self-managed nodes** and **Karpenter nodes**. Managed node +groups set this tag automatically. Without it, the cluster may not discover the +node even if it successfully registers. + +Note: If using EKS Auto Mode or Karpenter with `EC2NodeClass`, the tag is applied +automatically via the EC2NodeClass `tags` field. Check the EC2NodeClass spec if +missing. + +--- + +## Control Plane Audit Logs (NJ63–NJ64) + +| ID | Check | Source | Pass Criteria | Severity | +|----|-------|--------|---------------|----------| +| NJ63 | Node CSR not denied | CloudWatch Logs (`/aws/eks//cluster`) — filter `certificatesigningrequests` | No DENY entries for the node's CSR during join window. A denied CSR means the kubelet's certificate request was rejected — node cannot authenticate. | High | +| NJ64 | aws-auth / access entry not recently removed | CloudWatch Logs (`/aws/eks//cluster`) — filter `aws-auth` or CloudTrail `DeleteAccessEntry` | No recent deletion of the node role mapping within the failure window. Accidental removal is a common cause of sudden join failures across all nodes. | High | + +**NJ63 detail:** When a node joins, kubelet sends a CSR to the API server. The +CSR must be auto-approved by the `eks:node-bootstrapper` ClusterRoleBinding. If: +- The node role is not in `system:bootstrappers` group → CSR is never approved +- A custom admission webhook denies CSRs → CSR is explicitly denied +- Check: `aws logs filter-log-events --log-group-name /aws/eks//cluster + --filter-pattern "certificatesigningrequests"` + +**NJ64 detail:** Common scenario: another team accidentally removes entries from +aws-auth ConfigMap or deletes an access entry. All new nodes (and existing nodes +after token refresh) immediately fail with Unauthorized. Check: +- `aws logs filter-log-events --log-group-name /aws/eks//cluster + --filter-pattern "aws-auth"` (for ConfigMap changes) +- CloudTrail: `DeleteAccessEntry` or `UpdateAccessEntry` events