You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Window: ~24h ending 2026-09-11 (audit run §34562117697)
Clean day. Of 108 workflow runs analyzed, zero confirmed safe_outputs job failures were found — the second 100% day in this audit history (the first was 2026-09-04). None of the 11 actively-tracked recurring failure signatures (PR Sous Chef approve_workflow_run/update_pull_request issues, Design Decision Gate allowed-files decline, Smoke Copilot hardcoded discussion ID, Smoke Project/Smoke Issues credential failures, submit_pull_request_review non-PR-context, concurrent same-PR write race, create_discussion TypeError) occurred today.
Safe Output Job Statistics
Category
Runs
Confidence
safe_outputs succeeded (job-level data confirmed via audit.json)
55
High — direct
safe_outputs succeeded (inferred: run-level conclusion "success", no per-job artifact available)
30
Medium — architectural inference*
safe_outputs skipped (benign — upstream agent job itself skipped, e.g. yamllint found nothing)
1
High — direct
safe_outputs failed
0
High — direct
Not reached (workflow still in_progress at snapshot time)
3
N/A
Agent-job failure, zero safe-output items queued (out of scope — see below)
19
High for 2 sampled; extrapolated for remainder
Total examined
108
*Architectural inference: the job pipeline in this repo runs a conclusion job that depends on safe_outputs; a safe_outputs failure cascades to an overall run failure. Runs with overall conclusion "success" but no retrievable per-job breakdown are therefore treated as safe_outputs success with medium confidence, not fabricated as directly confirmed.
Success rate: 55/55 = 100% on directly-confirmed data; 85/85 = 100% including the medium-confidence inferred set. Zero confirmed failures either way.
The 19-run cluster: confirmed out of scope, not a safe-output problem
19 runs (13 across "Test Quality Sentinel", "Impeccable Skills Reviewer", "PR Code Quality Reviewer", "Ponytail Reviewer", "Design Decision Gate", "Matt Pocock Skills Reviewer"; plus "AI Moderator" ×2 and "Front Page Copy Guard" ×1) had an overall run conclusion of "failure" with only a generic audit.json message ("Workflow X failed with 1 error(s)") and no jobs array, making them look ambiguous at first pass. Fresh targeted log pulls (via agenticworkflows logs, artifacts:["all"]) against two representative samples — run §34552425032 (Test Quality Sentinel) and run §34556225318 (AI Moderator) — both show total_safe_items: 0 and write_runs: 0: the agent job failed (per prior history, likely at the CLI-execution step) before any safe-output item was ever queued. Six of the thirteen review-bot runs (all 34552425*) additionally share an identical trigger commit, branch, and duration fingerprint, consistent with one upstream event affecting a batch of PR-review workflows simultaneously rather than 13 independent issues. Per the scope of this monitor (safe-output jobs only, agent-job failures explicitly excluded), this cluster requires no action from this report, but is noted for awareness of the agent-job-failure monitor since it was not clear from summary-level data alone.
Methodology and full stats
Enumerated all 108 run directories under /tmp/gh-aw/aw-mcp/logs/run-* (run-34550588344 through run-34562144183) via Glob, then read audit.json per run (falling back to targeted Grep against run_summary.json where audit.json was absent) to extract the jobs array and the safe_outputs entry conclusion for each run, delegated to a subagent to keep this within budget.
Where audit.json lacked a jobs array entirely, cross-checked a 2-run sample with a fresh agenticworkflows logs pull (artifacts:["all"]) to confirm total_safe_items: 0 before concluding the cluster is an agent-job issue, not a safe_outputs job issue.
Excluded 3 in-progress runs (including the run for this monitor itself) from the failure-rate denominator rather than guessing the outcome.
No gh CLI or write APIs used; all reads via Read/Glob/Grep on cached logs plus the read-only agenticworkflows logs tool.
All previously-tracked open items remain unshipped and unchanged today (no new occurrences, no new information):
grant-actions-write-fork-pr-scope (PR Sous Chef fork-PR permission gap) — still proposed.
reclassify-protected-file-decline-as-non-failure (approve_workflow_run protected-files decline) — still proposed.
reclassify-allowed-files-decline-as-non-failure (push_to_pull_request_branch allowed-files decline, Design Decision Gate) — still proposed, unresolved since 2026-08-28.
fix-hardcoded-smoke-discussion-temp-id (Smoke Copilot variants) — still proposed.
reclassify-submit-pr-review-no-context-as-skip (Smoke Claude) — still proposed.
retry-update-branch-with-fresh-head-sha (PR Sous Chef update_pull_request stale SHA, first seen 2026-09-10) — still only 1 occurrence; no recurrence today.
investigate-smoke-project-bad-credentials, investigate-smoke-issues-jira-linear-credentials — no Smoke Project or Smoke Issues runs appeared in the 108-run window today at all (either did not run, or scheduled outside this window), so no new data either way.
Recommendations
No critical action required — zero confirmed safe_outputs job failures today.
Bug-fixes backlog (unchanged, still unshipped, no new occurrences today): the 6 open recurring-failure items above remain the actionable backlog; none regressed or recurred in this window.
Config/process: consider whether the 19-run degraded-audit cluster (no jobs array despite a real run) points to an audit.json generation gap worth flagging to whoever owns that tooling — not a safe-output bug, but it cost extra investigation effort today to rule out.
Next steps: no new work items opened this cycle. Continue monitoring retry-update-branch-with-fresh-head-sha for a 2nd occurrence before committing to a specific fix design.
Overall safe-output health remains excellent — 100% for the second time in the 3-week audit history.
References:
§34552425032 — Test Quality Sentinel, sampled to confirm zero safe-output items (out of scope)
§34556225318 — AI Moderator, sampled to confirm zero safe-output items (out of scope)
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Safe Output Health Monitor — 2026-09-11
Window: ~24h ending 2026-09-11 (audit run §34562117697)
Clean day. Of 108 workflow runs analyzed, zero confirmed
safe_outputsjob failures were found — the second 100% day in this audit history (the first was 2026-09-04). None of the 11 actively-tracked recurring failure signatures (PR Sous Chefapprove_workflow_run/update_pull_requestissues, Design Decision Gate allowed-files decline, Smoke Copilot hardcoded discussion ID, Smoke Project/Smoke Issues credential failures, submit_pull_request_review non-PR-context, concurrent same-PR write race, create_discussion TypeError) occurred today.Safe Output Job Statistics
safe_outputssucceeded (job-level data confirmed via audit.json)safe_outputssucceeded (inferred: run-level conclusion "success", no per-job artifact available)safe_outputsskipped (benign — upstreamagentjob itself skipped, e.g. yamllint found nothing)safe_outputsfailedin_progressat snapshot time)*Architectural inference: the job pipeline in this repo runs a
conclusionjob that depends onsafe_outputs; asafe_outputsfailure cascades to an overall run failure. Runs with overall conclusion "success" but no retrievable per-job breakdown are therefore treated assafe_outputssuccess with medium confidence, not fabricated as directly confirmed.Success rate: 55/55 = 100% on directly-confirmed data; 85/85 = 100% including the medium-confidence inferred set. Zero confirmed failures either way.
The 19-run cluster: confirmed out of scope, not a safe-output problem
19 runs (13 across "Test Quality Sentinel", "Impeccable Skills Reviewer", "PR Code Quality Reviewer", "Ponytail Reviewer", "Design Decision Gate", "Matt Pocock Skills Reviewer"; plus "AI Moderator" ×2 and "Front Page Copy Guard" ×1) had an overall run conclusion of "failure" with only a generic
audit.jsonmessage ("Workflow X failed with 1 error(s)") and nojobsarray, making them look ambiguous at first pass. Fresh targeted log pulls (viaagenticworkflows logs,artifacts:["all"]) against two representative samples — run §34552425032 (Test Quality Sentinel) and run §34556225318 (AI Moderator) — both showtotal_safe_items: 0andwrite_runs: 0: theagentjob failed (per prior history, likely at the CLI-execution step) before any safe-output item was ever queued. Six of the thirteen review-bot runs (all34552425*) additionally share an identical trigger commit, branch, and duration fingerprint, consistent with one upstream event affecting a batch of PR-review workflows simultaneously rather than 13 independent issues. Per the scope of this monitor (safe-output jobs only, agent-job failures explicitly excluded), this cluster requires no action from this report, but is noted for awareness of the agent-job-failure monitor since it was not clear from summary-level data alone.Methodology and full stats
/tmp/gh-aw/aw-mcp/logs/run-*(run-34550588344 through run-34562144183) via Glob, then readaudit.jsonper run (falling back to targeted Grep againstrun_summary.jsonwhere audit.json was absent) to extract thejobsarray and thesafe_outputsentry conclusion for each run, delegated to a subagent to keep this within budget.audit.jsonlacked ajobsarray entirely, cross-checked a 2-run sample with a freshagenticworkflows logspull (artifacts:["all"]) to confirmtotal_safe_items: 0before concluding the cluster is an agent-job issue, not a safe_outputs job issue.ghCLI or write APIs used; all reads via Read/Glob/Grep on cached logs plus the read-onlyagenticworkflows logstool.Trend vs. prior audit history
Recent daily
safe_outputssuccess rates: 08-22 99.45%, 08-23 98.94%, 08-25 98.92%, 08-26 99.34%, 08-28 98.6%, 08-29 98.3%, 08-30 99.66%, 08-31 99.52%, 09-02 98.84%, 09-03 98.95%, 09-04 100.0%, 09-05 97.4%, 09-06 99.57%, 09-10 98.1%, 09-11 100.0% (today).All previously-tracked open items remain unshipped and unchanged today (no new occurrences, no new information):
grant-actions-write-fork-pr-scope(PR Sous Chef fork-PR permission gap) — still proposed.reclassify-protected-file-decline-as-non-failure(approve_workflow_run protected-files decline) — still proposed.reclassify-allowed-files-decline-as-non-failure(push_to_pull_request_branch allowed-files decline, Design Decision Gate) — still proposed, unresolved since 2026-08-28.fix-hardcoded-smoke-discussion-temp-id(Smoke Copilot variants) — still proposed.reclassify-submit-pr-review-no-context-as-skip(Smoke Claude) — still proposed.retry-update-branch-with-fresh-head-sha(PR Sous Chefupdate_pull_requeststale SHA, first seen 2026-09-10) — still only 1 occurrence; no recurrence today.investigate-smoke-project-bad-credentials,investigate-smoke-issues-jira-linear-credentials— no Smoke Project or Smoke Issues runs appeared in the 108-run window today at all (either did not run, or scheduled outside this window), so no new data either way.Recommendations
jobsarray despite a real run) points to anaudit.jsongeneration gap worth flagging to whoever owns that tooling — not a safe-output bug, but it cost extra investigation effort today to rule out.retry-update-branch-with-fresh-head-shafor a 2nd occurrence before committing to a specific fix design.Overall safe-output health remains excellent — 100% for the second time in the 3-week audit history.
References:
All reactions