Skip to content

Send an alert webhook when a connector keeps failing - #976

Open
rangarajan19 wants to merge 1 commit into
HelpCode-ai:mainfrom
rangarajan19:feat/857-connector-failure-webhook
Open

rangarajan19 wants to merge 1 commit into
HelpCode-ai:mainfrom
rangarajan19:feat/857-connector-failure-webhook

Conversation

@rangarajan19

Copy link
Copy Markdown
Contributor

Summary

Backend half of #857 (PR 1 of the planned split, UI card to follow in PR 2): detects a connector failing repeatedly within a sliding window and dispatches a signed webhook (or Slack message), with a cooldown so a sustained outage sends one alert, not one per call.

Changes

  • New AlertsModule: per-organization webhook config stored via OrgSettingsService, signing secret generated server-side and encrypted (AES-256-GCM, AAD-bound to the org), shown once. Admin endpoints: GET/PUT/DELETE + POST .../test at /api/admin/settings/alert-webhook.
  • Detection hooks into AuditService.logInvocation, firing before the repeat-failure dedup check added since this plan was discussed, so a fast-repeating identical failure (the common case for a real outage) still counts toward the threshold instead of being silently skipped. Fire-and-forget: a dispatch failure is caught and logged, never affects the MCP call path.
  • Sliding-window failure counter and cooldown via RedisService (incr/expire), with an in-memory fallback when Redis isn't configured. Dispatch fires only when the counter hits the threshold exactly, and the cooldown key is written before the dispatch goes out, so two failures crossing the threshold at the same moment can't both fire.
  • Dispatch is HMAC-SHA256 signed (X-AnythingMCP-Signature) and goes through the existing SSRF-guarded outbound path. The webhook URL deliberately ignores both the env (SSRF_ALLOWED_HOSTS) and the admin-editable DB allowlist — ssrf.util.ts/outbound-http.ts gain an opt-in skipAllowlists option for this, since those allowlists exist for connectors to reach a specific internal host for their own reason, and a webhook target must not inherit that trust.
  • Docs: new section in docs/operations/observability.md — payload shape, headers, a Node verification snippet.

PR 2 (the Alerts card on /settings/admin) will follow once this is reviewed.

Testing

  • Added/updated unit tests
  • Tested manually (describe steps below)
  • All existing tests pass (npm test in packages/backend)

New: src/alerts/alerts.service.spec.ts, src/alerts/alerts.controller.spec.ts, and additions to src/audit/audit.service.spec.ts covering threshold→one dispatch, cooldown suppressing the next, cross-org isolation, signature verification, dispatch errors never affecting logInvocation, SSRF (private IP and the cloud metadata address rejected even when allowlisted), and non-admin → 403.

Ran from packages/backend:

npx tsc --noEmit -p tsconfig.json
npx eslint src/alerts src/audit/audit.module.ts src/audit/audit.service.ts src/common/ssrf.util.ts src/common/outbound-http.ts
npx jest src/alerts src/audit/audit.service.spec.ts src/audit/audit.breakdowns.spec.ts src/common/ssrf.util.spec.ts src/common/outbound-http.spec.ts src/common/outbound-fetch.util.spec.ts src/settings

All clean; 747 tests passing in the scoped run above (the files this PR touches and everything that depends on them).

Related Issues

Relates to #857

Detects a connector failing repeatedly (configurable threshold over a
sliding window) and dispatches a signed webhook or Slack message, with
a cooldown so a sustained outage sends one alert, not one per call.

- New AlertsModule: per-org webhook config (OrgSettings), encrypted
  secret shown once, GET/PUT/DELETE + POST .../test admin endpoints.
- Detection hooks into AuditService.logInvocation ahead of the
  repeat-failure dedup check, so a fast-repeating identical failure
  still counts toward the threshold. Fire-and-forget; never affects
  the MCP call path.
- Sliding-window counter and cooldown via Redis (incr/expire), with
  an in-memory fallback when Redis isn't configured.
- Dispatch is HMAC-signed and SSRF-guarded; the webhook URL ignores
  the env/DB connector allowlists so a workspace can't point it at a
  host allowed for an unrelated reason (ssrf.util.ts/outbound-http.ts
  gain an opt-in skipAllowlists flag for this).
- Docs: new section in docs/operations/observability.md.

Relates to HelpCode-ai#857.
: JSON.stringify(payload);
const timestamp = Math.floor(Date.now() / 1000).toString();
const signature = createHmac('sha256', secret)
.update(`${timestamp}.${body}`)
@rangarajan19

Copy link
Copy Markdown
Contributor Author

Hi @keysersoft, PR 1 for the connector-failure alert webhook is up here. While working on it I noticed #857 (and the whole #840-867 range, including the October Challenge issue #846) is no longer accessible — just a 404, not sure if that was intentional. Do you still want this feature added, or has priority changed? Happy to keep going as-is if so.

@rangarajan19

Copy link
Copy Markdown
Contributor Author

Hi @mirkopoloni , do we need this, because the issue is not showing , should I delete this or ?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants