Skip to content
← Public packages

@kentcdodds/sentry-triage

Sentry triage wakes Cole (Grok Bot) per issue: one active wake per repo (lease + queue), loop-safe Discord. Cole may spawn Cursor for isolated repo work.

src/index.ts

95 lines · 5.2 KB · TypeScript
import { readTriageSettings } from './settings.ts'
import type { DescribeSentryTriageResult } from './types.ts'

/**
 * Sentry triage automation.
 *
 * Flow (Seer-first): a new Sentry issue hits the HMAC-verified inbound
 * webhook; `handle-sentry-webhook` acknowledges quickly and dispatches
 * durable work via `workflows.create` to `process-sentry-webhook`, which
 * routes the issue to its triage project (see the registry in shared.ts),
 * requests a Seer root-cause analysis, and defers. When Seer finishes, the
 * `seer.root_cause_completed` webhook (same endpoint) is likewise
 * backgrounded; the processor routes by resource header to
 * `handle-seer-webhook` and wakes Cole (Grok Bot) seeded
 * with the RCA as a verified hypothesis — so Cole starts from Sentry's
 * investigation instead of redoing it. Cole fixes (possibly via Cursor), filters,
 * ignores, or recommends, then calls `record-outcome`, which edits the
 * issue's single Discord status message in place. If Seer cannot run (or
 * stalls past a grace window), the issue path falls back to from-scratch
 * triage so nothing is dropped.
 *
 * Per-repo lease: at most one active Cole wake per repository. Issues that
 * arrive while the lease is held enqueue; when Cole reports (or a
 * stale lease is swept), queued issues flush into one batch wake. This stops
 * deploy-induced error bursts from waking concurrent workers that collide on
 * the same fix (measured: 12 agents → 5 shared PRs, including a merge-conflict
 * that itself became a new Sentry issue).
 *
 * Seer path: the same internal integration also subscribes to Seer autofix
 * webhooks. Because a Sentry internal integration has a single webhook URL,
 * `seer.*` events (`sentry-hook-resource: seer`) arrive at the same
 * HMAC-verified endpoint and `handle-sentry-webhook` routes them to
 * `handle-seer-webhook`. On `seer.root_cause_completed` it either wakes
 * Cole with Sentry's RCA as a verified hypothesis (skipping
 * open-ended investigation) or — for projects whose Seer stopping point is
 * "open_pr" — holds the issue while Seer drafts a PR. On `seer.pr_created`
 * it wakes Cole to review Seer's draft PR
 * (adversarial review + ship-pr loop, same risk-gated merges), with a
 * 45-minute grace fallback to an RCA-seeded Cole wake when the PR
 * never arrives; when Cole is already on the issue it only annotates
 * the Discord message so Seer's PR is reviewed rather than duplicated.
 *
 * Guardrails: one Cole wake per issue ever, one active wake per repo (lease +
 * batch queue), 10 wakes/hour cap (shared across projects and both ingress
 * paths), anomaly breaker (>30 new issues/hour pauses triage with one alert),
 * and a loop guard Cole runs first via `get-issue-state`.
 *
 * Completeness backstop: webhook deliveries can be silently lost (observed
 * during worker deploys), so a daily package-manifest job runs `./reconcile` to
 * re-deliver recently-first-seen issues with no triage record — and issues
 * stuck awaiting a Seer RCA past the grace window — through the normal
 * webhook handler. Quiet days after reconcile (or after an issue.created
 * deferral) used to strand `awaiting-seer` records until the next cron;
 * `./sweep-stale-awaiting` runs every 15 minutes (paginated packageStorage list) so the grace window has
 * a driver that does not depend on another inbound issue.
 */
export default async function describeSentryTriage(): Promise<DescribeSentryTriageResult> {
	const settings = await readTriageSettings()
	return {
		name: settings.selfPackageName,
		webhooks: [
			'sentry (issue created + Seer autofix lifecycle, HMAC-verified; Seer-first)',
		],
		projects: settings.projects.map((project) => ({
			sentryProject: project.slug,
			repository: project.repoSlug,
		})),
		exports: {
			'./handle-sentry-webhook':
				'Inbound Sentry webhook entry (issue + Seer); acks and dispatches durable process workflow.',
			'./process-sentry-webhook':
				'Durable issue/Seer processor (prefetch, Discord, Seer handoff, Cole wake).',
			'./handle-seer-webhook':
				'Seer autofix lifecycle handler; wakes Cole from Seer RCA.',
			'./get-issue-state': 'Loop-guard lookup for Cole.',
			'./record-outcome':
				'Final Cole report; edits the Discord status message; flushes per-repo queue.',
			'./update-card':
				'Progressive Discord card links (started/pr/merged/deployed) without closing the issue.',
			'./flush-queue':
				'Maintenance: force-release a repo lease and flush its batch queue.',
			'./reconcile':
				'Webhook-loss backstop: re-deliver missed or Seer-stuck issues (daily job; manual backfill).',
			'./sweep-stale-awaiting':
				'Quiet-period driver: re-deliver awaiting-Seer records past the grace window (15m job).',
			'./reset-issue': 'Maintenance: clear one issue record + claim for re-triage.',
			'./triage-report': 'Maintenance: dump issue records and hourly counters.',
			'./settings': 'Read or persist owner Discord/Sentry/project settings in package storage.',
			'./adapt': 'After a community fork, return owner settings and exact swap steps.',
			'./configure-wake-webhook': 'Store minted grok-bot wake webhook URL in package storage.',
			'./probe-wake-webhook': 'Probe stored Cole wake webhook without returning the URL.',
		},
		discordChannel: settings.discordChannelId,
	}
}