Skip to content

Alerting & SLOs

Audience: SRE, Platform engineer

Part of the Observability guide. The metrics referenced below are catalogued in the Metrics reference; to scrape them, see Accessing metrics. For SLO targets, see Appendix A — Capacity Targets & SLOs.

Symptom → Metric Mapping

Symptom Metric(s) to check Notes
Jobs are slow to start pod_creation_latency_seconds p95/p99 SLO: p95 ≤ 15s, p99 ≤ 60s
Jobs are randomly cancelled renew_job_errors_total Each sustained error risks a job cancellation
Jobs are not being acquired active_sessions (should be ≥ 1 per RunnerGroup), job_acquisition_errors_total Zero sessions = no polling
Jobs are queuing but not starting active_sessions (OK) vs jobs_acquired_total not incrementing Check RateLimited condition
Scale-set jobs assigned but not starting scaleset_jobs_assigned_total rising vs scaleset_jobs_provisioned_total flat Tier wedged; check scaleset_provision_errors_total and worker-pod quota (scale-set has no active_sessions gauge)
Runner credentials are broken token_refresh_errors_total Spikes indicate Secret or GitHub App issue
Evictions causing re-runs eviction_retries_total, eviction_retries_exhausted_total Exhausted budget requires manual intervention
Quota rejecting worker pods quota_retries_total, quota_retries_exhausted_total Sustained retries mean tight quota headroom; exhausted budget requires manual intervention
Throughput decaying job by job agent_recycle_errors_total rising, active_sessions shrinking Agent re-registration failing; see the runbook
Jobs cancelled without ever starting worker_pods_reaped_total{reason="pending_deadline"} Worker pod stuck Pending past the deadline — fix the image/scheduling cause; see the runbook
Jobs running but ACTIVEJOBS shows 0 Check pod phase with kubectl get pods -l actions-gateway/runner-group=<name> (v1) or -l actions-gateway.com/runner-set=<name> (v2) ACTIVEJOBS is updated on pod phase-change events; the column reflects the last reconcile snapshot — not a real-time gauge. A pod that changed phase after the last reconcile will show up after the next event fires.
Proxy autoscaling not working HPA TARGETS showing <unknown> requests.cpu not set on proxy pods
GMC/AGC reconcile broken reconcile_errors_total Non-zero sustained rate indicates operator issue

Apply as code. These rules — and the SLO recording rules below — ship as a directly-appliable PrometheusRule at deploy/monitoring/prometheusrule.yaml. kubectl apply it into a namespace your Prometheus selects rules from instead of copying the YAML below by hand (see deploy/monitoring/README.md). The blocks here are the same rules, reproduced for reference.

The following Prometheus alerting rules map to the SLO targets in Appendix A. Adjust thresholds to match your environment.

groups:
  - name: actions-gateway
    rules:

      # Page: no sessions means no job acquisition
      - alert: ActionsGatewayNoActiveSessions
        expr: |
          actions_gateway_active_sessions == 0
        for: 2m
        labels:
          severity: critical
        annotations:
          runbook_url: "https://actions-gateway.com/operations/runbook/#actionsgatewaynoactivesessions"
          summary: "No active listener sessions for {{ $labels.runner_group }} in {{ $labels.namespace }}"
          description: "The AGC has no open long-poll sessions. Jobs queue indefinitely until sessions are restored."

      # Page: token refresh errors risk job failures within ~1 hour
      - alert: ActionsGatewayTokenRefreshErrors
        expr: |
          rate(actions_gateway_token_refresh_errors_total[5m]) > 0
        for: 5m
        labels:
          severity: critical
        annotations:
          runbook_url: "https://actions-gateway.com/operations/runbook/#actionsgatewaytokenrefresherrors"
          summary: "GitHub App token refresh errors in {{ $labels.namespace }}"
          description: "Token refresh has been failing for 5+ minutes. Sessions will fail once the current token expires (~1 hour)."

      # Page: sustained renewjob failures will cancel running jobs
      - alert: ActionsGatewayRenewJobErrors
        expr: |
          rate(actions_gateway_renew_job_errors_total[5m]) > 0.1
        for: 5m
        labels:
          severity: critical
        annotations:
          runbook_url: "https://actions-gateway.com/operations/runbook/#actionsgatewayrenewjoberrors"
          summary: "RenewJob errors in {{ $labels.namespace }}"
          description: "RenewJob is failing at >0.1/s for 5+ minutes. Running jobs may be cancelled."

      # Page: p99 pod creation latency SLO breach
      - alert: ActionsGatewayPodCreationLatencyP99
        expr: |
          histogram_quantile(0.99,
            rate(actions_gateway_pod_creation_latency_seconds_bucket[5m])
          ) > 60
        for: 5m
        labels:
          severity: critical
        annotations:
          runbook_url: "https://actions-gateway.com/operations/runbook/#actionsgatewaypodcreationlatencyp99"
          summary: "Pod creation p99 latency SLO breach in {{ $labels.namespace }}"
          description: "p99 pod creation latency exceeds 60s SLO. Check quota and node capacity."

      # Ticket: p95 pod creation latency SLO breach
      - alert: ActionsGatewayPodCreationLatencyP95
        expr: |
          histogram_quantile(0.95,
            rate(actions_gateway_pod_creation_latency_seconds_bucket[5m])
          ) > 15
        for: 10m
        labels:
          severity: warning
        annotations:
          runbook_url: "https://actions-gateway.com/operations/runbook/#actionsgatewaypodcreationlatencyp95"
          summary: "Pod creation p95 latency degraded in {{ $labels.namespace }}"
          description: "p95 pod creation latency exceeds 15s SLO. Investigate quota and scheduling."

      # Ticket: eviction budget exhausted — manual re-run required
      - alert: ActionsGatewayEvictionRetriesExhausted
        expr: |
          increase(actions_gateway_eviction_retries_exhausted_total[5m]) > 0
        labels:
          severity: warning
        annotations:
          runbook_url: "https://actions-gateway.com/operations/runbook/#actionsgatewayevictionretriesexhausted"
          summary: "Eviction retry budget exhausted for {{ $labels.runner_group }} in {{ $labels.namespace }}"
          description: "A job's eviction retry budget has been exhausted. Manual re-run required."

      # Ticket: quota retry budget exhausted — manual re-run required
      - alert: ActionsGatewayQuotaRetriesExhausted
        expr: |
          increase(actions_gateway_quota_retries_exhausted_total[5m]) > 0
        labels:
          severity: warning
        annotations:
          runbook_url: "https://actions-gateway.com/operations/runbook/#actionsgatewayquotaretriesexhausted"
          summary: "Quota retry budget exhausted for {{ $labels.runner_group }} in {{ $labels.namespace }}"
          description: "A job was abandoned after its namespace ResourceQuota retry budget was exhausted. Raise the quota or lower maxWorkers, then re-run."

      # Page: the namespace ResourceQuota is rejecting worker pods now (Q82)
      - alert: ActionsGatewayWorkerQuotaExceeded
        expr: |
          actions_gateway_worker_quota_exceeded == 1
        for: 5m
        labels:
          severity: critical
        annotations:
          runbook_url: "https://actions-gateway.com/operations/runbook/#actionsgatewayworkerquotaexceeded"
          summary: "Worker pods being rejected by ResourceQuota for {{ $labels.runner_group }} in {{ $labels.namespace }}"
          description: "The namespace ResourceQuota cannot admit another worker pod; acquired jobs will fail to schedule. Raise the quota or reduce maxWorkers."

      # Page: the ResourceQuota is rejecting proxy replicas now (Q82)
      - alert: ActionsGatewayProxyQuotaExceeded
        expr: |
          actions_gateway_proxy_quota_exceeded == 1
        for: 5m
        labels:
          severity: critical
        annotations:
          runbook_url: "https://actions-gateway.com/operations/runbook/#actionsgatewayproxyquotaexceeded"
          summary: "Proxy replica creation rejected by ResourceQuota for {{ $labels.name }} in {{ $labels.namespace }}"
          description: "The proxy pool is being held below the HPA's target by the namespace ResourceQuota. Raise the quota or lower proxy.maxReplicas."

      # Ticket: capacity can't reach the configured ceiling within quota headroom (Q82)
      - alert: ActionsGatewayQuotaPressure
        expr: |
          actions_gateway_worker_quota_pressure == 1 or actions_gateway_proxy_quota_pressure == 1
        for: 15m
        labels:
          severity: warning
        annotations:
          runbook_url: "https://actions-gateway.com/operations/runbook/#actionsgatewayquotapressure"
          summary: "Quota headroom too low to reach configured ceiling in {{ $labels.namespace }}"
          description: "A proxy or worker pool cannot scale to its configured maximum within the namespace ResourceQuota headroom. Plan a quota increase or lower the ceiling before the next load spike."

      # Ticket: reconcile errors need investigation
      - alert: ActionsGatewayReconcileErrors
        expr: |
          rate(controller_runtime_reconcile_errors_total[5m]) > 0.033
        for: 10m
        labels:
          severity: warning
        annotations:
          runbook_url: "https://actions-gateway.com/operations/runbook/#actionsgatewayreconcileerrors"
          summary: "Reconcile errors in {{ $labels.controller }} for {{ $labels.resource }}"
          description: "Reconcile errors at >2/minute for 10+ minutes. Resources may be stale."

      # Page: worker pods can't be scheduled — capacity is not materializing (Q157)
      - alert: ActionsGatewayWorkersUnschedulable
        expr: |
          actions_gateway_workers_unschedulable == 1
        for: 10m
        labels:
          severity: critical
        annotations:
          runbook_url: "https://actions-gateway.com/operations/runbook/#actionsgatewayworkersunschedulable"
          summary: "Worker pods unschedulable for {{ $labels.runner_group }} in {{ $labels.namespace }}"
          description: "Worker pods are stuck Pending past the scheduling grace because the scheduler can't place them (no matching node / affinity / taints  not quota). Capacity is not materializing; acquired jobs will not start."

      # Page: the GitHub egress IP-range allowlist has gone stale (Q157)
      - alert: ActionsGatewayEgressRulesStale
        expr: |
          actions_gateway_egress_rules_stale == 1
        for: 15m
        labels:
          severity: critical
        annotations:
          runbook_url: "https://actions-gateway.com/operations/runbook/#actionsgatewayegressrulesstale"
          summary: "GitHub egress allowlist stale for {{ $labels.name }} in {{ $labels.namespace }}"
          description: "The gateway's GitHub egress IP-range allowlist has not refreshed within the staleness window; a stalled refresh loop may let the proxy NetworkPolicy drift from GitHub's published ranges, silently dropping new ranges."

      # Ticket: agent re-registration failing — listener capacity decays job by job (Q267)
      - alert: ActionsGatewayAgentRecycleErrors
        expr: |
          rate(actions_gateway_agent_recycle_errors_total[5m]) > 0
        for: 10m
        labels:
          severity: warning
        annotations:
          runbook_url: "https://actions-gateway.com/operations/runbook/#actionsgatewayagentrecycleerrors"
          summary: "Agent recycle errors for {{ $labels.runner_group }} in {{ $labels.namespace }}"
          description: "Single-use JIT agent re-registration is failing. Sustained growth shrinks listener capacity and decays tenant throughput job by job."

      # Ticket: fan-out losers recycling on the fallback timeout — a class of stuck winners (Q266)
      - alert: ActionsGatewayFanoutFallbackTimeout
        expr: |
          rate(actions_gateway_fanout_loser_recycle_deferred_total{outcome="fallback_timeout"}[5m]) > 0
        for: 10m
        labels:
          severity: warning
        annotations:
          runbook_url: "https://actions-gateway.com/operations/runbook/#actionsgatewayfanoutfallbacktimeout"
          summary: "Fan-out recycle fallback timeouts for {{ $labels.runner_group }} in {{ $labels.namespace }}"
          description: "Deduped fan-out losers are recycling on the fallback timeout because their winner never concluded within the bound  a class of stuck winners. Investigate long-running or wedged winning jobs."

      # Ticket: fan-out completion (completejob on a deduped sibling) is failing (Q260)
      - alert: ActionsGatewayAbandonedDeliveryErrors
        expr: |
          rate(actions_gateway_abandoned_delivery_completions_total{outcome="error"}[5m]) > 0
        for: 10m
        labels:
          severity: warning
        annotations:
          runbook_url: "https://actions-gateway.com/operations/runbook/#actionsgatewayabandoneddeliveryerrors"
          summary: "Abandoned-delivery completion errors for {{ $labels.runner_group }} in {{ $labels.namespace }}"
          description: "The winner of a fanned-out job is failing to issue completejob on a deduped sibling delivery; the affected jobs may be cancelled at GitHub's ~15-minute unstarted-job timeout. Investigate the run service's completejob responses."

      # Ticket: egress proxy denying CONNECTs to off-allowlist destinations —
      # an SSRF / egress-policy signal, sharper than dial_errors (Q316)
      - alert: ActionsGatewayProxyConnectDenied
        expr: |
          rate(actions_gateway_proxy_connect_denied_total[5m]) > 0.1
        for: 10m
        labels:
          severity: warning
        annotations:
          runbook_url: "https://actions-gateway.com/operations/runbook/#actionsgatewayproxyconnectdenied"
          summary: "Egress proxy denying CONNECTs in {{ $labels.namespace }}"
          description: "The egress proxy is refusing CONNECT requests to off-allowlist destinations at >0.1/s for 10m  an SSRF / egress-policy signal (a workload probing blocked destinations, or a misconfigured egress target). Unlike dial errors, every denial here is an explicit allowlist rejection."

      # Page: scale-set tier wedged — jobs are being assigned but none are
      # getting provisioned (Q311). This is the scale-set analog of
      # ActionsGatewayNoActiveSessions: a ScaleSet-protocol RunnerSet never
      # emits actions_gateway_active_sessions, so throughput stalls are only
      # visible as demand (assigned) flowing while supply (provisioned) is flat.
      - alert: ActionsGatewayScaleSetProvisioningStalled
        expr: |
          (
            sum by (namespace, runner_set) (
              rate(actions_gateway_scaleset_jobs_assigned_total[15m])
            ) > 0
          )
          unless
          (
            sum by (namespace, runner_set) (
              rate(actions_gateway_scaleset_jobs_provisioned_total[15m])
            ) > 0
          )
        for: 10m
        labels:
          severity: critical
        annotations:
          runbook_url: "https://actions-gateway.com/operations/runbook/#actionsgatewayscalesetprovisioningstalled"
          summary: "Scale-set provisioning stalled for {{ $labels.runner_set }} in {{ $labels.namespace }}"
          description: "The scale-set acquisition tier is receiving JobAssigned messages but has provisioned no worker pods for 10+ minutes  the tier is wedged. Acquired jobs will not start. Check actions_gateway_scaleset_provision_errors_total, the worker-pod ResourceQuota, and the listener session health."

      # Ticket: scale-set provision attempts failing at a sustained rate (Q311)
      - alert: ActionsGatewayScaleSetProvisionErrors
        expr: |
          rate(actions_gateway_scaleset_provision_errors_total[5m]) > 0.1
        for: 10m
        labels:
          severity: warning
        annotations:
          runbook_url: "https://actions-gateway.com/operations/runbook/#actionsgatewayscalesetprovisionerrors"
          summary: "Scale-set provision errors for {{ $labels.runner_set }} in {{ $labels.namespace }}"
          description: "The scale-set acquisition tier is failing to provision worker pods (JIT-config mint or pod create) at >0.1/s for 10m. A transient failure retries on a later poll, but a sustained rate means provisioning is degraded  check the run service's generate-jitconfig responses and namespace quota headroom."

SLO Recording Rules

These recording rules pre-compute the metrics needed for burn-rate alerting against the SLO targets in Appendix A. Apply them alongside the alert rules above.

groups:
  - name: actions-gateway-slos
    interval: 30s
    rules:

      # Pod creation latency — p95 and p99 per namespace
      - record: actions_gateway:pod_creation_latency_seconds:p95
        expr: |
          histogram_quantile(0.95,
            sum by (namespace, le) (
              rate(actions_gateway_pod_creation_latency_seconds_bucket[5m])
            )
          )

      - record: actions_gateway:pod_creation_latency_seconds:p99
        expr: |
          histogram_quantile(0.99,
            sum by (namespace, le) (
              rate(actions_gateway_pod_creation_latency_seconds_bucket[5m])
            )
          )

      # Job duration — p50, p95, p99 per namespace and runner_group
      - record: actions_gateway:job_duration_seconds:p50
        expr: |
          histogram_quantile(0.50,
            sum by (namespace, runner_group, le) (
              rate(actions_gateway_job_duration_seconds_bucket[5m])
            )
          )

      - record: actions_gateway:job_duration_seconds:p95
        expr: |
          histogram_quantile(0.95,
            sum by (namespace, runner_group, le) (
              rate(actions_gateway_job_duration_seconds_bucket[5m])
            )
          )

      # Token refresh error rate (hourly) — compare against the <1/hr SLO
      - record: actions_gateway:token_refresh_errors:rate1h
        expr: |
          sum by (namespace) (
            increase(actions_gateway_token_refresh_errors_total[1h])
          )

      # Job acquisition success rate — fraction of acquisitions that succeed,
      # per namespace. Grouped by namespace only (not runner_group):
      # job_acquisition_errors_total is labelled namespace+reason with no
      # runner_group, so grouping the denominator by runner_group would leave
      # the error rate unmatched and the ratio would evaluate to empty.
      - record: actions_gateway:job_acquisition_success_rate:rate5m
        expr: |
          sum by (namespace) (
            rate(actions_gateway_jobs_acquired_total[5m])
          )
          /
          (
            sum by (namespace) (
              rate(actions_gateway_jobs_acquired_total[5m])
            )
            +
            sum by (namespace) (
              rate(actions_gateway_job_acquisition_errors_total[5m])
            )
          )

      # Scale-set provision success rate — fraction of provision attempts that
      # succeeded, per namespace. The scale-set analog of
      # job_acquisition_success_rate: a provision attempt either succeeds
      # (…_jobs_provisioned_total) or fails (…_provision_errors_total).
      - record: actions_gateway:scaleset_provision_success_rate:rate5m
        expr: |
          sum by (namespace) (
            rate(actions_gateway_scaleset_jobs_provisioned_total[5m])
          )
          /
          (
            sum by (namespace) (
              rate(actions_gateway_scaleset_jobs_provisioned_total[5m])
            )
            +
            sum by (namespace) (
              rate(actions_gateway_scaleset_provision_errors_total[5m])
            )
          )

← Back to Observability