Backup, Restore, and Disaster Recovery¶
Audience: SRE, Platform engineer
The ActionsGateway custom resource (CR) is the source of truth for a tenant gateway. The Gateway Manager Controller (GMC) provisions every per-gateway child resource — the Actions Gateway Controller (AGC) Deployment, ServiceAccounts, RoleBinding, Service, NetworkPolicies, and metrics TLS Secrets — from that single CR, and owner-references each child to it. Deleting the CR therefore cascades: the apiserver garbage-collects all owner-referenced children. This makes the gateway cheap to recreate, but it also means an accidental kubectl delete actionsgateway (or a deleted namespace) tears down the running control plane.
This document covers the backup posture that protects against that, and a runnable recovery procedure for restoring a deleted or corrupted gateway.
For component upgrade/rollback (a different concern — the binary version, not the CR) see upgrade.md. For symptom-driven diagnosis see troubleshooting.md.
Table of Contents¶
- What Is and Is Not Owned by the CR
- Backup Posture
- Primary: GitOps
- Secondary: etcd and per-resource backups
- What to Back Up
- Recovery Runbook
- Scenario A: Deleted or Corrupted ActionsGateway CR
- Scenario B: Whole Tenant Namespace Lost
- Scenario C: Full Cluster / etcd Restore
- CR Stuck Terminating
- Verification
What Is and Is Not Owned by the CR¶
Restore planning hinges on one distinction: GMC-generated children reconcile back automatically from the CR; operator-supplied inputs do not. Re-applying the CR rebuilds everything the GMC owns, but it cannot reconstruct a credential it never created.
Reconciles automatically from the CR (owner-referenced, garbage-collected on CR deletion, regenerated on re-apply):
| Child resource | Notes |
|---|---|
AGC Deployment |
Rebuilt from spec; the AGC reconstructs all in-memory session state on startup. |
AGC + worker ServiceAccount |
Per-gateway names. |
AGC RoleBinding |
Namespaced; binds the AGC ServiceAccount to the shipped tenant ClusterRole. |
AGC metrics Service |
mTLS metrics endpoint. |
Workload + AGC NetworkPolicy |
Egress lockdown, rebuilt per egress mode. |
Metrics TLS Secrets (server + client) |
Regenerated fresh by the GMC — new key material each restore. A ServiceMonitor referencing the published client Secret picks up the new material automatically (same Secret name). |
Does not reconcile from the CR — must already exist or be recreated separately:
| Resource | Why it is not restored by re-applying the CR | Recovery |
|---|---|---|
GitHub App credential Secret (spec.gitHubAppRef) |
Operator/tenant-supplied; the GMC only reads it (no owner reference), so it survives a CR-only deletion but is lost if the Secret or namespace is deleted. Without it the gateway sits Ready=False/CredentialUnavailable. |
Recreate from your secret backup before re-applying the CR. See Getting Started §3. |
Tenant Namespace + Pod Security Admission labels |
The namespace is operator-created; the GMC stamps PSA labels but does not create the namespace. | Recreate the namespace (the GMC re-stamps PSA labels on reconcile). |
Namespace ResourceQuota |
Platform-owned — not a field on the CR. See Runbook — Adjusting Tenant Quota. | Re-apply from version control. |
EgressProxy CR (spec.defaultProxyRef) |
A separate CR with its own owned children; not owned by the ActionsGateway. A defaultProxyRef pointing at a missing proxy fails closed (Ready=False/ProxyNotFound). |
Re-apply the EgressProxy CR before (or alongside) the gateway. |
The cluster-scoped ClusterRoleBinding granting the AGC read of ClusterRunnerTemplate |
Cluster-scoped objects cannot carry an owner reference to a namespaced CR, so the GMC deletes it explicitly via its cleanup finalizer and recreates it on reconcile. | Reconciles automatically from the CR — no action needed. |
RunnerSet CRs that target the gateway |
Reference the gateway but are not owned by it; they are never deleted by gateway teardown — they degrade to Ready=False/GatewayNotFound and recover when the gateway returns. |
No action needed; they re-bind automatically. |
In-flight jobs are not recoverable. Running jobs whose renew loop lapses during the outage are cancelled by GitHub and require a manual re-run. Queued (not-yet-acquired) jobs are redelivered to the next healthy session within ~2 minutes. This matches Runbook — AGC Total Failure.
Backup Posture¶
Primary: GitOps¶
Keep the desired state in version control and let a GitOps controller (Argo CD, Flux, or kubectl apply from a pinned manifest repo) be the restore mechanism. This is the recommended posture: the CR is small, declarative, and fully describes the gateway, so a Git repository plus a reconcile loop is your backup and your restore.
Track in version control, per tenant namespace:
- The
ActionsGatewayCR. - Any
EgressProxyCR referenced byspec.defaultProxyRef. - The tenant
Namespacemanifest (including its security-profile label). - The namespace
ResourceQuota.
Never commit the GitHub App credential Secret in plaintext. Manage it with a secrets-aware tool — Sealed Secrets, External Secrets Operator, or SOPS-encrypted manifests — so the encrypted form is safe in Git and the plaintext only ever exists in the cluster. The private key itself is never recoverable from the cluster (it is held only inside the Secret); keep the source .pem in a password manager or secrets vault so a new Secret can be minted if the namespace is lost. See Security Operations for credential-handling guidance.
Secondary: etcd and per-resource backups¶
GitOps protects the desired state. To also capture live status and resources not in Git, add one of:
- etcd snapshots — a cluster-level backup (
etcdctl snapshot save, or your managed-cluster provider's equivalent) captures every object, including the credential Secret and live CR status. This is the broadest safety net and the basis for Scenario C. Treat the snapshot as sensitive: it contains every Secret in the cluster. - Namespace-scoped backups — a tool such as Velero can back up and restore a tenant namespace (CRs, Secrets, quota) as a unit, which is convenient for Scenario B. For concrete Velero backup/restore commands grounded in the ownership model above — including the label selector that skips the GMC-owned children so the controllers rebuild them — see Velero Backup and Restore.
- Ad-hoc CR export — for a point-in-time copy of a single gateway:
Strip the runtime fields before re-applying (status, metadata.uid, metadata.resourceVersion, metadata.creationTimestamp, and the metadata.finalizers entry) — re-applying them can wedge the apply or resurrect a stale finalizer. The spec is the only part you need.
What to Back Up¶
| What | Where it should live | Restores |
|---|---|---|
ActionsGateway CR (spec) |
Git (GitOps) | The whole gateway control plane |
EgressProxy CR (spec), if used |
Git (GitOps) | The egress proxy pool |
Namespace + security-profile label |
Git (GitOps) | The tenant namespace |
ResourceQuota |
Git (GitOps) | Quota enforcement |
GitHub App credential Secret |
Encrypted (Sealed Secrets / ESO / SOPS) and the source .pem in a vault |
GitHub authentication |
| Everything, including live status | etcd snapshot / Velero | Whole-cluster or whole-namespace recovery |
Recovery Runbook¶
Pick the scenario that matches the blast radius.
Scenario A: Deleted or Corrupted ActionsGateway CR¶
The namespace and the GitHub App credential Secret are intact; only the CR is gone or has a broken spec.
- Confirm the blast radius. The owned children are already gone (garbage-collected with the CR) or about to be:
- Confirm the credential Secret survived (it is not owned by the CR, so a CR-only deletion leaves it in place): If it is missing, recreate it first — see Scenario B step 3.
- Re-apply the CR from version control (or your stripped CR export):
If a
defaultProxyRefis set, ensure theEgressProxyCR exists first — apply it in the same step if needed. - Wait for reconcile. The GMC re-provisions all owned children within ~30 seconds and regenerates the metrics TLS Secrets fresh.
- Verify.
Scenario B: Whole Tenant Namespace Lost¶
The namespace and everything in it — CR, credential Secret, quota — are gone.
- Recreate and mark the namespace (including its security-profile label) from version control. See Getting Started §2.
- Re-apply the
ResourceQuotafrom version control. - Recreate the GitHub App credential Secret from your encrypted backup (or mint a fresh one from the source
.pemin your vault). See Getting Started §3. - Re-apply the
EgressProxyCR, if the gateway references one. - Re-apply the
ActionsGatewayCR. - Verify.
RunnerSets that targeted this gateway flip fromReady=False/GatewayNotFoundback to ready automatically once the gateway is up.
Order matters only for the credential Secret and the
EgressProxy: the gateway fails closed (CredentialUnavailable/ProxyNotFound) until both exist, then reconciles on its own once they appear — no gateway re-apply is needed after they land.
Scenario C: Full Cluster / etcd Restore¶
For a control-plane loss, restore etcd from a snapshot per your Kubernetes distribution's documented procedure. After restore:
- Confirm the GMC is running:
kubectl rollout status deploy/gmc-controller-manager -n gmc-system. - The GMC reconciles every restored
ActionsGatewayCR idempotently — it compares desired vs. actual state and only applies what is missing. No resources are duplicated. This is the same idempotent recovery described in Runbook — GMC Total Failure. - Verify each tenant namespace.
CR Stuck Terminating¶
If a deleted CR hangs in Terminating, the GMC's cleanup finalizer (actions-gateway.github.com/gmc-cleanup) is retained because teardown of an owned or cluster-scoped child could not be confirmed (for example, the GMC is down). The finalizer exists precisely to avoid orphaning the cluster-scoped ClusterRoleBinding.
- Check the GMC is running and reconciling — restart it if needed (
kubectl rollout restart deploy/gmc-controller-manager -n gmc-system); it will complete teardown and clear the finalizer. - Only force-remove the finalizer as a last resort when the GMC is permanently gone, and understand the trade-off: it leaves the cluster-scoped
ClusterRoleBinding(agc-clusterrunnertemplate-reader.<namespace>.<name>) orphaned, which you must then delete by hand.
Verification¶
After any restore, confirm the gateway is healthy — the same checks as Runbook — Adding a Tenant:
- The GMC has re-provisioned the children:
- The AGC Deployment is available:
- The CR reports healthy:
Expect
kubectl get actionsgateway -n <namespace> <name> \ -o jsonpath='{.status.conditions[?(@.type=="Ready")].status}{"\n"}'True. IfCredentialUnavailableorProxyNotFound, the credential Secret orEgressProxyis still missing — recheck the relevant scenario step. - Sessions are polling —
actions_gateway_active_sessionsshould reach ≥ 1 per RunnerGroup within seconds. See Observability. - Trigger a test workflow and confirm a worker pod is provisioned and the job completes.
Reference Links¶
- Velero Backup and Restore — tool-specific how-to for namespace-level backup/restore with Velero
- Runbook — day-2 operations and incident response (AGC/GMC total failure)
- Troubleshooting — symptom → diagnosis → resolution
- Security Operations — credential handling and compromise response
- Getting Started — namespace, credential Secret, and CR creation steps
- Upgrade and Rollback — component version changes (distinct from CR restore)