Comparison · ARC alternative
Why GitHub Actions Gateway over ARC?¶
Actions Runner Controller (ARC) scale-set mode struggles with one job: running many runner sets, for many tenants, in one shared cluster — cost-effectively, with each tenant safely capped by its own ResourceQuota. GAG was built for exactly that, without giving up the self-service that makes a shared cluster worth running.
When a worker is evicted, preempted, or blocked by a full ResourceQuota
Failed and the job sits in GitHub's queue until someone reruns it by handThe problem ARC leaves you with¶
The failures compound, but they all trace back to one root: ARC's poor fit with
ResourceQuota makes per-tenant quotas unsafe — and unsafe quotas are what block
letting tenants run their own runners.
-
ResourceQuotais unsafe
A quota-blocked or evicted job can't recover on its own:
-
Critical jobs starve
No way to reserve capacity for expensive runners:
- each
AutoscalingRunnerSetonly caps itself withmaxRunners - no primitive for "GPU always keeps N slots"
- cheap CPU pods exhaust the quota; big tests stall
- each
-
Listener pods pile up
One always-on listener pod per scale set, running 24/7:
- a pod slot + a cluster IP each
- held alive just to long-poll GitHub
- 10 scale sets ≈ 10 always-on pods before a job runs
-
Platform team is the bottleneck
Every tenant is a manual checklist:
- namespace, quota, Role-Based Access Control (RBAC), scale sets, NetworkPolicies, egress
- per-team setup; every later change is a ticket
What changes with GAG¶
GAG vs ARC (scale-set mode)¶
GAG acquires jobs with the same runner-scale-set protocol ARC uses — a single acquirer per runner set, capacity-gated assignment, no many-acquirers fan-out — and it is the shipped default in the v2 API. So the comparison below is capability-for-capability against ARC's own model: every GAG row is additive, not a different-architecture trade-off. The difference is what surrounds the shared acquisition core — quota safety, priority tiers, per-tenant egress, and control-plane footprint.
| Capability | ARC (scale-set mode) | GitHub Actions Gateway |
|---|---|---|
| Runner-scale-set acquisition (single-acquirer, no fan-out) | yes | yes, by default v2 |
| Ephemeral, single-use runner pods | yes | yes |
| Custom runner pod template & image | yes | yes |
| Workers scale to zero between jobs | yes, with minRunners: 0 |
yes, by default |
Safe under a per-tenant ResourceQuota |
quota-blocked jobs stall; manual cleanup + rerun | won't take on a job it can't place live quota headroom is read before the job is claimed — on the default tier it bounds the capacity advertised to GitHub, on classic it declines the claim — so the job stays queued at GitHub. If headroom is lost afterwards, the pod create is retried in place while the lock is held |
| Auto-re-run jobs disrupted mid-flight (eviction / preemption / drain) | runner marked Failed; the job waits in GitHub's queue for a manual rerun |
re-run automatically, with a per-run retry budget kubelet evictions, scheduler preemptions under a priorityTiers floor, node drains, and hand-deleted workers all re-run through GitHub's own re-run API — retried until GitHub accepts it — with maxEvictionRetries capping the budget per run |
| Stop claiming jobs when the cluster can't place the worker | the runner claims its own job, so there is no seat before the claim to decide at | opt-in capacityGate on the runner setoff by default. When on, an unplaceable worker shape (drained pool, changed taint, spot gone) stops the gateway taking on more work, so jobs stay queued at GitHub instead of being claimed and cancelled. It bounds the rate of wasted claims — roughly one per pendingPodDeadline window — it does not eliminate the first one. The tenant turns it on; the platform states once, on the gateway, whether the cluster has a node autoscaler — so a runner set can never gate on a signal that is wrong for the cluster it runs in. Where a node may still arrive, the gate waits for cluster-autoscaler or Karpenter to say it will not add one |
| Guaranteed floor for critical runner types | no per-quota primitive | priority tiers per runner set |
| Throttle the rate new workers start (anti-stampede) | only maxRunners caps the count — a burst starts all at once |
opt-in scaleUp creation-rate limit per setfor shared-egress onset (NAT / firewall / VPN) |
| Per-tenant dedicated egress IPs | shared cluster egress | per-tenant proxy pool v2 proxy optional |
| Listener footprint, 10 runner sets at rest | 10 always-on pods + 10 cluster IPs | 1 shared pod — ~12 KiB per listener session |
| Per-tenant utilization metrics | scale-set metrics, not tenant-scoped | Prometheus per tenant + group job counts in kubectl get; ready-to-apply tenant dashboard + alerts as code |
| Right-size runner resources from measured usage | no feedback loop — runner resources stay a guess |
v2 measured recommendations in RunnerSet status+ opt-in profiles that auto-apply them at pod build ( Binpack/Throughput/NodeShare), a SizingDrift condition, and per-job peak metrics |
| Cross-tenant fleet health view (platform admin) | controller + per-scale-set metrics, aggregated by hand; no bundled dashboard | single-pane GMC fleet rollups degraded / egress-stale / quota per gateway, + a platform dashboard |
| Multiple gateways per namespace | multiple AutoscalingRunnerSets |
v2 multiple scoped gateways per namespace |
| Reusable runner pod templates | template inlined per AutoscalingRunnerSet |
v2 shared RunnerTemplatecluster-wide ClusterRunnerTemplate |
Every capability above is available today.
Disruptions re-run themselves — failures and cancels never do
The auto-re-run claim is scoped on purpose. What re-runs itself: a job whose
worker the cluster took away — a kubelet eviction, a scheduler preemption
under a priorityTiers floor, a node drain, even a stray
kubectl delete pod — each retried until GitHub accepts it, within a
per-run budget (maxEvictionRetries). What never does, by design: a job
that ran and failed (re-running it would mask real breakage), a run you
cancelled (a cancel is the intended stop), and workers the gateway's own
cleanup reaped as stuck (a re-run would loop them). The full boundary, with
the detection marks and metrics:
which disruptions auto-re-run a job.
Why measured right-sizing can't be bolted onto ARC
The right-sizing row is structural, not a feature race. Ephemeral runner
pods — ARC's and GAG's alike — run one job, live minutes, and have no
/scale-style controller to group them, so stock Vertical Pod Autoscaler
(and the dashboards built on it) cannot size them: its grouping, its
evict-and-resize actuation, and its long-running-service statistics all
miss this workload shape. The only place the loop can close is inside the
controller that creates the pods, at pod-build time — which is where GAG
runs it: sample per-job peaks, publish the recommendation on the
RunnerSet, and (opt-in) apply it to the next job's pod, with GPUs never
touched. The full alternatives analysis is in
Appendix D §D.7.
Onboarding: start on v2
New tenants should onboard on the recommended v2 API at
actions-gateway.com/v2beta1 — a decomposed ActionsGateway + RunnerSet +
RunnerTemplate, with an optional standalone EgressProxy; the rows marked
v2 are v2-only. The single-CR v1alpha1 shape
shown below is still fully served but
deprecated, and removed at v2.0.0 — see the
v1 → v2 migration guide and the
getting-started walkthrough for the v2 object set.
The numbers behind these claims
For limits and Service Level Objectives, see Appendix A — Capacity Targets & SLOs; for the utilization-and-cost argument, Appendix F — Cost model.
Where GAG is behind ARC
It's maturity, not capability. ARC is GA and widely deployed; GAG's v2 API
has only just reached beta (v2beta1, its first stability contract) and rides a
Public-Preview runner-scale-set
protocol. That is precisely why the v1 → v2 migration is handled on a committed,
documented schedule with a working
gag-migrate tool — the discipline is the
"won't strand you" signal while the track record accumulates.
Secure by default¶
Built for shared clusters running other teams' code: the multi-tenant hardening ships as reconciled defaults, not a post-install project.
-
Risk reduction
Untrusted job code is boxed in by default:
baselinePod Security Admission (PSA) per namespace- Default-deny network — DNS + own proxy only
- App keys read-only; never in env, never cached
- Controller writes confined to tenant namespaces
-
Lower operational cost
What you'd hand-build around ARC, reconciled from one CR:
- NetworkPolicies · PSA · RBAC · egress
- No Kyverno/OPA required — in-tree PodSecurity
- Kept in sync as tenants come and go
-
Ready out of the box
Secure by default; looser is an explicit opt-in:
- Default-deny ingress, cluster-only DNS
- Per-tenant egress IPs, mutual-TLS metrics
- Signed images + Software Bill of Materials (SBOM) + Supply-chain Levels for Software Artifacts (SLSA) provenance
Sandboxed runtimes compose with these defaults, and GAG's own CI runs that way
A sandboxed runtime is just a runtimeClassName on the worker pod template,
so GAG and ARC can both set one. The differentiator is not that field. It is
the layer underneath it, and the evidence that the combination holds.
Kata bounds the kernel, not the pod network. A micro-VM does not change the pod's network identity: the cloud metadata service still answers from inside the guest, so a compromised job can mint node credentials over the pod network even though the kernel it escaped into is disposable. GAG's default-deny NetworkPolicies are what close that path, reconciled per tenant rather than hand-built alongside. Kata is one layer, not a posture.
GAG's own end-to-end CI runs under it. The suite creates a kind cluster
inside a worker pod with runtimeClassName: kata and zero
privileged: true, validated on a nested-virtualization GKE node pool (node
kernel 6.8.0-1054-gke, guest kernel 6.18.35), and the default for that
suite ever since. The cluster-side how-to is written down, including
the capability set an unprivileged dockerd needs and
what Kata does not buy you.
This is not yet a claim about untrusted pull requests. That CI variant still runs a permissive egress policy, because its jobs pull from CDN-fronted public registries no CIDR allowlist can pin. Closing it takes an in-cluster pull-through mirror plus egress scoped to the mirror, GitHub, and DNS: see the roadmap.
For the full threat model, per-profile controls, and the abuse-response playbooks, see Security and Security operations.
Composable building blocks, not one giant CR¶
A tenant still declares only namespace-scoped resources, and the Gateway Manager
Controller (GMC) provisions the controller, proxy pool, RBAC, and network policies
to match — all within the platform-owned ResourceQuota the GMC never creates
or mutates, with no per-tenant cluster-admin after the initial install. What
changed with the recommended v2 API is that the single-CR monolith is
decomposed into small, reusable kinds — and that decomposition is a
differentiator ARC's inlined, per-scale-set model structurally can't express:
-
Reuse the pod shape
One
RunnerTemplate— or cluster-wideClusterRunnerTemplate— is referenced by everyRunnerSet. ARC inlines the pod template into eachAutoscalingRunnerSet, so N runner types means N copies to keep in sync. -
Clean ownership boundary
Platform owns the quota, the
PriorityClassallowlist, and cluster templates; the tenant composesRunnerSets within them. ARC has no primitive to separate platform-owned from tenant-owned concerns. -
Egress on purpose
A standalone
EgressProxyis referenced by the gateway (or perRunnerSet), or dropped entirely for direct — stillNetworkPolicy-restricted — egress. ARC has no per-tenant egress primitive at all. -
Many gateways, one namespace
Multiple scoped
ActionsGateways coexist in a namespace, each with its own GitHub binding and runner sets — not one CR that must own everything.
The v2 object set below is feature-equivalent to the legacy single-CR example — a proxied gateway with a GPU runner set (priority tiers) and a Linux runner set:
apiVersion: actions-gateway.com/v2beta1
kind: EgressProxy # (1)!
metadata:
name: team-a-egress
namespace: team-a
spec:
minReplicas: 2
maxReplicas: 10
---
apiVersion: actions-gateway.com/v2beta1
kind: RunnerTemplate # (2)!
metadata:
name: default
namespace: team-a
spec:
podTemplate:
spec:
containers:
- name: runner
---
apiVersion: actions-gateway.com/v2beta1
kind: ActionsGateway # (3)!
metadata:
name: team-a-gateway
namespace: team-a
spec:
credentials:
type: GitHubApp
githubApp:
name: my-github-app # name-only Secret ref in this namespace
githubURL: https://github.com/team-a-org
defaultProxyRef:
name: team-a-egress # every RunnerSet inherits this unless it sets proxyRef
---
apiVersion: actions-gateway.com/v2beta1
kind: RunnerSet
metadata:
name: gpu
namespace: team-a
spec:
gatewayRef: { name: team-a-gateway }
templateRef: { name: default } # (4)!
runnerLabels: ["gpu"] # (5)!
priorityTiers: # (6)!
- priorityClassName: runner-critical
threshold: 5
- priorityClassName: runner-standard
threshold: 20
---
apiVersion: actions-gateway.com/v2beta1
kind: RunnerSet
metadata:
name: linux
namespace: team-a
spec:
gatewayRef: { name: team-a-gateway }
templateRef: { name: default }
runnerLabels: ["linux"]
maxWorkers: 30
- Optional. A standalone per-tenant egress proxy pool, Horizontal Pod Autoscaler
(HPA)-managed between these bounds; all GitHub traffic exits through it on
dedicated IPs. Drop it (and
defaultProxyRef) for direct, stillNetworkPolicy-restricted egress — collapsing the minimum to three objects. - A reusable pod shape referenced by both
RunnerSets below viatemplateRef. Define it once; a cluster-scopedClusterRunnerTemplateshares one shape across every namespace. The Pod Security Admission level is a namespace label in v2, not a CR field — all gateways in a namespace share one level. credentials.githubApp.namereferences aSecretin this namespace holding the GitHub AppappId,installationId, andprivateKey. The GMC watches the reference name, not the Secret contents — see credential rotation.WorkloadIdentityis the opt-in no-PEM credential member.- Both runner sets reference the same
RunnerTemplate. There is noResourceQuotafield on any of these CRs — the single quota every runner set shares is platform-owned, set on the namespace by the platform admin, so it is a real cap the tenant cannot raise. Priority tiers decide who wins when it is contended. - Exactly one label per runner set: it is the set's scale-set name at GitHub and
its single
runs-onmatch target (runs-on: gpu), unique across the sets under one gateway. This is the same single-name routing ARC scale sets use, soruns-onlines carry across from ARC unedited. - The first 5 GPU pods get the higher-priority
PriorityClass; the next tier bursts opportunistically; the final threshold caps total concurrency. ThepriorityClassNamevalues must be on the platform's allowlist (the GMC--allowed-priority-classesflag), and whether a tier preempts is set on the platform-ownedPriorityClassobject — a tenant cannot name a class that evicts other tenants' pods.
The legacy single-CR v1alpha1 shape — which expresses this whole gateway in one
ActionsGateway CR — is still fully served but
deprecated, and removed at v2.0.0; see the
getting-started walkthrough
for it and the v1 → v2 migration guide to move
across without changing how your jobs are acquired.
Ready to try it? Follow the getting-started guide. Already running ARC? The Migrating from ARC guide maps every concept above onto GAG and walks one runner group across with zero downtime.