Autoscaling with KEDA
KEDA can scale authentik servers on CPU usage and workers on queued tasks.
A KEDA ScaledObject defines which Deployment to scale and which metrics to use. KEDA creates a Horizontal Pod Autoscaler (HPA) for that Deployment. Disable the chart's HPA for any Deployment managed by KEDA so that two HPAs do not compete. If you only need CPU-based server scaling, the chart HPA requires fewer components.
Prerequisites
- An existing authentik Helm chart installation and its
values.yamlfile. - KEDA installed at a version compatible with your Kubernetes cluster. Installing the Helm chart without
--versionselects the latest chart version from the repository. The KEDA links in this guide refer to the latest version; use the documentation for your installed version if it differs. - For server CPU scaling, a working resource-metrics API, usually provided by Metrics Server. Every container in the server pods needs a CPU request for pod CPU utilization to be defined.
- For worker queue scaling, a Prometheus query endpoint that KEDA can reach and that contains authentik's worker metrics. This example assumes that you have Prometheus managed by the Prometheus Operator. The Operator uses the chart's
ServiceMonitorto configure Prometheus to scrape worker metrics. KEDA queries Prometheus, not authentik's metrics endpoint.
Configure the Helm chart
Add the following settings to the server and worker sections of your existing values.yaml. If KEDA will manage only one Deployment, update only that section. Preserve your other settings.
server:
replicas: 2
autoscaling:
enabled: false
resources:
requests:
cpu: 500m
worker:
replicas: 1
autoscaling:
enabled: false
metrics:
enabled: true
serviceMonitor:
enabled: true
With the chart defaults and release name authentik, worker.metrics.enabled: true creates a ClusterIP Service named authentik-worker-metrics on port 9300. The ServiceMonitor specifies /metrics as the scrape path. Keep the unauthenticated metrics Service internal.
The Prometheus resource must select both the ServiceMonitor's namespace and its labels:
serviceMonitorNamespaceSelectormust include the namespace containing the ServiceMonitor, which defaults to authentik's namespace. An omitted or null selector searches only the Prometheus resource's own namespace.serviceMonitorSelectormust match the ServiceMonitor's labels. Setworker.metrics.serviceMonitor.labelsin your Helm values to add any labels that Prometheus requires.
When switching from the chart HPA, disable it before applying the ScaledObjects. Without the chart HPA, Helm writes spec.replicas from server.replicas and worker.replicas. Switching autoscalers and later Helm upgrades can temporarily reset the replica counts until the HPA reconciles them. Choose replica values that provide enough capacity during that interval, and check the counts after upgrades.
Configure the ScaledObjects
Create ScaledObject in the same namespace as its target Deployment. These examples use authentik-server and authentik-worker in the current namespace. Adjust the names and namespaces (including the worker query's namespace selector) to match your installation. Tune replica limits and thresholds for your workload.
Server CPU usage
This trigger targets an average CPU utilization of 75% across server pods, measured against their CPU requests. For single-container pods requesting 500m CPU each, the target is an average of 375m per pod. The example keeps at least two server replicas running.
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: authentik-server
spec:
scaleTargetRef:
name: authentik-server
minReplicaCount: 2
maxReplicaCount: 5
advanced:
horizontalPodAutoscalerConfig:
behavior:
scaleDown:
stabilizationWindowSeconds: 300
triggers:
- type: cpu
metricType: Utilization
metadata:
value: "75"
Worker queue length
authentik_tasks_queued counts tasks in the QUEUED state by queue and actor. It includes delayed tasks that are not yet due to run (eta). It excludes tasks in progress, tasks waiting for dependencies, and future recurring runs that have not been enqueued. A large number of delayed tasks can therefore cause workers to scale up even when there is little work ready to run.
Each worker exports the queue length from the shared database. The query below takes the maximum per queue and actor before summing, so it does not count the same tasks once for every worker. Because workers are scraped at different times, the result can temporarily reflect an older, higher count.
This example queries a Prometheus Service named prometheus in the monitoring namespace. Set serverAddress to your Prometheus endpoint.
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: authentik-worker
spec:
scaleTargetRef:
name: authentik-worker
minReplicaCount: 1
maxReplicaCount: 5
advanced:
horizontalPodAutoscalerConfig:
behavior:
scaleDown:
stabilizationWindowSeconds: 300
triggers:
- type: prometheus
metricType: AverageValue
metadata:
serverAddress: http://prometheus.monitoring.svc:9090
query: 'sum(max by (queue_name, actor_name) (authentik_tasks_queued{service="authentik-worker-metrics"}))'
threshold: "20"
ignoreNullValues: "false"
Adjust the service selector (and potentially add a namespace selector if running multiple authentik instances) if your Prometheus series use different labels. Workers publish zero for registered actors when their queues are empty.
If the query returns no series, ignoreNullValues: "false" makes KEDA report an error. This does not detect every scrape failure: another worker can still provide matching series. Monitor Prometheus target health as well. If Prometheus requires authentication, configure a KEDA TriggerAuthentication.
With metricType: AverageValue, threshold: "20" sets a target of 20 queued tasks per worker replica. A total queue length of 100 produces a desired count of five replicas, subject to HPA tolerance, stabilization, and replica limits. Adjust the threshold based on how quickly one replica processes tasks and how long tasks can wait.
Deploy and verify
Upgrade your Helm release with the updated values.yaml. This command assumes a release named authentik in the current namespace; adjust both to match your installation:
helm upgrade authentik authentik/authentik -f values.yaml
If scaling workers, check that the ServiceMonitor exists:
kubectl get servicemonitor authentik-worker
Before applying the worker ScaledObject, check in Prometheus that the worker targets are UP and the query above returns one value. A ServiceMonitor alone does not confirm that Prometheus scraped the metrics.
Apply the ScaledObjects and inspect the HPAs that KEDA creates. If scaling only one Deployment, omit the other file and its HPA inspection command:
kubectl apply -f server-scaledobject.yaml -f worker-scaledobject.yaml
kubectl get scaledobject,hpa,deployment
kubectl describe hpa keda-hpa-authentik-server
kubectl describe hpa keda-hpa-authentik-worker
For worker scaling, watch kubectl get deployment authentik-worker --watch as tasks enter and leave the queue. For CPU scaling, watch kubectl get deployment authentik-server --watch under server load.
Capacity and tuning
- Keep at least one worker for scheduled tasks and file monitoring. For high availability, run at least two replicas of each component and schedule them on different nodes.
- Worker concurrency also depends on
AUTHENTIK_WORKER__PROCESSESandAUTHENTIK_WORKER__THREADS. Additional replicas increase PostgreSQL connection demand. They do not resolve bottlenecks in the database or external services. - Tune
scaleDown.stabilizationWindowSecondsto avoid frequent changes in replica count. KEDA'scooldownPeriodapplies only when scaling to zero, so it has no effect with the minimum replica counts shown here. - The queue metric excludes tasks in progress. Scaling down can interrupt active work, and the stabilization window does not guarantee that tasks finish. Review task time limits and
worker.terminationGracePeriodSeconds(30 seconds by default) for your workload. - Each worker metrics scrape refreshes the shared queued-task count from PostgreSQL. More worker replicas mean more database queries during scraping. Monitor database load when scaling on this metric.