KEDA Explained: Event-Driven Kubernetes Autoscaling
KEDA event-driven Kubernetes autoscaling helps teams scale workers from the demand waiting in a queue, stream, or other external system instead of relying only on CPU and memory. This explainer is for developers and platform operators who already run containerized workloads in Kubernetes and need to understand where KEDA fits, how it activates workloads, and what to check before using it in production.
What Is KEDA?
KEDA, the Kubernetes Event-Driven Autoscaling project, is an autoscaler that connects Kubernetes workloads to signals from external event sources. A scaler reads a source-specific measurement, such as the number of messages waiting in a queue, and exposes that demand to Kubernetes so the workload can add or remove replicas. The KEDA documentation describes the supported scalers and the Kubernetes resources used to configure them.
KEDA is not a message broker, job queue, or consumer runtime. Your application still receives and processes messages using its own client library. KEDA observes a metric about that workload and adjusts its replica count; it does not deliver, acknowledge, or retry application messages.
The Problem KEDA Solves
Kubernetes’ built-in Horizontal Pod Autoscaler (HPA) can scale a Deployment or StatefulSet from resource metrics such as CPU and memory. Those signals are useful for services whose work grows with compute use, but they can lag behind demand in a backlog. A worker may be mostly idle while messages accumulate, or a newly created worker may need several minutes of backlog before its CPU crosses a target.
Scaling an asynchronous worker presents a second problem: when its queue is empty, there may be no running pods to produce CPU or memory metrics. Keeping a minimum number of idle workers avoids that activation gap but consumes resources even when there is nothing to process.
KEDA addresses both cases by querying external sources for demand and coordinating with Kubernetes autoscaling. For a queue consumer, the queue depth can signal that more workers are needed, while an empty queue can eventually allow the Deployment to scale to zero. The consumer remains responsible for processing and acknowledging messages, and the source remains responsible for retaining them.
How KEDA Works
KEDA watches its custom resources and uses configured scalers to query event sources. A ScaledObject associates one or more triggers with a target workload. KEDA’s operator handles activation between zero and one replica. When the workload is active, KEDA provides external metrics through its metrics API and manages a Kubernetes HPA for scaling between one and the configured maximum. This extends rather than replaces the native scaling loop; see the Kubernetes HPA documentation for how the HPA calculates replica counts from metrics.
The distinction between pod scaling and node scaling matters. KEDA can request more pods, but it does not provision Kubernetes nodes. If the cluster has no room for those pods, a cluster autoscaler or a managed node-provisioning system must add capacity separately.
| Concern | Kubernetes HPA | KEDA | Cluster autoscaler |
|---|---|---|---|
| What it scales | Pods for a workload | Pods for a workload through a generated HPA, with activation from zero | Worker nodes or node groups |
| Typical input | CPU, memory, or configured metrics | Queue depth, stream lag, cloud-service metrics, or other scaler signals | Pending pods and node capacity |
| Scale from zero | Not from CPU or memory alone | Supported when a trigger can detect activity | Does not create workload replicas |
| Main configuration | HorizontalPodAutoscaler |
ScaledObject or ScaledJob, triggers, and optional authentication |
Cluster or provider-specific node group configuration |
| Responsibility boundary | Chooses replica count from metrics | Connects event-source demand to Kubernetes scaling | Supplies capacity when pods cannot be scheduled |
| Common use | Long-running services with a useful resource or metric target | Event consumers, queue workers, and scheduled or externally signaled work | Clusters whose requested pods exceed available node resources |
An event trigger does not automatically create a one-message-per-pod relationship. The metric, target threshold, processing rate, and workload’s concurrency determine how Kubernetes interprets demand. Set them using observed throughput and acceptable processing delay rather than assuming that each additional message requires another replica.
Components and Key Concepts
- Scaler: A source-specific adapter that obtains a metric or activation signal. KEDA supports integrations for systems such as queues, streams, cloud services, and Prometheus, as well as extension mechanisms for custom sources.
ScaledObject: The resource for scaling an existing Deployment, StatefulSet, or other supported target. It holds the target reference, minimum and maximum replica counts, polling and cooldown settings, and one or more triggers.ScaledJob: A resource for workloads where KEDA creates Kubernetes Jobs in response to event demand, rather than changing the replica count of one long-running worker Deployment.- KEDA operator and metrics server: The operator watches KEDA resources and handles activation and HPA lifecycle. The metrics server makes external metrics available to the HPA for active scaling.
- Authentication: Credentials and identity settings let a scaler read a protected source. Use Kubernetes Secrets or supported workload identity mechanisms instead of committing credentials in a manifest.
- CloudEvents: Some systems standardize event envelopes using the CloudEvents specification. KEDA does not require CloudEvents or inspect every event body; a scaler reads the source-specific signal configured for that integration.
The KEDA project site provides project information and deployment options. Before selecting a scaler, check its documentation for the exact metric semantics, authentication requirements, and supported resource types.
Real-World Use Cases
Queue-backed background work is a common fit. An order-processing or image-conversion worker can add replicas when its queue grows and scale down after the backlog is handled. Configure a maximum based on broker limits and downstream capacity, not only Kubernetes capacity.
Stream consumers can scale from lag or another source-specific signal. Partition count often limits useful parallelism: adding replicas beyond the number of independently consumable partitions may add idle consumers rather than throughput. Also consider ordering guarantees and rebalance costs.
Intermittent workloads may use a source that signals activity only during certain hours or in response to external demand. A ScaledJob can be appropriate when each unit of work should run as a Job with its own completion and retry behavior.
In each case, scaling improves how quickly available capacity follows work; it does not repair slow handlers, poison messages, broker outages, or inadequate downstream services. Pair scaling with queue-age alerts, retry limits, dead-letter handling, and application-level idempotency.
Getting Started
Install KEDA into a cluster where you have permission to create its custom resources and cluster-level components. The following Helm commands use the official chart repository:
helm repo add kedacore https://kedacore.github.io/charts
helm repo update
helm install keda kedacore/keda --namespace keda --create-namespace
kubectl rollout status deployment/keda-operator -n keda
The next example assumes a Deployment named orders-worker, a RabbitMQ broker reachable in the default namespace, and an orders queue. Store an authenticated broker connection string in a Kubernetes Secret before applying the manifest; do not commit credentials in source control. The TriggerAuthentication reads that connection string, and the ScaledObject uses queue length as its scaling signal.
Create the Secret from an environment variable that you set outside source control:
kubectl create secret generic rabbitmq-connection \
--namespace default \
--from-literal=host="$RABBITMQ_HOST"
RABBITMQ_HOST should contain the full AMQP connection URL for the broker and queue namespace used by your application.
apiVersion: keda.sh/v1alpha1
kind: TriggerAuthentication
metadata:
name: rabbitmq-auth
namespace: default
spec:
secretTargetRef:
- parameter: host
name: rabbitmq-connection
key: host
---
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: orders-worker
namespace: default
spec:
scaleTargetRef:
name: orders-worker
pollingInterval: 30
cooldownPeriod: 120
minReplicaCount: 0
maxReplicaCount: 20
triggers:
- type: rabbitmq
metadata:
protocol: amqp
queueName: orders
mode: QueueLength
value: "20"
authenticationRef:
name: rabbitmq-auth
Apply the manifest after the Deployment and broker are available:
kubectl apply -f orders-worker-scaling.yaml
kubectl get scaledobject,hpa -n default
kubectl describe scaledobject orders-worker -n default
kubectl logs deployment/keda-operator -n keda
Publish enough test messages to cross the configured threshold, then watch the Deployment’s replica count and queue depth. After draining the queue, allow the cooldown period before expecting a scale-down to zero. Check the KEDA operator logs and the ScaledObject conditions if the trigger cannot be read or the target does not scale. Also confirm that the worker’s Deployment resource requests, consumer concurrency, broker limits, and cluster capacity permit the requested replica count.
Common Misconceptions
KEDA replaces the HPA. It does not. For ScaledObject workloads, KEDA creates and manages an HPA so Kubernetes can perform active replica calculations; KEDA adds event-source metrics and handles activation from zero.
KEDA reads and processes each event. It generally polls or queries a metric through a scaler. The application still owns message consumption, acknowledgements, retries, and business logic.
Scale-to-zero is instant and free. The scaler must detect activity, the pod must be scheduled and become ready, and the application must connect to its source. Polling and cold-start time affect latency. A scaled-down pod count also does not guarantee that the underlying cluster or external services stop incurring costs.
Related Articles
- Kubernetes Architecture Explained
- Container Orchestration Best Practices
- Event-Driven Microservices
- Infrastructure Monitoring with Prometheus
Changelog
- Initial publication; configuration and behavior checked against the official KEDA and Kubernetes documentation.

