Control Planes, Part 2: When to Stop Configuring Autoscalers and Build Your Own
This is the second in my series on control plane architecture -  read the first one on control plane 2026-10-1 23:56:56 Author: hackernoon.com(查看原文) 阅读量:1 收藏

This is the second in my series on control plane architecture -  read the first one on control plane fundamentals here

In my last piece, I broke down the control plane as the "brain" of a distributed system. It’s the architectural component responsible for scaling, load balancing, reliability, and fault tolerance behind the scenes. Scaling got one paragraph. It deserves a lot more than that.

Here's the uncomfortable truth about scaling: almost every team eventually asks "should we build this ourselves, or just configure what the cloud gives us for free?" The answer isn't the same for everyone, and getting it wrong in either direction is expensive. Either you're reinventing a wheel AWS already built, or you're bolting a general-purpose autoscaler onto a problem it was never designed to solve.

Let's go through the actual options.

Cloud-Native Autoscaling: The Default, and Why It Works Until It Doesn't

Most cloud providers ship autoscaling as a first-class feature, and for a huge share of workloads, it's genuinely the right call. Let’s go through this with some popularly used scaling solutions.

AWS Auto Scaling Groups

An Auto Scaling Group (ASG) manages a fleet of EC2 instances against targets such as CPU utilization, request count per target, or a custom CloudWatch metric. You define:

  • A minimum, maximum, and desired instance count
  • A scaling policy e.g. target tracking, step scaling, or simple scaling
  • Health checks so unhealthy instances get replaced automatically

Target tracking is the one most teams reach for first: you pick a metric (say, 60% average CPU) and the ASG adds or removes instances to hold that target. It's simple, it's declarative, and it works well when your workload's resource usage correlates cleanly with your scaling needs.

Where it gets shaky: request latency doesn't always track CPU. A service can be CPU-idle and still falling over because of connection limits, downstream dependency slowness, or memory pressure. If your bottleneck isn't the metric you're scaling on, the ASG will confidently do nothing while you burn.

Beyond EC2: Autoscaling for Managed Databases

ASGs solve scaling for fleet management of the hosts in the fleet running your code. Le’ts go through another example of what scaling looks like when we’re not managing a fleet but rather throughput for a distributed cloud-based application. Take for instance managed database services like DynamoDB and Amazon Keyspaces (Cassandra-compatible). These need a fundamentally different kind of autoscaling that we discuss below.

DynamoDB and Amazon Keyspaces expose their scaling mechanism through Application Auto Scaling, which works conceptually like target tracking on an ASG but against a different metric entirely: consumed read/write capacity units as a percentage of provisioned capacity. You set a target utilization (commonly 70%), and Application Auto Scaling adjusts your table's (or index's) provisioned throughput up or down to hold that target. A few things make this meaningfully different from EC2-style scaling:

  • It's reactive, not predictive, and it's specifically tuned to avoid over-reacting to short spikes. It scales up faster than it scales down, similar in spirit to the "conservative scale-down, aggressive scale-up" pattern from EC2, but baked into the service rather than something you configure yourself.
  • It operates per-table (and per-index), not per-fleet. A table with a hot secondary index can have its index scale independently of the base table's capacity.

This is worth sitting with for a second: normally, scaling a distributed database system typically means the kind of partition-level re-sharding I describe in the bespoke-systems section below. A managed service like DynamoDB or Amazon Keyspaces absorbs that complexity into the service itself, exposing only a capacity-and-target interface at the table level. You're still benefiting from partition-aware scaling under the hood, you just don't have to build or operate it.

Taking a step back and looking at the broader pattern shows that for managed, stateful services, "autoscaling" usually means target tracking against a throughput or capacity metric specific to that service, not a generic infrastructure metric like CPU. Because the thing being scaled isn't a host, it's a logical unit of provisioned capacity that the service itself maps onto physical resources on your behalf.

Kubernetes: Three Different Scalers, Three Different Jobs

Kubernetes doesn't have one autoscaler, it has three, and conflating them is one of the more common mistakes I see.

  • Horizontal Pod Autoscaler (HPA) adds or removes pod replicas based on observed metrics (CPU, memory, or custom/external metrics via the metrics API). This is your "more copies of the same thing" lever.
  • Vertical Pod Autoscaler (VPA) adjusts the resource requests and limits of existing pods, giving a pod more CPU or memory rather than spinning up more copies. Useful for workloads that don't horizontally shard well.
  • Cluster Autoscaler operates one level up: it adds or removes nodes from the cluster based on whether pods are pending (unschedulable) or nodes are underutilized.

These three operate independently and can fight each other if misconfigured. HPA scaling out pods while VPA is simultaneously trying to resize them underneath, for instance. A coherent scaling strategy on Kubernetes usually means picking HPA or VPA for a given workload, not both, and treating Cluster Autoscaler as the layer that reacts to whatever HPA/VPA already decided.

Configuring These Well: A Few Patterns That Actually Matter

  • Scale on the metric that reflects your actual bottleneck, not the easiest one to measure. If your service is I/O bound, CPU-based scaling will lag reality.
  • Set conservative scale-down, aggressive scale-up. Under-provisioning during a spike costs you availability; over-provisioning during a lull costs you money. Those aren't symmetric risks.
  • Use cooldown periods deliberately. Without them, you get flip flopping of scale up and scale down operations. Which could often cause more damage than doing nothing via cascading failures and long recovery times.
  • Combine metrics where you can. Target tracking on a single metric is a starting point, not an end state. Step scaling policies or custom metrics (queue depth, request latency, connection count) usually model reality better than CPU alone.

When Off-the-Shelf Autoscaling Isn't Enough

Varsha Ganesh's image-db9b48

Cloud-native autoscaling assumes a fairly generic shape: stateless-ish workloads where adding a copy of the same unit of work solves your scaling problem. A lot of systems don't fit that shape.

Distributed databases are the clearest example. Scaling a Cassandra or DynamoDB-style system isn't "add more instances", it's re-partitioning data that has already been sharded and replicated across a fleet, live, without losing availability or consistency guarantees. An ASG has no concept of a partition, let alone how to split one safely.

This is where teams end up building a bespoke scaling system. A purpose-built control plane component instead of a general-purpose one.

What a Bespoke Scaling System Actually Needs

If you're architecting one from scratch, there are a few components that show up again and again:

  1. A monitoring layer that understands your resource model. Generic infrastructure metrics (CPU, memory) aren't enough. You need visibility into the system you're actually scaling: partition throughput, queue depth, shard size, hot-key detection.
  2. A decision engine that translates "this resource is over/under its threshold" into a concrete action such as split this partition, merge these two shards, add a replica here. This is usually where the real complexity lives, because the decision has to account for the cost and risk of the action itself, not just the trigger. In addition, “cost” needs to be measured across multiple dimensions - such as time taken for the operation, the actual cost of the physical resource on which the new shard, partition will be placed, allowed reaction time etc.
  3. An execution layer that can carry out the action safely. Like with cost, safety has many dimensions such as idempotence, appropriate granularity, and the ability to halt or roll back mid-operation. Re-partitioning a live dataset that's serving production traffic is not a one-shot script; it's a state machine.
  4. Backpressure and rate limiting on the scaling system itself. A control plane component that decides to re-shard everything at once because ten partitions crossed their threshold simultaneously can do more damage than the problem it was solving. Bespoke scalers need to throttle their own actions.
  5. Observability into the scaler, not just the system it scales. When your scaling logic is custom, "why did it do that" needs to be answerable after the fact. To this end, the system needs to maintain and audit detailed logs, decision traces, and metrics on its own behavior.

Where Else This Shows Up

Distributed databases are the most typical example, but the same pattern i.e. a scaling unit with placement constraints, statefulness, or a cost profile that a single infrastructure metric can't capture, recurs across several other categories of systems:

  • Message queues and log-based systems. Scaling a Kafka cluster isn't as simple as "add brokers". It's rebalancing existing partitions across a larger broker set without violating replication placement or disrupting in-flight consumer offsets. Partition count is typically fixed at topic-creation time (the same constraint I covered in the load balancing piece), so tools like Cruise Control exist specifically because generic autoscaling has no model for which partition belongs on which broker, or what moving it costs in network I/O mid-rebalance.
  • Search and indexing systems. Elasticsearch and OpenSearch scale by adding shards and rebalancing them across nodes, but shard count per index is usually fixed at index-creation time too. Over-sharding wastes overhead, under-sharding creates a rebalancing wall later. Bespoke tooling here tends to focus on index lifecycle: rolling indices, deciding when to reindex with a different shard count, and moving hot recent-data shards to faster storage while aging data cools onto cheaper tiers.
  • Stateful ML serving and feature stores. Scaling decisions here are often driven by GPU memory footprint and model-loading time rather than request volume. A cold replica can take minutes to become useful, which makes predictive, pre-warmed scaling matter more than reactive target tracking.
  • Connection-oriented systems. Anything holding long-lived client connections - systems like chat backends, multiplayer game servers, live video, distributed load balancers, WebSocket gateways etc. can't scale by just adding capacity, because existing connections are pinned to specific hosts. Scaling down means gracefully draining and migrating live sessions, not simply terminating instances, which requires custom connection-draining logic most generic autoscalers simply don't model.
  • Batch and workflow orchestration. Systems like Airflow or Spark scale against job queue depth and the resource shape of individual jobs. A job needing 64GB won't fit on a node sized for the median job,  which makes the scaling decision closer to bin-packing than target tracking on a single metric.

The common thread across all of these, including the database case above: bespoke scaling shows up wherever the unit being scaled carries placement constraints, state, or a cost profile a single infrastructure metric can't capture.

The Real Question: Build vs. Configure

Before building any of this, it's worth being honest about which category your system actually falls into:

  • If your scaling unit is stateless and interchangeable e.g. web servers, workers pulling from a queue, most application-tier services,  cloud-native autoscaling (ASG, HPA, Cluster Autoscaler) will very likely get you further than a custom system, faster and with less operational risk.
  • If your scaling unit is stateful, partitioned, or has placement constraints e.g.  a distributed database, a sharded cache, anything where "which specific piece of data or state moves where" matters, you're probably looking at a bespoke system, because generic autoscalers have no model for that.

The mistake I've seen most often isn't picking the wrong tool. It's not asking the question at all, and defaulting to whatever the cloud provider offers because it's already there. That works fine for a while. It stops working exactly at the point your scaling unit stops being interchangeable, and by then the migration is a lot more painful than the upfront design decision would have been.

Coming Up Next

In the next piece in this series, I'll dig into load balancing and routing patterns,  including how systems handle the "hot partition" problem I touched on above, and what re-partitioning for load distribution actually looks like under the hood.


文章来源: https://hackernoon.com/control-planes-part-2-when-to-stop-configuring-autoscalers-and-build-your-own?source=rss
如有侵权请联系:admin#unsafe.sh