Server
Autoscaling
Autoscaling is the platform adding or removing copies of your app based on a metric it watches, such as CPU, queue depth or requests in flight. It reacts after load shows up, not before.
How it is measured
You read it from the replica count over time next to the metric that drives it. Plot instance count and average CPU on the same chart at one-minute resolution, then look at the gap between the load rising and the new instance taking traffic.
The settings that matter are the target (say 60 percent CPU), the minimum and maximum replica counts, and the cooldown in seconds. Most 'autoscaling did not help' stories come from a boot time of 90 seconds against a spike that lasted 40.
Worked example
An Astro storefront on a container platform runs 2 replicas, scaling to 8 at 65 percent CPU. A newsletter goes out at 09:00 and requests jump from 40 to 520 per second within a minute. The platform notices at 09:01, starts three more replicas, and they are serving by 09:02:30.
For those 150 seconds the first two replicas sat above 95 percent CPU and p95 climbed from 280 ms to 4.1 seconds. By 09:10 the fleet had 7 replicas and p95 was back to 310 ms. The scaling worked, but too slowly for the first wave.
How it differs
Autoscaling decides how many copies run, automatically. Horizontal scaling is the shape of the change, adding copies, and you can do it by hand. Autoscaling excludes manual intervention and usually the pre-warming of capacity. Horizontal scaling excludes bigger single machines but says nothing about who triggers it.
Common errors
Setting the maximum so low the fleet pins at the ceiling during the spike. Scaling on CPU for a service that waits on a database. Ignoring boot and warm-up time. Letting scale-in kill instances that still hold requests. Assuming it fixes a slow query when more copies only add more connections to the same database.
In practice
Raise the minimum replica count before a known launch rather than relying on reaction time, and load-test once with a 5x jump to see how long the gap lasts. Check that your database connection limit can survive the maximum replica count multiplied by the pool size.