A distributed denial-of-service attack does not need a vulnerability. It needs more traffic, more packets or more expensive requests than your weakest component can handle, sent from enough sources that blocking one of them does nothing. Booter services rent hundreds of gigabits per second for little money. Reflection and amplification turn small spoofed requests into large responses aimed at you, and residential-proxy botnets make an HTTP flood look like thousands of ordinary browsers.
This article explains the architecture that actually withstands that, from first principles. It covers why a shared anycast edge inverts the economics, how detection works on flow telemetry, what each mitigation layer can and cannot stop, when scrubbing centres and BGP diversion are still needed, and why most attacks against protected sites succeed by going around the protection rather than through it. You get a worked attack timeline, detection code, AWS and Google Cloud commands, failure modes and a checklist. Figures in examples are illustrative.
Three attack classes and economic denial of service
Attacks fall into three classes, and each one exhausts a different resource. Mixing them up is the most common design mistake, because a control that stops one class is useless against the others.
| Class | Resource exhausted | Typical vectors | Where it must be stopped |
|---|---|---|---|
| Volumetric (L3) | Link and router capacity, measured in bits per second | UDP reflection: DNS, NTP, memcached, CLDAP, SSDP; raw UDP floods | Upstream of your links: the edge or a scrubbing provider |
| Protocol (L4) | Connection state, measured in packets per second | SYN floods, ACK floods, fragment floods, TCP state exhaustion | A stateless or SYN-cookie front that terminates TCP |
| Application (L7) | CPU, database, cache, measured in requests per second | HTTP GET/POST floods, cache-busting query strings, login and search abuse | A proxy that sees HTTP: WAF, rate limits, bot scoring |
There is a fourth outcome that is not a class but a result: economic denial of service. If your answer to load is autoscaling, an attack that fails technically can still succeed financially, because you pay for every instance, every gigabyte of egress and every serverless invocation it triggers. A design that keeps up but quadruples your bill has only moved the damage.
The defender only wins by pooling capacity larger than any single attack and sharing its cost across many customers. That is what a cloud DDoS architecture is.
The architecture, piece by piece
Anycast edge. The same IP prefix is announced over BGP from hundreds of points of presence (PoPs), and each packet is routed to the topologically nearest one. Botnet sources are scattered, so their traffic is automatically partitioned: a 2 Tbps flood spread over 250 PoPs is roughly 8 Gbps per PoP, which each site can absorb and filter. Anycast also removes the aiming point, because there is no single data centre to concentrate on, though a regionally concentrated botnet loads nearby PoPs more than average.
Stateless edge filtering. The cheapest drop happens first. Packets from source ports of known reflectors (UDP 53 responses you never asked for, UDP 123, UDP 11211) aimed at a web prefix have no legitimate purpose and are dropped at line rate, typically in hardware ACLs or in XDP/eBPF programs that run before the kernel network stack allocates anything.
Detection. Routers export sampled flow records into a pipeline that keeps per-destination baselines (bits and packets per second, protocol mix, SYN-to-ACK ratio, source diversity) with daily and weekly seasonality. Deviation triggers mitigation in seconds, because links saturate faster than humans react.
Scrubbing. Non-HTTP services (game servers, VoIP, raw TCP) cannot sit behind an HTTP proxy, so a provider announces your prefix, scrubs it and returns clean traffic over GRE tunnels or private interconnects.
Detection from flow telemetry
Detection does not need machine learning to catch the bulk of attacks. A 50x jump in UDP source port 123 toward one prefix is NTP amplification by definition. What it needs is a baseline per destination and per feature, a robust deviation measure, and a minimum absolute floor so that a quiet service does not alert when traffic goes from 10 to 40 packets per second. The sketch below keeps an exponentially weighted mean and variance per key and flags sustained deviation; production systems add seasonality by keeping one baseline per hour-of-week bucket.
import math
from collections import defaultdict
ALPHA, Z_ALERT = 0.05, 6.0 # EWMA weight, deviations that count as attack
MIN_PPS, SUSTAIN = 50_000, 3 # absolute floor, consecutive intervals
class Baseline:
def __init__(self):
self.mean, self.var, self.hot = 0.0, 1.0, 0
def observe(self, x):
z = (x - self.mean) / math.sqrt(self.var + 1e-9)
attack = x > MIN_PPS and z > Z_ALERT
self.hot = self.hot + 1 if attack else 0
if not attack: # never learn from attacks
d = x - self.mean
self.mean += ALPHA * d
self.var = (1 - ALPHA) * (self.var + ALPHA * d * d)
return self.hot >= SUSTAIN, z
baselines = defaultdict(Baseline)
def on_interval(flows, sample_rate=1000):
"""flows: (dst_prefix, proto, src_port, packets) from sampled sFlow/IPFIX."""
agg = defaultdict(int)
for dst, proto, sport, pkts in flows:
agg[(dst, proto, sport if proto == "udp" else None)] += pkts * sample_rate
for key, pps in agg.items():
fire, z = baselines[key].observe(pps)
if fire:
yield {"key": key, "pps": pps, "z": round(z, 1)}Two details matter. The baseline must not learn during an attack, or a slow ramp teaches it the flood is normal. And sampled flows must be scaled by the sampling rate before comparison with thresholds. For L7, watch cache hit ratio and origin latency too: a cache-busting flood can melt the origin while request rate barely moves.
Mitigation by layer
Volumetric. Reflection floods are dropped by edge ACLs and rate shapers in the PoP that received them. The defence is capacity plus cheap filtering; nothing behind the edge can help once a link is full.
Protocol. A SYN flood tries to fill the server's half-open connection table. SYN cookies encode the connection parameters into the initial sequence number of the SYN-ACK, so the server keeps no state until the client's ACK proves it received the reply. Spoofed sources never complete the handshake, so they cost the edge one stateless computation each. Because the edge terminates TCP and TLS itself, only established connections ever travel inward.
Application. This is the hard fight, because each request is individually valid. The tools are per-key rate limits (per IP, per session, per API key, per path), bot scoring on TLS and HTTP fingerprints, and graduated challenges: invisible JavaScript or proof-of-work first, interactive challenges last. In Google Cloud Armor, a rate-based ban on an expensive path looks like this:
gcloud compute security-policies rules create 1000 \
--security-policy=edge-policy \
--expression="request.path.matches('/search')" \
--action=rate-based-ban \
--rate-limit-threshold-count=600 \
--rate-limit-threshold-interval-sec=60 \
--ban-duration-sec=300 \
--conform-action=allow \
--exceed-action=deny-429 \
--enforce-on-key=IPPer-IP limits underperform against residential-proxy botnets; keying on a session, API key or fingerprint cluster catches what they miss.
Always-on versus on-demand. Always-on means traffic permanently flows through the protection layer, so mitigation is a policy decision measured in seconds. On-demand means BGP diversion to a scrubbing provider after detection, which commonly takes minutes. That suits whole network ranges you do not want proxied, but for user-facing services the outage usually costs more than always-on.
Origin shielding
Most successful attacks on protected properties do not beat the edge. They find the origin and hit it directly. Origin addresses leak through historical DNS records from before the CDN was added, through mail servers on the same host that put the address in message headers, through verbose error pages and redirects, and through forgotten subdomains such as a staging API that was never put behind the proxy. Certificate transparency logs leak those forgotten hostnames, which attackers then resolve, but they do not contain IP addresses themselves.
Shielding therefore has three parts. Re-address the origin when you introduce protection, because the old address is already in DNS history databases. Accept traffic only from the protection layer, or better, make the origin reachable only over private connectivity. And put every hostname behind the same front door. On AWS, CloudFront publishes a managed prefix list for its origin-facing addresses, so a security group can admit only CloudFront:
PL=$(aws ec2 describe-managed-prefix-lists \
--filters Name=prefix-list-name,Values=com.amazonaws.global.cloudfront.origin-facing \
--query 'PrefixLists[0].PrefixListId' --output text)
aws ec2 authorize-security-group-ingress --group-id sg-0abc123 \
--ip-permissions "IpProtocol=tcp,FromPort=443,ToPort=443,PrefixListIds=[{PrefixListId=$PL}]"A network allow-list proves the packet came from the provider's ranges, not that it came from your distribution; another customer of the same CDN can point at your origin. Add a second check the attacker cannot forge: a secret header injected by the edge and verified at the origin, mutual TLS between edge and origin, or a private link that has no public address at all.
Worked example: one attack, four pivots
Consider an API serving 80,000 requests per second through an anycast CDN. The timeline below is illustrative but every step reflects a mechanism described above.
- 14:07. A booter sends 400 Gbps of NTP reflection toward the API prefix. Anycast spreads it over about 200 PoPs; the busiest sees under 10 Gbps. Edge ACLs already drop UDP source port 123 toward web prefixes, so the flood dies at ingress. The detector above fires on its third interval and logs an event. Origin graphs do not move.
- 14:20. The attacker pivots to a 40 million packet-per-second SYN flood from spoofed sources. Edge TCP termination answers with SYN cookies; no state accumulates and legitimate handshakes complete normally.
- 14:31. The real fight: 600,000 requests per second of
GET /search?q=<random>from residential proxies, each unique so the cache cannot help. Per-IP limits barely trigger, but the requests share a TLS fingerprint that does not match their claimed browser. Bot scoring flags the cluster, a JavaScript challenge goes on/search, and the flood collapses because the bots cannot execute it. Real users see a brief latency bump. - 14:55. The attacker hits an origin address recovered from a years-old DNS history record. The packets die against a security group that admits only the CDN's prefix list, because the origin was re-addressed when shielding was introduced.
- Afterwards: a tighter search rate limit, captures added to the game-day library, and a bill check.
Managed services compared
Every major cloud provides basic network-layer protection at no extra charge and sells a premium tier. Tier names change, so confirm them against current documentation before you design around one.
| Service | Baseline tier | Premium tier adds | Notes |
|---|---|---|---|
| AWS Shield | Shield Standard, automatic for all customers | Shield Advanced: response team access, cost protection, automatic L7 mitigation with AWS WAF | Strongest when traffic enters via CloudFront, Global Accelerator or Route 53 |
| Google Cloud Armor | Standard, pay per use: DDoS and WAF on the external load balancer | Enterprise (formerly Managed Protection Plus): Adaptive Protection, threat intelligence, DDoS cost protection | Adaptive Protection learns baselines and proposes L7 rules |
| Azure DDoS Protection | Basic infrastructure protection | IP Protection (per public IP) or Network Protection (per virtual network), which adds rapid response and cost protection | Network Protection is the tier with response support |
| Independent edges | Varies | Network-layer prefix protection over BGP, plus L7 proxying | Useful for multi-cloud or on-premises prefixes |
Cost protection, which credits scaling charges caused by a documented attack, is the only contractual answer to economic denial of service; check its conditions before you rely on it.
Failure modes
| Failure | What happens | Prevention |
|---|---|---|
| Origin exposed | Attacker bypasses the edge and saturates the origin directly | Re-address, allow-list provider ranges, verify an edge secret, scan your own DNS history |
| Unprotected subdomain | A staging or legacy API host has no proxy and takes the whole backend down | Inventory hostnames from CT logs and DNS; one front door for all |
| Autoscaling absorbs the attack | Service stays up, bill multiplies | Scaling caps, budget alerts, cost-protection enrolment, rate limits before the scaler |
| Cache-busting L7 flood | Request rate looks fine, origin melts | Alert on cache hit ratio and origin latency; normalise or ignore unknown query parameters |
| Over-aggressive rules | Mitigation blocks real users, completing the attack for them | Preview mode first, challenge before block, per-path scope |
| Shared fate with stateful firewall | A firewall's connection table fills before the server's | Keep stateful devices behind the stateless edge, never in front of it |
Trade-offs
Proxying everything versus network-layer protection. An L7 proxy sees requests and can stop application floods, but it terminates TLS, so it holds your certificates and sees plaintext. Network-layer protection preserves end-to-end encryption but cannot tell a valid login from a credential-stuffing request.
Single provider versus multi-CDN. One provider gives one policy plane and simpler origin lock-down; two providers add resilience against the provider itself failing, at the cost of keeping two rule sets in sync and allow-listing two sets of ranges.
Challenges versus friction. Every challenge costs real users some latency and breaks some API clients. Scope challenges to the attacked path and exempt authenticated machine clients by key.
What to do next
- List every public hostname and address you own, including staging, and confirm each sits behind the edge.
- Check whether your origin's current address appears in public DNS history; if it does, re-address it.
- Restrict origin ingress to the provider's published ranges and add an edge-to-origin secret or mutual TLS.
- Write rate limits for your three most expensive paths and run them in preview mode for a week.
- Put scaling caps and budget alerts on everything an attack could scale, and enrol in cost protection if you buy a premium tier.
- Add alerts on cache hit ratio and origin latency, not just request rate.
- Write the runbook: who declares an incident, who can deploy an emergency rule, how to reach the provider's response team.
- Run a game day against a staging stack with permission from your provider, and replay captures from real incidents.
- Keep learning: AWS Shield, Google Cloud Armor, cloud WAF design, Alibaba Anti-DDoS and cloud load balancers.