A distributed denial-of-service attack does not need a vulnerability. It needs more traffic, more packets or more expensive requests than your weakest component can handle, sent from enough sources that blocking one of them does nothing. Booter services rent hundreds of gigabits per second for little money. Reflection and amplification turn small spoofed requests into large responses aimed at you, and residential-proxy botnets make an HTTP flood look like thousands of ordinary browsers.

This article explains the architecture that actually withstands that, from first principles. It covers why a shared anycast edge inverts the economics, how detection works on flow telemetry, what each mitigation layer can and cannot stop, when scrubbing centres and BGP diversion are still needed, and why most attacks against protected sites succeed by going around the protection rather than through it. You get a worked attack timeline, detection code, AWS and Google Cloud commands, failure modes and a checklist. Figures in examples are illustrative.

Three attack classes and economic denial of service

Attacks fall into three classes, and each one exhausts a different resource. Mixing them up is the most common design mistake, because a control that stops one class is useless against the others.

ClassResource exhaustedTypical vectorsWhere it must be stopped
Volumetric (L3)Link and router capacity, measured in bits per secondUDP reflection: DNS, NTP, memcached, CLDAP, SSDP; raw UDP floodsUpstream of your links: the edge or a scrubbing provider
Protocol (L4)Connection state, measured in packets per secondSYN floods, ACK floods, fragment floods, TCP state exhaustionA stateless or SYN-cookie front that terminates TCP
Application (L7)CPU, database, cache, measured in requests per secondHTTP GET/POST floods, cache-busting query strings, login and search abuseA proxy that sees HTTP: WAF, rate limits, bot scoring

There is a fourth outcome that is not a class but a result: economic denial of service. If your answer to load is autoscaling, an attack that fails technically can still succeed financially, because you pay for every instance, every gigabyte of egress and every serverless invocation it triggers. A design that keeps up but quadruples your bill has only moved the damage.

The defender only wins by pooling capacity larger than any single attack and sharing its cost across many customers. That is what a cloud DDoS architecture is.

The architecture, piece by piece

Cloud DDoS protection: absorb at the edge, detect from telemetry, mitigate per layer, shield the originAttack trafficbotnets, reflectors, L7Anycast edge100s of PoPs, one prefixFlow telemetrysFlow, NetFlow, IPFIXDetectiondeviation vs baselineTbps insamplesVolumetric mitigationedge ACLs, drop amplificationProtocol mitigationSYN cookies, state limitsL7 mitigationWAF, rate limits, challengesL3 floodL4L7policyScrubbing centersBGP diversion, GRE returnManaged protectionShield, Cloud Armor, AzureOrigin shieldingallow-list, hidden IPs, private linkdivert prefixclean traffic onlyOperations: runbooks, game days, baseline reviews, cost-of-attack monitoringFloods die where they land; only policy-checked requests reach the origin.
Defence in depth. Anycast spreads the flood, telemetry-driven detection selects per-layer mitigation, scrubbing covers non-HTTP prefixes, and the origin only ever accepts traffic from the protection layer.

Anycast edge. The same IP prefix is announced over BGP from hundreds of points of presence (PoPs), and each packet is routed to the topologically nearest one. Botnet sources are scattered, so their traffic is automatically partitioned: a 2 Tbps flood spread over 250 PoPs is roughly 8 Gbps per PoP, which each site can absorb and filter. Anycast also removes the aiming point, because there is no single data centre to concentrate on, though a regionally concentrated botnet loads nearby PoPs more than average.

Stateless edge filtering. The cheapest drop happens first. Packets from source ports of known reflectors (UDP 53 responses you never asked for, UDP 123, UDP 11211) aimed at a web prefix have no legitimate purpose and are dropped at line rate, typically in hardware ACLs or in XDP/eBPF programs that run before the kernel network stack allocates anything.

Detection. Routers export sampled flow records into a pipeline that keeps per-destination baselines (bits and packets per second, protocol mix, SYN-to-ACK ratio, source diversity) with daily and weekly seasonality. Deviation triggers mitigation in seconds, because links saturate faster than humans react.

Scrubbing. Non-HTTP services (game servers, VoIP, raw TCP) cannot sit behind an HTTP proxy, so a provider announces your prefix, scrubs it and returns clean traffic over GRE tunnels or private interconnects.

Detection from flow telemetry

Detection does not need machine learning to catch the bulk of attacks. A 50x jump in UDP source port 123 toward one prefix is NTP amplification by definition. What it needs is a baseline per destination and per feature, a robust deviation measure, and a minimum absolute floor so that a quiet service does not alert when traffic goes from 10 to 40 packets per second. The sketch below keeps an exponentially weighted mean and variance per key and flags sustained deviation; production systems add seasonality by keeping one baseline per hour-of-week bucket.

import math
from collections import defaultdict

ALPHA, Z_ALERT = 0.05, 6.0       # EWMA weight, deviations that count as attack
MIN_PPS, SUSTAIN = 50_000, 3      # absolute floor, consecutive intervals

class Baseline:
    def __init__(self):
        self.mean, self.var, self.hot = 0.0, 1.0, 0

    def observe(self, x):
        z = (x - self.mean) / math.sqrt(self.var + 1e-9)
        attack = x > MIN_PPS and z > Z_ALERT
        self.hot = self.hot + 1 if attack else 0
        if not attack:                      # never learn from attacks
            d = x - self.mean
            self.mean += ALPHA * d
            self.var = (1 - ALPHA) * (self.var + ALPHA * d * d)
        return self.hot >= SUSTAIN, z

baselines = defaultdict(Baseline)

def on_interval(flows, sample_rate=1000):
    """flows: (dst_prefix, proto, src_port, packets) from sampled sFlow/IPFIX."""
    agg = defaultdict(int)
    for dst, proto, sport, pkts in flows:
        agg[(dst, proto, sport if proto == "udp" else None)] += pkts * sample_rate
    for key, pps in agg.items():
        fire, z = baselines[key].observe(pps)
        if fire:
            yield {"key": key, "pps": pps, "z": round(z, 1)}

Two details matter. The baseline must not learn during an attack, or a slow ramp teaches it the flood is normal. And sampled flows must be scaled by the sampling rate before comparison with thresholds. For L7, watch cache hit ratio and origin latency too: a cache-busting flood can melt the origin while request rate barely moves.

Mitigation by layer

Volumetric. Reflection floods are dropped by edge ACLs and rate shapers in the PoP that received them. The defence is capacity plus cheap filtering; nothing behind the edge can help once a link is full.

Protocol. A SYN flood tries to fill the server's half-open connection table. SYN cookies encode the connection parameters into the initial sequence number of the SYN-ACK, so the server keeps no state until the client's ACK proves it received the reply. Spoofed sources never complete the handshake, so they cost the edge one stateless computation each. Because the edge terminates TCP and TLS itself, only established connections ever travel inward.

Application. This is the hard fight, because each request is individually valid. The tools are per-key rate limits (per IP, per session, per API key, per path), bot scoring on TLS and HTTP fingerprints, and graduated challenges: invisible JavaScript or proof-of-work first, interactive challenges last. In Google Cloud Armor, a rate-based ban on an expensive path looks like this:

gcloud compute security-policies rules create 1000 \
  --security-policy=edge-policy \
  --expression="request.path.matches('/search')" \
  --action=rate-based-ban \
  --rate-limit-threshold-count=600 \
  --rate-limit-threshold-interval-sec=60 \
  --ban-duration-sec=300 \
  --conform-action=allow \
  --exceed-action=deny-429 \
  --enforce-on-key=IP

Per-IP limits underperform against residential-proxy botnets; keying on a session, API key or fingerprint cluster catches what they miss.

Always-on versus on-demand. Always-on means traffic permanently flows through the protection layer, so mitigation is a policy decision measured in seconds. On-demand means BGP diversion to a scrubbing provider after detection, which commonly takes minutes. That suits whole network ranges you do not want proxied, but for user-facing services the outage usually costs more than always-on.

Origin shielding

Most successful attacks on protected properties do not beat the edge. They find the origin and hit it directly. Origin addresses leak through historical DNS records from before the CDN was added, through mail servers on the same host that put the address in message headers, through verbose error pages and redirects, and through forgotten subdomains such as a staging API that was never put behind the proxy. Certificate transparency logs leak those forgotten hostnames, which attackers then resolve, but they do not contain IP addresses themselves.

Shielding therefore has three parts. Re-address the origin when you introduce protection, because the old address is already in DNS history databases. Accept traffic only from the protection layer, or better, make the origin reachable only over private connectivity. And put every hostname behind the same front door. On AWS, CloudFront publishes a managed prefix list for its origin-facing addresses, so a security group can admit only CloudFront:

PL=$(aws ec2 describe-managed-prefix-lists \
  --filters Name=prefix-list-name,Values=com.amazonaws.global.cloudfront.origin-facing \
  --query 'PrefixLists[0].PrefixListId' --output text)

aws ec2 authorize-security-group-ingress --group-id sg-0abc123 \
  --ip-permissions "IpProtocol=tcp,FromPort=443,ToPort=443,PrefixListIds=[{PrefixListId=$PL}]"

A network allow-list proves the packet came from the provider's ranges, not that it came from your distribution; another customer of the same CDN can point at your origin. Add a second check the attacker cannot forge: a secret header injected by the edge and verified at the origin, mutual TLS between edge and origin, or a private link that has no public address at all.

Worked example: one attack, four pivots

Consider an API serving 80,000 requests per second through an anycast CDN. The timeline below is illustrative but every step reflects a mechanism described above.

  1. 14:07. A booter sends 400 Gbps of NTP reflection toward the API prefix. Anycast spreads it over about 200 PoPs; the busiest sees under 10 Gbps. Edge ACLs already drop UDP source port 123 toward web prefixes, so the flood dies at ingress. The detector above fires on its third interval and logs an event. Origin graphs do not move.
  2. 14:20. The attacker pivots to a 40 million packet-per-second SYN flood from spoofed sources. Edge TCP termination answers with SYN cookies; no state accumulates and legitimate handshakes complete normally.
  3. 14:31. The real fight: 600,000 requests per second of GET /search?q=<random> from residential proxies, each unique so the cache cannot help. Per-IP limits barely trigger, but the requests share a TLS fingerprint that does not match their claimed browser. Bot scoring flags the cluster, a JavaScript challenge goes on /search, and the flood collapses because the bots cannot execute it. Real users see a brief latency bump.
  4. 14:55. The attacker hits an origin address recovered from a years-old DNS history record. The packets die against a security group that admits only the CDN's prefix list, because the origin was re-addressed when shielding was introduced.
  5. Afterwards: a tighter search rate limit, captures added to the game-day library, and a bill check.

Managed services compared

Every major cloud provides basic network-layer protection at no extra charge and sells a premium tier. Tier names change, so confirm them against current documentation before you design around one.

ServiceBaseline tierPremium tier addsNotes
AWS ShieldShield Standard, automatic for all customersShield Advanced: response team access, cost protection, automatic L7 mitigation with AWS WAFStrongest when traffic enters via CloudFront, Global Accelerator or Route 53
Google Cloud ArmorStandard, pay per use: DDoS and WAF on the external load balancerEnterprise (formerly Managed Protection Plus): Adaptive Protection, threat intelligence, DDoS cost protectionAdaptive Protection learns baselines and proposes L7 rules
Azure DDoS ProtectionBasic infrastructure protectionIP Protection (per public IP) or Network Protection (per virtual network), which adds rapid response and cost protectionNetwork Protection is the tier with response support
Independent edgesVariesNetwork-layer prefix protection over BGP, plus L7 proxyingUseful for multi-cloud or on-premises prefixes

Cost protection, which credits scaling charges caused by a documented attack, is the only contractual answer to economic denial of service; check its conditions before you rely on it.

Failure modes

FailureWhat happensPrevention
Origin exposedAttacker bypasses the edge and saturates the origin directlyRe-address, allow-list provider ranges, verify an edge secret, scan your own DNS history
Unprotected subdomainA staging or legacy API host has no proxy and takes the whole backend downInventory hostnames from CT logs and DNS; one front door for all
Autoscaling absorbs the attackService stays up, bill multipliesScaling caps, budget alerts, cost-protection enrolment, rate limits before the scaler
Cache-busting L7 floodRequest rate looks fine, origin meltsAlert on cache hit ratio and origin latency; normalise or ignore unknown query parameters
Over-aggressive rulesMitigation blocks real users, completing the attack for themPreview mode first, challenge before block, per-path scope
Shared fate with stateful firewallA firewall's connection table fills before the server'sKeep stateful devices behind the stateless edge, never in front of it

Trade-offs

Proxying everything versus network-layer protection. An L7 proxy sees requests and can stop application floods, but it terminates TLS, so it holds your certificates and sees plaintext. Network-layer protection preserves end-to-end encryption but cannot tell a valid login from a credential-stuffing request.

Single provider versus multi-CDN. One provider gives one policy plane and simpler origin lock-down; two providers add resilience against the provider itself failing, at the cost of keeping two rule sets in sync and allow-listing two sets of ranges.

Challenges versus friction. Every challenge costs real users some latency and breaks some API clients. Scope challenges to the attacked path and exempt authenticated machine clients by key.

What to do next

  1. List every public hostname and address you own, including staging, and confirm each sits behind the edge.
  2. Check whether your origin's current address appears in public DNS history; if it does, re-address it.
  3. Restrict origin ingress to the provider's published ranges and add an edge-to-origin secret or mutual TLS.
  4. Write rate limits for your three most expensive paths and run them in preview mode for a week.
  5. Put scaling caps and budget alerts on everything an attack could scale, and enrol in cost protection if you buy a premium tier.
  6. Add alerts on cache hit ratio and origin latency, not just request rate.
  7. Write the runbook: who declares an incident, who can deploy an emergency rule, how to reach the provider's response team.
  8. Run a game day against a staging stack with permission from your provider, and replay captures from real incidents.
  9. Keep learning: AWS Shield, Google Cloud Armor, cloud WAF design, Alibaba Anti-DDoS and cloud load balancers.
Key takeaway: DDoS defence is an architecture, not a product. Pool capacity at an anycast edge so floods die where they land, detect from per-destination baselines in seconds, match each attack class to the layer that can stop it, and make sure the origin accepts traffic from nowhere else. Then cap what an attack can cost you, because staying up while the bill explodes is still a loss.