Route aggregation reduces table size and churn by advertising a summarized prefix instead of its components. In a single-tenant network that's mostly a scaling decision. In a multi-tenant network — where different tenants' address space rides the same aggregate, and different peers need different visibility into the tenant-level detail — aggregation stops being a scaling decision and becomes a policy decision, with real failure modes: suppressed more-specifics that were load-balancing traffic, AS-SET loss that breaks loop detection, and MED/path-attribute collisions once multiple original paths get folded into one route.
Why Aggregate at All
Every prefix a router advertises costs table space and processing on every neighbor that receives it, and every change to that prefix (a link flap, a tenant re-provisioning a subnet) generates an update that propagates outward. In a multi-tenant network — a provider or enterprise hosting many customers' or business units' address blocks behind a common edge — the number of distinct prefixes grows with tenant count, not with the network's actual topology. Left unaggregated, a large tenant base turns into a large, noisy routing table for every external peer, most of whom don't need visibility into individual tenant subnets — they need a route to the network, not a map of its internals.
Aggregation collapses a set of more-specific prefixes into a single summary route, typically configured via aggregate-address (or the vendor equivalent) with a summary-only option that suppresses advertisement of the components. The mechanism is simple. The consequences of applying it in a shared-infrastructure context are where the actual engineering is.
Failure Mode 1: Suppressing the More-Specifics That Were Doing Work
summary-only suppression assumes the more-specific routes were purely additive noise. In multi-tenant topologies they frequently aren't — a tenant with multiple upstream connections, or a subnet deliberately split across two edge routers for load distribution, relies on the more-specifics being visible so traffic actually takes the intended path. Once you aggregate and suppress, external peers see one route to the aggregate via whichever next-hop your BGP policy prefers, and any traffic engineering that depended on granular route visibility stops working — silently, because the aggregate is still reachable, just not through the path that was actually provisioned for that tenant.
The fix isn't "don't aggregate." It's identifying which more-specifics are structural (they encode an actual routing decision that needs to survive) versus incidental (they're an artifact of how the address space happens to be carved up), and aggregating only the second category, or using selective suppression (suppress-map matching only the prefixes safe to fold in) instead of blanket summary-only.
Failure Mode 2: AS-SET Loss and Loop Detection
When you aggregate routes that originated from different autonomous systems (common in multi-tenant environments where tenants bring their own ASN, or where the network sits between multiple upstream transit providers), the aggregate by default carries only the aggregating router's AS path — the individual AS paths of the components are dropped unless you explicitly configure the aggregate to carry an AS-SET.
This matters for loop prevention. BGP's core loop-detection mechanism is a router refusing to accept a route whose AS path already contains its own AS number. If aggregation drops the AS-SET, a route that should have been rejected as a loop can instead be accepted, because the aggregate's simplified path no longer shows the AS that would have triggered the rejection. The as-set keyword on the aggregate command preserves the set of contributing ASes in the path attribute specifically to keep this detection intact — omitting it is a common oversight because the aggregate still "works" under normal conditions; the failure only shows up when a loop condition that should have been caught, isn't.
Failure Mode 3: Attribute Collisions Across Folded Paths
An aggregate is built from multiple original routes, and those routes don't necessarily agree on MED (Multi-Exit Discriminator), origin type, or community values. BGP has defined behavior for this — MED and origin type are dropped from the aggregate unless explicitly configured to be preserved (as-set again plays a role here on some implementations, along with attribute-map configuration), and communities require explicit merging via route-map if you want tenant-specific tagging (used for downstream policy like blackholing or traffic-engineering communities) to survive aggregation.
If tenants are relying on communities for downstream policy — a not-uncommon multi-tenant pattern, where a tenant's traffic gets tagged so an upstream provider applies a specific policy to it — and the aggregation step silently drops those communities because nobody configured the merge, the downstream policy stops applying with no error, no log entry indicating why, and no obvious correlation to the aggregation change that caused it.
Failure Mode 4: Aggregation and Route Leaks in Multi-Tenant Peering
A subtler problem: aggregation can turn an otherwise-contained route leak into a much larger blast radius. If a single tenant's more-specific route gets leaked to an unintended peer, the damage is scoped to that tenant's traffic. If that same leak happens after aggregation — the leaked route is the aggregate, not the tenant's individual prefix — the blast radius is every tenant whose address space falls inside that aggregate's CIDR boundary, even tenants who had no involvement in whatever misconfiguration caused the leak.
This is a reason to be deliberate about where in the topology aggregation happens relative to policy enforcement points. Aggregating too early (before tenant-specific route filters and peer-specific policy are fully applied) collapses the ability to reason about a single tenant's blast radius independently of the others sharing the aggregate.
Aggregation Is a Set of Boundary Decisions, Not One
Aggregation is not a single decision made once at the network edge. It's a set of per-boundary decisions: which prefixes are safe to fold, which need AS-SET preservation because multiple origin ASes are involved, which attributes (MED, communities) need explicit merge configuration because downstream policy depends on them, and where in the topology aggregation happens relative to per-tenant route filtering so a leak stays scoped. Getting this wrong doesn't usually produce an outage — aggregated routes are still reachable. It produces traffic taking the wrong path, policy silently not applying, or a route leak affecting more tenants than the fault that caused it, all of which are considerably harder to diagnose after the fact than a routing table that's a bit larger than optimal.