What is MC-LAG
MC-LAG (Multi-Chassis Link Aggregation Group) is a mechanism for implementing link aggregation across two physically separate devices, instead of just across ports on a single switch. It keeps every benefit of ordinary link aggregation — bandwidth pooling, load balancing across member links — while adding something standard LAG can't provide on its own: device-level redundancy, since the two aggregated ports live on entirely separate physical switches.
MC-LAG does this through horizontal virtualization: two physical switches are virtualized into what looks, from the outside, like a single logical device. A server or downstream switch connecting into an MC-LAG pair simply sees one aggregation partner, unaware that its two "member ports" actually terminate on two different chassis. Because the MC-LAG group presents as a single node externally, it avoids the loop risk that naive dual-homed connections would otherwise create, while still making full use of every link — none of them sit idle the way a blocked backup path under STP would.
How MC-LAG Works
An MC-LAG deployment is built from a defined set of link types, each with a specific job. The two switches participating in the aggregation are MC-LAG peers, one taking the Active role and one the Standby role — a distinction that matters for the control-plane TCP session between them, but not for data forwarding, where both peers independently forward traffic and hold equal status. The peer-link is a direct physical connection between the two peers, used to carry traffic across when a downstream link fails. The keep-alive link is a heartbeat connection — commonly sharing the same physical link as the peer-link — that carries the control protocol traffic, synchronizes tables, and establishes the peer relationship. And an optional DAD link (Dual-Active Detection) gives the pair a way to detect a true dual-active condition — both peers acting as Active simultaneously — specifically for scenarios where peer-link and keep-alive share the same physical connection; if the keep-alive link goes down, the system automatically shuts down every interface on the standby node except logical interfaces, management ports, and the peer-link itself.
The control plane runs on ICCP (Inter-Chassis Communication Protocol), a lightweight protocol defined in RFC 7275, using TCP port 8888 between peers. Beyond the initial neighbor establishment, ICCP continuously synchronizes several categories of information: the system MAC address (so LACP packets sent to a downstream device carry an identical system ID from both peers, making cross-device aggregation actually work), MC-LAG member port configuration and status (for consistency checking and coordinated failure handling), and ARP and FDB table entries (so both peers make consistent forwarding decisions). Heartbeat packets are sent every second by default; if 15 consecutive heartbeats go unanswered, the session is declared timed out and the ICCP connection is considered broken.
To catch configuration drift before it causes a problem, MC-LAG runs a consistency check comparing settings like local/peer IP symmetry and member port configuration across the two peers, in one of three modes: idle (report only, take no action — the default), default (shut down the failing member port on both peers), or graceful (shut down the failing member port only on the standby side, keeping active service uninterrupted).
Why MC-LAG is Beneficial
Device-level redundancy, not just link-level: Because the two aggregated ports terminate on separate physical switches, MC-LAG survives an entire device failure — something a single switch's standard LAG configuration fundamentally cannot do.
No idle backup links: Presenting as a single logical device externally means MC-LAG avoids the loop-prevention trade-off of blocking redundant paths — every member link stays actively forwarding traffic.
Fast, coordinated failure handling: Continuous ARP, FDB, and status synchronization over ICCP means the standby peer already has the state it needs to take over forwarding the moment something changes, rather than rebuilding tables from scratch.
Configuration drift gets caught, not discovered in production: The consistency-check mechanism actively compares peer configuration and can act automatically — in idle, default, or graceful mode — depending on how much automatic intervention an operator wants.
Scales beyond a single pair: Two-level (cascaded) MC-LAG topologies let a network expand well past what one MC-LAG pair alone could support, while keeping the same redundancy model throughout.