Glossary

MC-LAG

Multi-Chassis Link Aggregation Group

What is MC-LAG

MC-LAG (Multi-Chassis Link Aggregation Group) is a mechanism for implementing link aggregation across two physically separate devices, instead of just across ports on a single switch. It keeps every benefit of ordinary link aggregation — bandwidth pooling, load balancing across member links — while adding something standard LAG can't provide on its own: device-level redundancy, since the two aggregated ports live on entirely separate physical switches.

MC-LAG does this through horizontal virtualization: two physical switches are virtualized into what looks, from the outside, like a single logical device. A server or downstream switch connecting into an MC-LAG pair simply sees one aggregation partner, unaware that its two "member ports" actually terminate on two different chassis. Because the MC-LAG group presents as a single node externally, it avoids the loop risk that naive dual-homed connections would otherwise create, while still making full use of every link — none of them sit idle the way a blocked backup path under STP would.

How MC-LAG Works

An MC-LAG deployment is built from a defined set of link types, each with a specific job. The two switches participating in the aggregation are MC-LAG peers, one taking the Active role and one the Standby role — a distinction that matters for the control-plane TCP session between them, but not for data forwarding, where both peers independently forward traffic and hold equal status. The peer-link is a direct physical connection between the two peers, used to carry traffic across when a downstream link fails. The keep-alive link is a heartbeat connection — commonly sharing the same physical link as the peer-link — that carries the control protocol traffic, synchronizes tables, and establishes the peer relationship. And an optional DAD link (Dual-Active Detection) gives the pair a way to detect a true dual-active condition — both peers acting as Active simultaneously — specifically for scenarios where peer-link and keep-alive share the same physical connection; if the keep-alive link goes down, the system automatically shuts down every interface on the standby node except logical interfaces, management ports, and the peer-link itself.

The control plane runs on ICCP (Inter-Chassis Communication Protocol), a lightweight protocol defined in RFC 7275, using TCP port 8888 between peers. Beyond the initial neighbor establishment, ICCP continuously synchronizes several categories of information: the system MAC address (so LACP packets sent to a downstream device carry an identical system ID from both peers, making cross-device aggregation actually work), MC-LAG member port configuration and status (for consistency checking and coordinated failure handling), and ARP and FDB table entries (so both peers make consistent forwarding decisions). Heartbeat packets are sent every second by default; if 15 consecutive heartbeats go unanswered, the session is declared timed out and the ICCP connection is considered broken.

To catch configuration drift before it causes a problem, MC-LAG runs a consistency check comparing settings like local/peer IP symmetry and member port configuration across the two peers, in one of three modes: idle (report only, take no action — the default), default (shut down the failing member port on both peers), or graceful (shut down the failing member port only on the standby side, keeping active service uninterrupted).

Why MC-LAG is Beneficial

  • Device-level redundancy, not just link-level: Because the two aggregated ports terminate on separate physical switches, MC-LAG survives an entire device failure — something a single switch's standard LAG configuration fundamentally cannot do.

  • No idle backup links: Presenting as a single logical device externally means MC-LAG avoids the loop-prevention trade-off of blocking redundant paths — every member link stays actively forwarding traffic.

  • Fast, coordinated failure handling: Continuous ARP, FDB, and status synchronization over ICCP means the standby peer already has the state it needs to take over forwarding the moment something changes, rather than rebuilding tables from scratch.

  • Configuration drift gets caught, not discovered in production: The consistency-check mechanism actively compares peer configuration and can act automatically — in idle, default, or graceful mode — depending on how much automatic intervention an operator wants.

  • Scales beyond a single pair: Two-level (cascaded) MC-LAG topologies let a network expand well past what one MC-LAG pair alone could support, while keeping the same redundancy model throughout.

At Asteraix

What We Can Do at Asteraix

AsterNOS implements a complete MC-LAG stack, configurable through the CLI, with dedicated support for the full range of deployment patterns operators actually build.

  • Straightforward domain and peer-link setup: mclag domain <domain-id> creates the MC-LAG domain, and a dedicated static-aggregation peer-link (recommended on a high-speed interface, with interface delayed startup to reduce packet loss on reboot) is bound with peer-link link-aggregation <lag-id> — the foundation every other MC-LAG feature builds on.

  • Flexible keep-alive and ICCP backup paths: The keep-alive link can share the peer-link or run on its own physical connection, with configurable heartbeat-interval and session-timeout, and a separate backup-channel vlan <vlan-id> lets operators configure an automatic ICCP fallback path — verified in AsterNOS's own documented example, where show mclag state visibly switches from Primary Channel to Backup Channel the moment the keep-alive link is disconnected.

  • Dual-Active Detection for shared-link scenarios: When peer-link and keep-alive share a physical connection, a separate dad local-address/dad peer-address pair with configurable dad detection-delay and recovery-delay timers protects against a true dual-active condition, automatically isolating the standby side's interfaces if the keep-alive link fails.

  • Built-in loop protection and fast convergence: monitor-link-group links uplink and downlink port state so a downstream port goes down immediately when its uplink fails, while loopback-detect (in shutdown or block-vlan mode) catches accidental loops directly on MC-LAG member interfaces — both confirmed in AsterNOS's documented examples, including a show loopback-detect output showing a live loop caught and blocked on a specific VLAN.

  • Dual-Active Gateway and Unique-IP for Layer 3/VXLAN access: When MC-LAG peers act as Layer 3 gateways, matching VLANIF IP/MAC (and matching VRF MAC for VXLAN L3 VNIs) let both peers serve as an identical gateway, while unique-ip (in diff_mac or same_mac mode) lets the MC-LAG pair still run individual routing-protocol sessions with the access side despite sharing a gateway identity.

  • Documented two-level and Layer 3 backup topologies: AsterNOS's MC-LAG documentation includes a complete cascaded (two-level) MC-LAG example for scaling host connectivity, and a Layer 3 backup-link example using OSPF and BFD over the keep-alive VLAN — giving operators a proven reference for both large-scale and high-availability-hardened designs, not just the basic single-pair case.