Blog

Packet Broker DPU Selection Guide: Marvell OCTEON 10 vs OCTEON TX2

In Network Packet Broker (NPB) deployments, DPU selection determines device power consumption, processing efficiency, and product lifecycle. This guide compares 2 Marvell DPU chips — OCTEON TX2 CN9670 and OCTEON 10 CN10308 — against 100G NPB scenario requirements.

1. Scenario Constraints and Requirements

Why 100G Is the Bandwidth Ceiling

Bandwidth is capped at 100G by 2 factors: the PX306P-48Y-M switch chip's internal topology, which fixes DPU-to-switch interconnect bandwidth at 100G with no hardware path beyond it, and current DPU performance, which is matched to 100G line rate without requirement for 200G or higher.

Traffic forwarded to the DPU has already been filtered by the switch chip. The DPU handles only FusionNOS-layer processing — keyword filtering, packet masking. At 100G, hardware capability, chip performance, and workload requirements align without excess capacity.

Core Requirements for 100G NPB

  • Bandwidth: stable 100G line rate, zero packet loss, low latency, matched to the PX306P-48Y-M with DPU bandwidth design

  • Processing: switch chip handles base NPB forwarding; DPU handles FusionNOS advanced processing; hardware specification (coprocessor, memory) affects processing performance — coprocessor capability primarily affects CRC and compression throughput

  • Deployment and lifecycle: power control, cost control, long-term supply, expansion capacity, avoidance of selection obsolescence, compatibility with hardware specification differences across existing and new deployments

2. Parameter Comparison

Dimension

CN10308 (GHC103 card)

CN9670 (GHC96 card)

Impact on 100G NPB

Process node

5nm (TSMC)

16nm

CN10308: lower leakage, higher power efficiency, less heat. CN9670: older process, lower efficiency, more heat, higher cooling requirement

CPU

8-core Arm Neoverse N2 (Armv9), 2.5GHz

24-core custom OCTEON TX2 core (Armv8.2), 1.8GHz

CN10308 single-core IPC is 3× CN9670; lower per-core processing latency for FusionNOS. CN9670 has more cores for multi-core concurrent workloads, but lower clock speed and lower per-core efficiency

Card power

~40W

~80W

CN10308: lower cooling requirement, lower long-term operating cost, reduced cooling hardware investment. CN9670: higher cooling requirement, higher operating cost

Memory slots

2× DDR5, 48GB max

3× DDR4, 96GB max

CN10308: DDR5, higher speed, 48GB ceiling. CN9670: DDR4, 96GB ceiling

DPU-to-switch interconnect

100G

100G

PX306P-48Y-M card supports 100G only; interconnect bandwidth difference between the 2 chips has no practical effect at this scenario's line rate

Lifecycle

Current-generation platform, long-term supply, active software updates and fixes

Previous-generation platform, being phased out, limited future software support, increasing component sourcing difficulty

CN10308: lower obsolescence risk for new projects, longer-term support. CN9670: limited future support, suited to short-term or transitional deployment only

Forwarding at 100G

Sustains line rate, no bottleneck, lower forwarding latency

Sustains line rate, no bottleneck, slightly higher forwarding latency

Both sustain 100G; CN9670's lower core efficiency and higher power result in slightly higher forwarding latency

Hardware compatibility

Full compatibility with PX306P-48Y-M, supports FusionNOS

Full compatibility with PX306P-48Y-M, supports FusionNOS

No compatibility difference

Storage

64GB eMMC, optional NVMe

64GB eMMC, optional NVMe

Identical

Additional

Hardware ML/AI acceleration engine

None

CN10308's ML/AI engine supports traffic anomaly detection, adaptive QoS, and security threat identification — a capability CN9670 does not have

3. Power Consumption: Contributing Factors

The power gap between the 2 chips (CN9670 ~80W+, CN10308 ~40W) traces to 2 structural differences.

Primary: Process Node and CPU Architecture

Process node: CN10308's 5nm process has lower leakage current and higher power efficiency, giving it a lower base power draw than CN9670's 16nm process.

CPU architecture: FusionNOS relies on a global hash mechanism. CN10308's 8-core Armv9 architecture delivers higher per-core efficiency with lower lock-contention overhead. CN9670's 24-core Armv8.2 architecture has lower per-core efficiency and higher multi-core coordination overhead, keeping power draw elevated.

Secondary: Interfaces and Acceleration Modules

CN10308's 50G SerDes, DDR5 memory, and hardware acceleration modules further reduce power draw and raise FusionNOS processing efficiency. CN9670's equivalent modules operate at lower efficiency, adding power consumption.

4. Selection Rationale: CN10308

  • Efficiency: at 100G, CN10308's higher per-core performance avoids multi-core lock-contention overhead; core utilization exceeds CN9670, raising FusionNOS processing efficiency

  • Power: 5nm process brings CN10308 power draw to approximately 50% of CN9670, lowering long-term operating cost

  • Memory: DDR5 support provides higher frequency and bandwidth than CN9670's DDR4

  • Extensibility: hardware ML/AI acceleration engine supports future traffic anomaly detection, adaptive QoS, and security threat identification

5. Selection Conclusion

CN10308 — Recommended for New Deployments

Fits new 100G NPB designs meeting these conditions:

  1. Power and cooling requirements are a priority; lower long-term operating cost and reduced cooling hardware investment are desired

  2. Workload does not depend on coprocessor CRC/compression functions; memory expansion requirement is low

  3. Product longevity, stable long-term supply, and obsolescence avoidance are priorities

  4. FusionNOS processing efficiency, per-core performance, and forwarding latency are priorities

  5. Future expansion into ML/AI-driven traffic anomaly detection, adaptive QoS, or security threat identification is planned

CN9670 — Applicable in Specific Scenarios

Fits deployments meeting these conditions:

  1. Workload depends on coprocessor functions such as CRC or compression

  2. Memory capacity requirement exceeds 48GB and requires large-memory deployment

  3. Existing deployment upgrade requiring compatibility with current hardware/software design, with no power optimization requirement and no budget for hardware/software rework

  4. No near-term product iteration plan; higher power draw and limited future support are acceptable for short-term or transitional deployment

6. Common Selection Errors

More cores does not mean higher performance. At 100G, CN9670's 24 cores carry low per-core efficiency and high lock-contention overhead; core utilization is low and additional cores add power draw without adding throughput. Per-core performance is the determining factor.

An older platform is not inherently more stable. CN10308, as the current-generation platform, matches CN9670 on stability and exceeds it on supply availability and expansion capacity.

NPB and FusionNOS are not the same system. The two operate at different layers: NPB base forwarding runs on the switch chip under AsterNOS; FusionNOS advanced processing runs on the DPU. Each handles a distinct function.

Keep reading