In Network Packet Broker (NPB) deployments, DPU selection determines device power consumption, processing efficiency, and product lifecycle. This guide compares 2 Marvell DPU chips — OCTEON TX2 CN9670 and OCTEON 10 CN10308 — against 100G NPB scenario requirements.
1. Scenario Constraints and Requirements
Why 100G Is the Bandwidth Ceiling
Bandwidth is capped at 100G by 2 factors: the PX306P-48Y-M switch chip's internal topology, which fixes DPU-to-switch interconnect bandwidth at 100G with no hardware path beyond it, and current DPU performance, which is matched to 100G line rate without requirement for 200G or higher.
Traffic forwarded to the DPU has already been filtered by the switch chip. The DPU handles only FusionNOS-layer processing — keyword filtering, packet masking. At 100G, hardware capability, chip performance, and workload requirements align without excess capacity.
Core Requirements for 100G NPB
Bandwidth: stable 100G line rate, zero packet loss, low latency, matched to the PX306P-48Y-M with DPU bandwidth design
Processing: switch chip handles base NPB forwarding; DPU handles FusionNOS advanced processing; hardware specification (coprocessor, memory) affects processing performance — coprocessor capability primarily affects CRC and compression throughput
Deployment and lifecycle: power control, cost control, long-term supply, expansion capacity, avoidance of selection obsolescence, compatibility with hardware specification differences across existing and new deployments
2. Parameter Comparison
Dimension | CN10308 (GHC103 card) | CN9670 (GHC96 card) | Impact on 100G NPB |
|---|---|---|---|
Process node | 5nm (TSMC) | 16nm | CN10308: lower leakage, higher power efficiency, less heat. CN9670: older process, lower efficiency, more heat, higher cooling requirement |
CPU | 8-core Arm Neoverse N2 (Armv9), 2.5GHz | 24-core custom OCTEON TX2 core (Armv8.2), 1.8GHz | CN10308 single-core IPC is 3× CN9670; lower per-core processing latency for FusionNOS. CN9670 has more cores for multi-core concurrent workloads, but lower clock speed and lower per-core efficiency |
Card power | ~40W | ~80W | CN10308: lower cooling requirement, lower long-term operating cost, reduced cooling hardware investment. CN9670: higher cooling requirement, higher operating cost |
Memory slots | 2× DDR5, 48GB max | 3× DDR4, 96GB max | CN10308: DDR5, higher speed, 48GB ceiling. CN9670: DDR4, 96GB ceiling |
DPU-to-switch interconnect | 100G | 100G | PX306P-48Y-M card supports 100G only; interconnect bandwidth difference between the 2 chips has no practical effect at this scenario's line rate |
Lifecycle | Current-generation platform, long-term supply, active software updates and fixes | Previous-generation platform, being phased out, limited future software support, increasing component sourcing difficulty | CN10308: lower obsolescence risk for new projects, longer-term support. CN9670: limited future support, suited to short-term or transitional deployment only |
Forwarding at 100G | Sustains line rate, no bottleneck, lower forwarding latency | Sustains line rate, no bottleneck, slightly higher forwarding latency | Both sustain 100G; CN9670's lower core efficiency and higher power result in slightly higher forwarding latency |
Hardware compatibility | Full compatibility with PX306P-48Y-M, supports FusionNOS | Full compatibility with PX306P-48Y-M, supports FusionNOS | No compatibility difference |
Storage | 64GB eMMC, optional NVMe | 64GB eMMC, optional NVMe | Identical |
Additional | Hardware ML/AI acceleration engine | None | CN10308's ML/AI engine supports traffic anomaly detection, adaptive QoS, and security threat identification — a capability CN9670 does not have |
3. Power Consumption: Contributing Factors
The power gap between the 2 chips (CN9670 ~80W+, CN10308 ~40W) traces to 2 structural differences.
Primary: Process Node and CPU Architecture
Process node: CN10308's 5nm process has lower leakage current and higher power efficiency, giving it a lower base power draw than CN9670's 16nm process.
CPU architecture: FusionNOS relies on a global hash mechanism. CN10308's 8-core Armv9 architecture delivers higher per-core efficiency with lower lock-contention overhead. CN9670's 24-core Armv8.2 architecture has lower per-core efficiency and higher multi-core coordination overhead, keeping power draw elevated.
Secondary: Interfaces and Acceleration Modules
CN10308's 50G SerDes, DDR5 memory, and hardware acceleration modules further reduce power draw and raise FusionNOS processing efficiency. CN9670's equivalent modules operate at lower efficiency, adding power consumption.
4. Selection Rationale: CN10308
Efficiency: at 100G, CN10308's higher per-core performance avoids multi-core lock-contention overhead; core utilization exceeds CN9670, raising FusionNOS processing efficiency
Power: 5nm process brings CN10308 power draw to approximately 50% of CN9670, lowering long-term operating cost
Memory: DDR5 support provides higher frequency and bandwidth than CN9670's DDR4
Extensibility: hardware ML/AI acceleration engine supports future traffic anomaly detection, adaptive QoS, and security threat identification
5. Selection Conclusion
CN10308 — Recommended for New Deployments
Fits new 100G NPB designs meeting these conditions:
Power and cooling requirements are a priority; lower long-term operating cost and reduced cooling hardware investment are desired
Workload does not depend on coprocessor CRC/compression functions; memory expansion requirement is low
Product longevity, stable long-term supply, and obsolescence avoidance are priorities
FusionNOS processing efficiency, per-core performance, and forwarding latency are priorities
Future expansion into ML/AI-driven traffic anomaly detection, adaptive QoS, or security threat identification is planned
CN9670 — Applicable in Specific Scenarios
Fits deployments meeting these conditions:
Workload depends on coprocessor functions such as CRC or compression
Memory capacity requirement exceeds 48GB and requires large-memory deployment
Existing deployment upgrade requiring compatibility with current hardware/software design, with no power optimization requirement and no budget for hardware/software rework
No near-term product iteration plan; higher power draw and limited future support are acceptable for short-term or transitional deployment
6. Common Selection Errors
More cores does not mean higher performance. At 100G, CN9670's 24 cores carry low per-core efficiency and high lock-contention overhead; core utilization is low and additional cores add power draw without adding throughput. Per-core performance is the determining factor.
An older platform is not inherently more stable. CN10308, as the current-generation platform, matches CN9670 on stability and exceeds it on supply availability and expansion capacity.
NPB and FusionNOS are not the same system. The two operate at different layers: NPB base forwarding runs on the switch chip under AsterNOS; FusionNOS advanced processing runs on the DPU. Each handles a distinct function.