HPE – Certified Genuine Parts InfiniBand Switch – 40 Ports – P08358-001
Forty ports.
One non-blocking fabric.
A field guide to the Mellanox InfiniBand HDR 40-port QSFP56 switch — the workhorse leaf/spine building block behind a lot of the HPC and AI clusters we quote.
Every high-performance cluster eventually runs into the same wall: the compute is only as fast as the fabric connecting it. That’s the problem this switch was built to remove. Whether it lands in your rack under an HPE part number or straight from NVIDIA/Mellanox as the Quantum QM8700, it’s the same 1U engine underneath — and it’s still one of the most cost-effective ways to build a truly non-blocking InfiniBand fabric.
01 What HDR actually buys you
InfiniBand generations are named for their signaling rate, and HDR sits at 200 Gb/s per port — double the previous EDR generation. This switch packs 40 of those ports into a single rack unit, built around Mellanox’s Quantum switch silicon. Do the math and you get up to 16 Tb/s of non-blocking bandwidth moving through the chassis with sub-90-nanosecond port-to-port latency, which is the number that matters most once your workload starts caring about tail latency instead of just throughput.
02 The splitter trick
Not every node needs a full 200Gb/s lane. Using splitter cables, each HDR port can fan out into two HDR100 (100Gb/s) connections, effectively doubling the port count to 80 for environments where per-node bandwidth needs are more modest than the fabric’s ceiling. It’s a cheap way to stretch a single switch across more racks without giving up headroom for the nodes that do need full HDR.
03 Compute happening inside the switch
What separates this from a “dumb” high-speed switch is SHArP — Mellanox’s Scalable Hierarchical Aggregation and Reduction Protocol. It lets the switch itself perform reduction operations (the kind used constantly in MPI collectives and distributed training) as data passes through the fabric, instead of shipping everything back to compute nodes to be aggregated. For HPC and AI training clusters leaning on all-reduce operations, that’s latency and CPU cycles given back to the workload.
04 Managed vs. unmanaged, and what ships in the box
The switch comes in managed and unmanaged variants, and in both standard-depth and back-to-front airflow configurations to match your row’s hot/cold aisle orientation. The managed version carries its own x86 dual-core subsystem for fabric management, independent of your compute nodes.
| Part number | P08358-001 |
| Port count | 40x QSFP56 |
| Signaling rate | HDR — 200 Gb/s per port |
| Aggregate bandwidth | Up to 16 Tb/s, non-blocking |
| Port-to-port latency | Sub-90 ns |
| Switch silicon | Mellanox Quantum, SHArP-capable |
| Form factor | 1U, standard or back-to-front airflow |
| Power | Dual AC power supplies |
| Management | Managed (x86 dual-core) or unmanaged |
| Weight | ~30.9 lb |
05 Why it still matters
- It’s a proven, widely-deployed building block — plenty of secondary-market and refurbished supply exists, which keeps cost per port down.
- Non-blocking design means you’re not oversubscribing the fabric as you scale racks.
- In-network SHArP computing offloads work that would otherwise burn CPU cycles on your compute nodes.
- Splitter cabling gives you flexibility to mix full-HDR and HDR100 nodes on the same leaf switch.
If you’re speccing out an InfiniBand fabric — new or refurbished — this is usually where we start the conversation. Reach out and we’ll help you size the leaf/spine layout against your node count and budget.
