January 27, 2026
Article
Modern data centres supporting AI/ML, distributed storage, HPC, and cloud-native workloads require extreme throughput, ultra-low latency, and predictable performance, but as Ethernet scales to 400G and 800G, traditional host-centric networking struggles due to CPU overhead from protocol handling, queue management, and completion processing, increasing tail latency and resource contention. RDMA addresses this by enabling zero-copy, kernel-bypass data transfers that remove the CPU from the critical data path while preserving application-level work requests, and iWave Fibre SmartNICs extend this approach by offloading RDMA and storage data paths entirely into hardware, enabling autonomous, wire-speed operation with deterministic latency and minimal CPU involvement.
As Ethernet speeds scale to 400G and 800G, the inefficiencies of software-based networking stacks become impossible to ignore:
While RDMA fundamentally changes data movement between servers, end-to-end system performance is still constrained if protocol processing and storage translation remain software-driven within the infrastructure.
Compared to traditional NICs, iWave Fibre SmartNICs take acceleration further by moving RDMA transport processing, DMA orchestration, and storage data paths into programmable hardware, enabling:
At 400GbE and 800GbE speeds, network bandwidth is no longer the constraint CPU-centric protocol handling is. Interrupts, software queues, and kernel mediation introduce latency, jitter, and throughput collapse under load.
The iWave Fibre SmartNIC eliminates this bottleneck by implementing a full RDMA Core and NVMe-oF Host/Target Controller directly inside the Agilex™ 7 FPGA, enabling a fully hardware-managed, CPU-free data path.
The FPGA-resident RDMA Core fully terminates RoCE traffic in hardware, implementing transport-level functions such as connection context management, reliable delivery, congestion handling, queue pair (QP) state machines, DMA scheduling, and work completion generation without host CPU or kernel involvement. Memory registration, doorbell handling, and data placement are executed directly by the FPGA, enabling true kernel bypass and zero-copy transfers at line rate.
Figure 1: Traditional NIC data path with CPU-mediated memory copies.
Figure 1 illustrates data movement using a traditional NIC. Data originating in the application memory space (host DDR) is first copied by the CPU into an intermediate networking buffer in host memory, and then copied again into the NIC for transmission. On the receiving client, the process is reversed: the NIC transfers the data into a CPU-managed networking buffer, after which the CPU copies it into the application memory space. This multi-stage data movement results in additional memory copies, increased CPU involvement, and higher end-to-end latency.
Figure 2: RDMA NIC data path with direct application memory access.
Figure 2 shows data movement using iWave’s RDMA-enabled NIC. Here, the RDMA NIC directly accesses application memory by registering a region of host DDR during initialization. This enables direct data transfers between application memory and the NIC, eliminating CPU-mediated memory copies. On the receiving client, the RDMA-enabled NIC writes incoming data directly into the application memory space without CPU intervention. This architecture significantly reduces CPU overhead, memory bandwidth consumption, and overall latency.
RDMA payloads are then handed off internally to a hardware NVMe-oF Host/Target Controller instantiated in the same FPGA fabric. This controller performs NVMe command decoding, submission and completion queue management, namespace mapping, and PCIe transaction generation toward NVMe SSDs or downstream PCIe switch fabrics. The entire NVMe command lifecycle is handled deterministically in hardware, eliminating software queues, interrupts, and polling loops.
By tightly coupling the RDMA Core and NVMe-oF Controller within the FPGA, the SmartNIC creates a single, continuous data path from Ethernet to NVMe. Network packets are transformed directly into NVMe operations with no intermediate buffering in host memory and no dependency on a host OS or embedded CPU. This hardware-only pipeline removes OS jitter, CPU scheduling effects, and tail-latency spikes, delivering predictable latency and sustained throughput at 400GbE and 800GbE.
While RDMA solves transport inefficiencies on the wire, the architecture of the storage enclosure itself has remained a bottleneck. Traditional Just-a-Bunch-of-Flash (JBoF) systems act as dense storage shelves but suffer from a critical design flaw: they rely on embedded x86 CPUs to manage data flow.
iWave, in collaboration with Altera, has introduced a CPU-Free JBoF architecture that fundamentally alters this paradigm. By removing the general-purpose processor from the storage shelf, the architecture shifts all control and data plane logic to the SmartNIC, creating a system that is leaner, faster, and more secure.
In a standard JBoF implementation, the data path is interrupted by software:
The iWave architecture eliminates the embedded CPU entirely. The Agilex™ 7-based iWave Fibre SmartNIC becomes the central control engine and data mover, connecting directly to the PCIe switches and NVMe drives.
Enabling Next-Gen Workloads
This architectural shift is specifically designed for environments where average performance is insufficient:
For more information contact us at mktg@iwave-global.com and visit www.iwave-global.com.
We appreciate you contacting iWave.
Our representative will get in touch with you soon!