September 17, 2026
Article
Storage and compute have been drifting apart for a decade. Hyperscale data centres, AI training clusters, and high performance computing systems no longer want storage locked inside a single chassis they want a pool of NVMe drives that any compute node can reach at near local latency. NVMe over Fabrics (NVMe-oF) is the protocol that makes this possible, and RDMA over Converged Ethernet v2 (RoCEv2) is one of the most widely deployed transports for carrying it. iWave’s FPGA based RoCEv2 and NVMe Host IP cores are built specifically to implement this stack in hardware, where software based initiators and targets run out of headroom.
Local NVMe drives are fast hundreds of thousands of IOPS and microsecond class latency but that performance is trapped behind the PCIe slot it’s plugged into. NVMe-oF extends the NVMe command set and queueing model across a network fabric, so a remote drive looks, to the host software stack, almost identical to a local one. Instead of the SSD being physically connected to the host, the host can access an NVMe namespace located in a remote storage subsystem.
The two most common transports for this are:
Among these, NVMe over RDMA using RoCEv2 is particularly attractive for FPGA based systems because it combines the high bandwidth of Ethernet with RDMA’s low latency, low CPU overhead data movement.
This is where iWave’s FPGA IP cores bridge the performance gap. Instead of relying on host CPU software to process the network and storage stack, key transport and initiator functions can be implemented directly in the FPGA fabric. This creates a hardware accelerated end-to-end data path, reducing CPU overhead while enabling low-latency, high-throughput access to remote storage resources.
| IP Core | Purpose | Major Capabilities |
|---|---|---|
| RoCEv2 IP Core | FPGA-based RDMA networking | RDMA transport, QP/CQ, segmentation/reassembly, retransmission, integrity checking, AXI4 DMA, up to 800G |
| NVMe-oF Bridge Core | Local-to-remote NVMe storage bridge | NVMe command processing, RDMA, buffering, arbitration, multi queue support, remote NVMe/JBOF connectivity |
| NVMe Host IP Core | Hardware NVMe initiator | NVMe SQ/CQ, command processing, local PCIe NVMe, remote NVMe-oF, unified storage interface |
Together, these cores enable an FPGA platform to implement NVMe-oF without routing every I/O operation through the general purpose CPU’s network and storage stacks.
The practical deployment model looks like this:
NVMe-oF over RoCEv2 addresses the need to separate storage resources from individual compute systems while maintaining high speed data access. By implementing RDMA transport and storage initiator functions directly in FPGA fabric, iWave’s IP cores reduce host CPU involvement and provide an efficient hardware data path from the network to host memory.
The architecture supports current high speed Ethernet deployments and provides a flexible foundation for future storage fabrics and emerging technologies such as UEC/UET.
To learn more about iWave embedded solutions and design services, visit www.iwave-global.com or contact us at mktg@iwave-global.com
We appreciate you contacting iWave.
Our representative will get in touch with you soon!