November 28, 2025
Article
The surge in data intensive applications has elevated storage to a critical performance component. With legacy interfaces like SATA unable to meet modern bandwidth demands, the industry has shifted toward NVMe over PCIe for low-latency, high-throughput. iWave’s NVMe Host Controller IP Core brings this capability to FPGA platforms, delivering up to 7 GB/s performance with a scalable architecture and broad SSD interoperability
NVMe defines two primary queue types, the Admin Queue and the I/O Queue, each comprising a Submission Queue (SQ) and a Completion Queue (CQ) that work together to manage command execution and completion reporting. The Admin Queue pair handles controller initialization, configuration, and management operations, while the I/O Queue pairs are dedicated to data transfer commands between the host and non-volatile memory.
The SQ-CQ mechanism operates as an efficient command passing structure. The SQ where the host posts commands, and a CQ, where the SSD reports completed commands. NVMe can support up to 64K I/O queues, each capable of holding 64K commands, providing deep parallelism that can scale efficiently with multi-core processors. This architecture, with support for multiple queues, dedicated doorbell registers, and circular buffers, ensures low-latency command processing, high parallelism, and efficient utilization of both host and SSD resources.
When the host issues a new command, it formats it per the NVMe specification and places it in the appropriate Submission Queue (SQ).Each SQ has a dedicated Tail Doorbell Register, which the host updates to notify the SSD’s NVMe controller that new commands are available. The controller fetches the commands from the SQ, executes them, and places the corresponding completion entries into the associated Completion Queue (CQ).
The controller may then generate an interrupt to signal command completion. After the host processes the completion entry, it updates the CQ Head Doorbell Register, indicating that the entry has been consumed and the queue slot can be reused.
Implementing NVMe on an FPGA unlocks ultra-high throughput and fully leverages modern SSD performance. Unlike fixed-function controllers, FPGA-based designs allow fine-tuning of queue depth, outstanding command count, and on-chip buffer allocation, optimizing performance for different SSD models. Peak throughput can often be achieved with modest queue depths and minimal buffering, ensuring efficient use of FPGA resources. These designs are also highly portable across FPGA families and vendors.
To maximize throughput and efficient data transfer, the NVMe Host controller must be optimized to efficiently handle data transfer from either the user logic or the SSD, with minimum propagating delays upstream or downstream. This can be achieved through a combination of techniques:
By implementing these strategies, FPGA-based NVMe host controllers can sustain high data throughput while maintaining deterministic latency, fully exploiting the capabilities of modern SSDs.
iWave’s NVMe Host Controller IP Core is compliant with the NVM Express Base Specification 1.4, enabling seamless, high-speed memory transfers to and from NVMe storage devices. It uses internal memory for buffering and leverages external DDR for applications that cannot tolerate backpressure. The IP Core interfaces via PCIe, making it ideal for high-performance, large-capacity storage systems
Key Features & Highlights :
iWave NVMe Host controller Performance Validation and Results
iWave’s NVMe Host Controller IP Core has been extensively validated across multiple FPGA platforms and vendors, supporting both PCIe Gen3 and Gen4 implementations. Engineered for efficiency and maximum throughput, the IP consistently delivers near-optimal performance across a wide range of configurations. Key performance results with different SSDs are summarized below, showcasing its reliability and high-speed capabilities.
Throughput analysis of iWave NVMe Gen3 host controller IP with Samsung 970 evo plus SSD
| 1. | Capacity | : 2TB |
| 2. | NAND Flash Memory | : 3bit MLC(TLC) |
| 3. | Sequential Write Speed | : Before SLC buffer is filled : 3300 MB/s : After SLC buffer is filled : 1750 MB/s |
| 4. | Sequential Read Speed | : 3500 MB/s |
| 5. | SLC Buffer Size | : 78GB |
Throughput analysis of iWave NVMe Gen4 host controller IP with WD SN850X SSD
| 1. | Capacity | : 2TB |
| 2. | NAND Flash Memory | : 3D TLC |
| 3. | Sequential Write Speed | : Before SLC buffer is filled : 6600 MB/s : After SLC buffer is filled : 1600 MB/s |
| 4. | Sequential Read Speed | : 7300 MB/s |
| 5. | SLC Buffer Size | : 600 GB |
The rise of data-intensive applications is pushing storage to its limits, and NVMe’s scalable, low-latency architecture is key to unlocking modern SSD performance. iWave’s FPGA-optimized NVMe Host Controller IP Core maximizes throughput reaching up to 7 GB/s on Gen4 SSDs through intelligent queue management, buffering, and PCIe integration. Proven across multiple SSD vendors and FPGA families, it provides a high-performance foundation for next-generation storage acceleration and real-time data capture.
We appreciate you contacting iWave.
Our representative will get in touch with you soon!