IEICE Transactions on Fundamentals of Electronics, Communications and Computer Sciences
Online ISSN : 1745-1337
Print ISSN : 0916-8508
An HLS Library for Dynamically Adaptive SSSP Accelerator with Hybrid Node-Edge Parallelism for Irregular Graphs
Haopeng MENG, Kazutoshi WAKABAYASHI, Makoto IKEDA
Author information
JOURNAL FREE ACCESS Advance online publication

Article ID: 2026VLP0004

Details
Abstract

Single Source Shortest Path (SSSP) is a fundamental primitive in graph analytics, yet efficient hardware acceleration remains challenging due to irregular graph structures, skewed vertex degree distributions, and frequent update conflicts during parallel edge relaxation. This paper presents a parameterized High-Level Synthesis (HLS) library for dynamically adaptive SSSP acceleration based on the Delta-Stepping scheduling framework. Instead of designing dedicated hardware for individual shortest-path algorithms, the proposed library abstracts representative SSSP algorithms into a unified four-stage processing pipeline with reusable hardware modules and configurable scheduling policies, enabling customized accelerator generation while preserving a common hardware architecture. The proposed design extends the conventional bucket-based Delta-Stepping method by introducing a hierarchical priority queue that preserves local ordering within buckets while supporting flexible scheduling. Based on this framework, a hybrid node-edge parallel execution model dynamically switches between multi-vertex parallelism for sparse regions and edge-partition parallelism for high-degree vertices to improve hardware utilization for irregular graph structures. A triangular systolic-array-based conflict detection module is further integrated into the streaming pipeline to efficiently resolve concurrent updates generated by parallel edge relaxation. Implemented on a Xilinx UltraScale+ FPGA, the proposed architecture achieves up to 1042 MTEPS at 200 MHz with 32 Edge Relax functional units, providing 1.76× higher throughput than our previous FPGA implementation based on conventional Delta-Stepping scheduling, 4.10× improvement over a CPU-based implementation, and up to 22.2× higher performance than representative earlier accelerator systems. The measured peak memory bandwidth utilization reaches 62% under off-chip DRAM execution.

Content from these authors
© 2026 The Institute of Electronics, Information and Communication Engineers
Previous article Next article
feedback
Top