Shuyi Xiong

dblp:374/4242 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2024
0009-0001-6998-877XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 62% Electronic design automation · 19% Reconfigurable computing and FPGAs · 19%
Databases, data mining, and information retrieval
1 paper
Graph data management · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory management
dynamic memory management
0.812024
A Scalable, Efficient, and Robust Dynamic Memory Management Library for HLS-based FPGAs · MICRO 2024
Reconfigurable computing and FPGAs › FPGA accelerator
FPGA-based graph processing
0.812024
A Scalable, Efficient, and Robust Dynamic Memory Management Library for HLS-based FPGAs · MICRO 2024
Electronic design automation
high-level synthesis
0.812024
A Scalable, Efficient, and Robust Dynamic Memory Management Library for HLS-based FPGAs · MICRO 2024
Memory systems
processing-in-memory
0.812024
PhGraph: A High-Performance ReRAM-Based Accelerator for Hypergraph Applications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024
Memory systems › in-memory computing
ReRAM-based accelerator
0.812024
PhGraph: A High-Performance ReRAM-Based Accelerator for Hypergraph Applications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024
Graph data management › hypergraph
hypergraph processing
0.212024
PhGraph: A High-Performance ReRAM-Based Accelerator for Hypergraph Applications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024
Memory systems › memory management
memory allocation
0.212024
A Scalable, Efficient, and Robust Dynamic Memory Management Library for HLS-based FPGAs · MICRO 2024

Methods — techniques the papers use, named apart from their topics

load-balanced scheduling · 1.5hypergraph partitioning · 1.5digital memristor PIM · 1.5analog memristor PIM · 1.5memory defragmentation · 0.8graph analytics · 0.8concurrent traversal · 0.8
YearPublicationVenuePosition
2024 A Scalable, Efficient, and Robust Dynamic Memory Management Library for HLS-based FPGAs
abstract
Nowadays, high-level synthesis (HLS) has gained prominence for FPGA-based architecture prototyping, enhancing productivity significantly. Despite this advancement, HLS tools are impeded by a critical drawback: they lack support for dynamic memory management (DMM), leading to static mem-ory allocation and suboptimal use of memory resources. In response, numerous efforts have been made to develop DMM solutions compatible with HLS. However, our analysis indicates that existing solutions fail to concurrently meet the desired trifecta of scalability (efficient management of memory of any size), efficiency (minimal latency in memory (de-)allocation), and robustness (low allocation failure rates). This limitation hampers their applicability in real-world scenarios. In this paper, we introduce GraDMM, a “three-birds-one- stone” solution that comprehensively enhances the scalability, efficiency, and robustness of DMM. The key insight is to formulate memory (de-)allocation as graph analytics and lever-age sophisticated FPGA-based graph processing techniques. To achieve scalability, GraDMM specializes a simplified pipeline that significantly suppresses resource utilization expansion caused by managed memory scaling. This is crucial for managing arbitrarily sized memory on resource-limited FPGA platforms. For efficiency, GraDMM implements a data-centric concurrent traversal scheme and a shortcut-assisted fast traversal policy to accelerate (de-)allocation-guided graph traversal, reducing mem-ory (de-)allocation latency. To enhance robustness, GraDMM incorporates an adaptive memory defragmenter that defragments managed memory to minimize fragmentation-induced allocation failures. GraDMM is encapsulated as a library, providing high- level interfaces for users and ensuring synthesizability with Vi- vado HLS. Experimental results demonstrate that GraDMM out-performs three state-of-the-art HLS-compatible DMM solutions by significant margins: 56.71 %-85.59% in resource consumption savings, 78.94 % -99.99 % in (de-)allocation latency improvement, and 10.71 %-65.75% in allocation failure reduction.
Qinggang Wang, Long Zheng 0003, Zhaozeng An, Shuyi Xiong, Yu Huang 0013, Pengcheng Yao, Xiaofei Liao, Hai Jin 0001, Jingling Xue
MICRO4
2024 PhGraph: A High-Performance ReRAM-Based Accelerator for Hypergraph Applications
abstract
Hypergraph processing has emerged as an effective approach to analyze complex multilateral relationships in real-world scenarios. Existing hypergraph processing solutions based on conventional architectures are severely bottlenecked by off-chip memory accesses. In this paper, we propose the first Processing-In-Memory (PIM)-featured ReRAM-based hypergraph accelerator, dubbed PhGraph, which facilitates performance-and energy-efficient hypergraph processing. On the hardware level, PhGraph integrates analog memristor-based PIM (with high matrix-grained parallelism) and digital memristor-based PIM (for high bipartite-edge-grained efficiency) into one standalone solution. On the software level, an overlap-aware hypergraph partitioning mechanism is proposed to polarize hypergraph workloads into matrix-formatted dense and bipartite-edge-formatted sparse partitions for performance acceleration using analog memristor-based PIM and digital ones, respectively. In addition, PhGraph is equipped with load-balanced partition scheduling and algorithm mapping co-designs to boost hardware utilization and efficiency. Experimental results show that PhGraph outperforms the state-of-the-art CPU-, FPGA-, and ASIC-based solutions by up to 4,309.81×, 547.13×, and 166.76× in terms of performance, and 36,416.11×, 924.12×, and 41.44× in terms of energy-savings, respectively.
Long Zheng 0003, Ao Hu, Qinggang Wang, Yu Huang 0013, Haoqin Huang, Pengcheng Yao, Shuyi Xiong, Xiaofei Liao, Hai Jin 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7