EDBT 2026 Demo / reviewers in the wild / expert
Hsin-Hsuan Sung
dblp:248/6038
· DBLP profile ↗
12ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-6186-7669ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | N:M sparsity-oriented graph reordering for accelerating GNNs on GPU sparse tensor coresabstractRecent GPUs incorporate Sparse Tensor Cores (SPTC) to accelerate computations on matrices that satisfy structured N : M sparsity; software stacks further extend support to generalized V : N : M patterns. Although graphs in graph neural networks (GNNs) are typically sparse, their sparsity is irregular and rarely conforms to these patterns. This paper introduces a lossless graph reordering algorithm that reshapes irregular graph data into the required sparse patterns, allowing GNN workloads to exploit SPTC. The transformation preserves model accuracy and maintains the symmetry of graph adjacency matrices, ensuring compatibility with symmetry-based graph algorithms. On the SuiteSparse collection, our method removes 93.77%–98.51% of vector-level N : M violations and increases the fraction of conforming graphs from 5%–7% to 79.3%–91.3%. On NVIDIA A100 GPUs, it accelerates SpMM by up to 50.7 × (geometric-mean speedups 3.44 × –7.25 × ) over cuSPARSE, and speeds up key GNN graph operations on real graphs by as much as 6.7 × (2.6 × on average). Ruifeng Zhang 0008, Jou-An Chen, Hsin-Hsuan Sung, Ang Li 0006, Xipeng Shen |
J. Parallel Distributed Comput. | 3 |
| 2025 | Accelerating GNNs on GPU Sparse Tensor Cores through N: M Sparsity-Oriented Graph ReorderingabstractRecent GPUs have introduced Sparse Tensor Cores (SPTC) to accelerate computations on sparse matrices meeting the N:M sparse patterns. Software tools expand the support to more general V:N:M patterns. Graphs in Graph Neural Networks (GNNs) are typically sparse, but the sparsity is often irregular, not conforming to the required V:N:M sparse patterns. This paper proposes a novel graph reordering algorithm to transform irregular graph data into the required sparse patterns for GNNs to benefit from SPTC. The optimization is lossless, maintaining the accuracy of GNN. It at the same time keeps the symmetry of the adjacency matrices of the graphs so that the same matrices can remain compatible with many symmetry-based graph algorithms. The optimization successfully removes 98-100% violations of the N:M sparse patterns at the vector level and increases the portion of conforming graphs in the SuiteSparse collection from 5-9% to 88.7-93.5%. On A100 GPUs, the optimization accelerates Sparse Matrix Matrix (SpMM) by up to 43X (a geomean speedup of 2.3X - 7.5X) over cuSPARSE and speeds up the key graph operations in GNNs on real graphs by as much as 8.6X (3.5X on average). Jou-An Chen, Hsin-Hsuan Sung, Ruifeng Zhang 0008, Ang Li 0006, Xipeng Shen |
PPoPP | 2 |
| 2024 | Enabling Efficient Deep Learning on MCU With Transient Redundancy EliminationabstractDeploying deep neural networks (DNNs) with satisfactory performance in resource-constrained environments is challenging. This is especially true of microcontrollers due to their tight space and computational capabilities. However, there is a growing demand for DNNs on microcontrollers, as executing large DNNs on microcontrollers is critical to reducing energy consumption, increasing performance efficiency, and eliminating privacy concerns. This paper presents a novel and systematic data redundancy elimination method to implement efficient DNNs on microcontrollers through innovations in computation and space optimization. By making the optimization itself a trainable component in the target neural networks, this method maximizes performance benefits while keeping the DNN accuracy stable. Experiments are performed on two microcontroller boards with three popular DNNs, namely CifarNet, ZfNet and SqueezeNet. Experiments show that this solution eliminates more than 96% of computations in DNNs and makes them fit well on microcontrollers, yielding 3.4-5$\times$speedup with little loss of accuracy. Jiesong Liu, Feng Zhang 0007, Jiawei Guan, Hsin-Hsuan Sung, Xiaoguang Guo, Saiqin Long, Xiaoyong Du 0001, Xipeng Shen |
IEEE Trans. Computers | 4 |
| 2023 | Space-Efficient TREC for Enabling Deep Learning on MicrocontrollersabstractDeploying deep neural networks (DNNs) for a resource-constrained environment and achieving satisfactory performance is challenging. It is especially so on microcontrollers for their stringent space and computing power. This paper focuses on new ways to make TREC, an optimization recently proposed to enable computation reuse in DNNs, space and time efficient on Microcontrollers. The solution maximizes the performance benefits while keeping the DNN accuracy stable. Experiments show that the solution eliminates over 96% computations in DNNs and makes them fit well into microcontrollers, producing 3.4-5× speedups with only marginal accuracy loss. Jiesong Liu, Feng Zhang 0007, Jiawei Guan, Hsin-Hsuan Sung, Xiaoguang Guo, Xiaoyong Du 0001, Xipeng Shen |
ASPLOS (3) | 4 |
| 2023 | BitGNN: Unleashing the Performance Potential of Binary Graph Neural Networks on GPUsabstractRecent studies have shown that Binary Graph Neural Networks (GNNs) are promising for saving computations of GNNs through binarized tensors. Prior work, however, mainly focused on algorithm designs or training techniques, leaving it open to how to materialize the performance potential on accelerator hardware fully. This work redesigns the binary GNN inference backend from the efficiency perspective. It fills the gap by proposing a series of abstractions and techniques to map binary GNNs and their computations best to fit the nature of bit manipulations on GPUs. Results on real-world graphs with GCNs, GraphSAGE, and GraphSAINT show that the proposed techniques outperform state-of-the-art binary GNN implementations by 8-22X with the same accuracy maintained. BitGNN code is publicly available.1. Jou-An Chen, Hsin-Hsuan Sung, Xipeng Shen, Sutanay Choudhury, Ang Li 0006 |
ICS | 2 |
| 2023 | Decentralized Application-Level Adaptive Scheduling for Multi-Instance DNNs on Open Mobile Devices
Hsin-Hsuan Sung, Jou-An Chen, Wei Niu 0002, Jiexiong Guan, Bin Ren 0002, Xipeng Shen |
USENIX ATC | 1 |
| 2023 | Accelerating matrix-centric graph processing on GPUs through bit-level optimizations
Jou-An Chen, Hsin-Hsuan Sung, Xipeng Shen, Nathan R. Tallent, Kevin J. Barker, Ang Li 0006 |
J. Parallel Distributed Comput. | 2 |
| 2022 | Bit-GraphBLAS: Bit-Level Optimizations of Matrix-Centric Graph Processing on GPUabstractIn a general graph data structure like an adjacency matrix, when edges are homogeneous, the connectivity of two nodes can be sufficiently represented using a single bit. This insight has, however, not yet been adequately exploited by the existing matrix-centric graph processing frameworks. This work fills the void by systematically exploring the bit-level representation of graphs and the corresponding optimizations to the graph operations. It proposes a two-level representation named Bit-Block Compressed Sparse Row (B2SR) and presents a series of optimizations to the graph operations on B2SR by leveraging the intrinsics of modern GPUs. Evaluations on NVIDIA Pascal and Volta GPUs show that the optimizations bring up to 40× and 6555× for essential GraphBLAS kernels SpMV and SpGEMM, respectively, making GraphBLAS-based BFS accelerate up to 433×, SSSP, PR, and CC up to 35×, and TC up to 52×. Jou-An Chen, Hsin-Hsuan Sung, Xipeng Shen, Nathan R. Tallent, Kevin J. Barker, Ang Li 0006 |
IPDPS | 2 |
| 2022 | TREC: Transient Redundancy Elimination-based ConvolutionabstractThe intensive computations in convolutional neural networks (CNNs) pose challenges for resource-constrained devices; eliminating redundant computations from convolution is essential. This paper gives a principled method to detect and avoid transient redundancy, a type of redundancy existing in input data or activation maps and hence changing across inferences. By introducing a new form of convolution (TREC), this new method makes transient redundancy detection and avoidance an inherent part of the CNN architecture, and the determination of the best configurations for redundancy elimination part of CNN backward propagation. We provide a rigorous proof of the robustness and convergence of TREC-equipped CNNs. TREC removes over 96% computations and achieves 3.51x average speedups on microcontrollers with minimal (about 0.7%) accuracy loss. Jiawei Guan, Feng Zhang 0007, Jiesong Liu, Hsin-Hsuan Sung, Xiaoyong Du 0001, Xipeng Shen |
NeurIPS | 4 |
| 2022 | Brief Industry Paper: Enabling Level-4 Autonomous Driving on a Single $1k Off-the-Shelf CardabstractIn the past few years we have developed hardware computing systems for commercial autonomous vehicles, but inevitably the high development cost and long turn-around time have been major roadblocks for commercial deployment. Hence we also explored the potential of software optimization. This paper, for the first-time, shows that it is feasible to enable full leve1-4 autonomous driving workloads on a single off-the-shelf card (Jetson AGX Xavier) for less than ${\$}1\mathrm{k}$, an order of magnitude less than the state-of-the-art systems, while meeting all the requirements of latency. The success comes from the resolution of some important issues shared by existing practices through a series of measures and innovations. Hsin-Hsuan Sung, Yuanchao Xu 0001, Jiexiong Guan, Wei Niu 0002, Bin Ren 0002, Yanzhi Wang 0001, Shaoshan Liu, Xipeng Shen |
RTAS | 1 |
| 2021 | Brief Industry Paper: Towards Real-Time 3D Object Detection for Autonomous Vehicles with Pruning SearchabstractIn autonomous driving, 3D object detection is es-sential as it provides basic knowledge about the environment. However, as deep learning based 3D detection methods are usually computation intensive, it is challenging to support realtime 3D object detection on edge-computing devices in selfdriving cars with limited computation and memory resources. To facilitate this, we propose a compiler-aware pruning search framework, to achieve real-time inference of 3D object detection on the resource-limited mobile devices. Specifically, a generator is applied to sample better pruning proposals in the search space based on current proposals with their performance, and an evaluator is adopted to evaluate the sampled pruning proposal performance. To accelerate the search, the evaluator employs Bayesian optimization with an ensemble of neural predictors. We demonstrate in experiments that for the first time, the pruning search framework can achieve real-time 3D object detection on mobile (Samsung Galaxy S20 phone) with state-of-the-art detection performance. Pu Zhao 0001, Wei Niu 0002, Geng Yuan, Yuxuan Cai 0001, Hsin-Hsuan Sung, Shaoshan Liu, Sijia Liu 0001, Xipeng Shen, Bin Ren 0002, Yanzhi Wang 0001, Xue Lin 0001 |
RTAS | 5 |
| 2019 | Finding Documents Related to Taiwan in the Veritable Records of Qing Using Relevance Feedback
Hsin-Hsuan Sung, Jou-An Chen, Jieh Hsiang |
TPDL | 1 |