EDBT 2026 Demo / reviewers in the wild / expert
Jou-An Chen
dblp:248/6223
· DBLP profile ↗
10ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-9820-9210ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | N:M sparsity-oriented graph reordering for accelerating GNNs on GPU sparse tensor coresabstractRecent GPUs incorporate Sparse Tensor Cores (SPTC) to accelerate computations on matrices that satisfy structured N : M sparsity; software stacks further extend support to generalized V : N : M patterns. Although graphs in graph neural networks (GNNs) are typically sparse, their sparsity is irregular and rarely conforms to these patterns. This paper introduces a lossless graph reordering algorithm that reshapes irregular graph data into the required sparse patterns, allowing GNN workloads to exploit SPTC. The transformation preserves model accuracy and maintains the symmetry of graph adjacency matrices, ensuring compatibility with symmetry-based graph algorithms. On the SuiteSparse collection, our method removes 93.77%–98.51% of vector-level N : M violations and increases the fraction of conforming graphs from 5%–7% to 79.3%–91.3%. On NVIDIA A100 GPUs, it accelerates SpMM by up to 50.7 × (geometric-mean speedups 3.44 × –7.25 × ) over cuSPARSE, and speeds up key GNN graph operations on real graphs by as much as 6.7 × (2.6 × on average). Ruifeng Zhang 0008, Jou-An Chen, Hsin-Hsuan Sung, Ang Li 0006, Xipeng Shen |
J. Parallel Distributed Comput. | 2 |
| 2025 | Accelerating GNNs on GPU Sparse Tensor Cores through N: M Sparsity-Oriented Graph ReorderingabstractRecent GPUs have introduced Sparse Tensor Cores (SPTC) to accelerate computations on sparse matrices meeting the N:M sparse patterns. Software tools expand the support to more general V:N:M patterns. Graphs in Graph Neural Networks (GNNs) are typically sparse, but the sparsity is often irregular, not conforming to the required V:N:M sparse patterns. This paper proposes a novel graph reordering algorithm to transform irregular graph data into the required sparse patterns for GNNs to benefit from SPTC. The optimization is lossless, maintaining the accuracy of GNN. It at the same time keeps the symmetry of the adjacency matrices of the graphs so that the same matrices can remain compatible with many symmetry-based graph algorithms. The optimization successfully removes 98-100% violations of the N:M sparse patterns at the vector level and increases the portion of conforming graphs in the SuiteSparse collection from 5-9% to 88.7-93.5%. On A100 GPUs, the optimization accelerates Sparse Matrix Matrix (SpMM) by up to 43X (a geomean speedup of 2.3X - 7.5X) over cuSPARSE and speeds up the key graph operations in GNNs on real graphs by as much as 8.6X (3.5X on average). Jou-An Chen, Hsin-Hsuan Sung, Ruifeng Zhang 0008, Ang Li 0006, Xipeng Shen |
PPoPP | 1 |
| 2025 | Mobile-3DCNN: An Acceleration Framework for Ultra-Real-Time Execution of Large 3D CNNs on Mobile DevicesabstractIt is challenging to deploy 3D Convolutional Neural Networks (3D CNNs) on mobile devices, specifically if both real-time execution and high inference accuracy are in demand, because the increasingly large model size and complex model structure of 3D CNNs usually require tremendous computation and memory resources. Weight pruning is proposed to mitigate this challenge. However, existing pruning is either not compatible with modern parallel architectures, resulting in long inference latency or subject to significant accuracy degradation. This article proposes an end-to-end 3D CNN acceleration framework based on pruning/compilation co-design called Mobile-3DCNN that consists of two parts: a novel, fine-grained structured pruning enhanced by a prune/Winograd adaptive selection (that is mobile-hardware-friendly and can achieve high pruning accuracy), and a set of compiler optimization and code generation techniques enabled by our pruning (to fully transform the pruning benefit to real performance gains). The evaluation demonstrates that Mobile-3DCNN outperforms state-of-the-art end-to-end DNN acceleration frameworks that support 3D CNN execution on mobile devices, Alibaba Mobile Neural Networks and Pytorch-Mobile with speedup up to 34× with minor accuracy degradation, proving it is possible to execute high-accuracy large 3D CNNs on mobile devices in real-time (or even ultra-real-time). Wei Niu 0002, Mengshu Sun, Zhengang Li 0001, Jou-An Chen, Jiexiong Guan, Xipeng Shen, Jun Liu 0075, Yanzhi Wang 0001, Xue Lin 0001, Bin Ren 0002 |
ACM Trans. Archit. Code Optim. | 4 |
| 2023 | BitGNN: Unleashing the Performance Potential of Binary Graph Neural Networks on GPUsabstractRecent studies have shown that Binary Graph Neural Networks (GNNs) are promising for saving computations of GNNs through binarized tensors. Prior work, however, mainly focused on algorithm designs or training techniques, leaving it open to how to materialize the performance potential on accelerator hardware fully. This work redesigns the binary GNN inference backend from the efficiency perspective. It fills the gap by proposing a series of abstractions and techniques to map binary GNNs and their computations best to fit the nature of bit manipulations on GPUs. Results on real-world graphs with GCNs, GraphSAGE, and GraphSAINT show that the proposed techniques outperform state-of-the-art binary GNN implementations by 8-22X with the same accuracy maintained. BitGNN code is publicly available.1. Jou-An Chen, Hsin-Hsuan Sung, Xipeng Shen, Sutanay Choudhury, Ang Li 0006 |
ICS | 1 |
| 2023 | Decentralized Application-Level Adaptive Scheduling for Multi-Instance DNNs on Open Mobile Devices
Hsin-Hsuan Sung, Jou-An Chen, Wei Niu 0002, Jiexiong Guan, Bin Ren 0002, Xipeng Shen |
USENIX ATC | 2 |
| 2023 | Accelerating matrix-centric graph processing on GPUs through bit-level optimizations
Jou-An Chen, Hsin-Hsuan Sung, Xipeng Shen, Nathan R. Tallent, Kevin J. Barker, Ang Li 0006 |
J. Parallel Distributed Comput. | 1 |
| 2022 | Bit-GraphBLAS: Bit-Level Optimizations of Matrix-Centric Graph Processing on GPUabstractIn a general graph data structure like an adjacency matrix, when edges are homogeneous, the connectivity of two nodes can be sufficiently represented using a single bit. This insight has, however, not yet been adequately exploited by the existing matrix-centric graph processing frameworks. This work fills the void by systematically exploring the bit-level representation of graphs and the corresponding optimizations to the graph operations. It proposes a two-level representation named Bit-Block Compressed Sparse Row (B2SR) and presents a series of optimizations to the graph operations on B2SR by leveraging the intrinsics of modern GPUs. Evaluations on NVIDIA Pascal and Volta GPUs show that the optimizations bring up to 40× and 6555× for essential GraphBLAS kernels SpMV and SpGEMM, respectively, making GraphBLAS-based BFS accelerate up to 433×, SSSP, PR, and CC up to 35×, and TC up to 52×. Jou-An Chen, Hsin-Hsuan Sung, Xipeng Shen, Nathan R. Tallent, Kevin J. Barker, Ang Li 0006 |
IPDPS | 1 |
| 2021 | RT3D: Achieving Real-Time Execution of 3D Convolutional Neural Networks on Mobile DevicesabstractMobile devices are becoming an important carrier for deep learning tasks, as they are being equipped with powerful, high-end mobile CPUs and GPUs. However, it is still a challenging task to execute 3D Convolutional Neural Networks (CNNs) targeting for real-time performance, besides high inference accuracy. The reason is more complex model structure and higher model dimensionality overwhelm the available computation/storage resources on mobile devices. A natural way may be turning to deep learning weight pruning techniques. However, the direct generalization of existing 2D CNN weight pruning methods to 3D CNNs is not ideal for fully exploiting mobile parallelism while achieving high inference accuracy. This paper proposes RT3D, a model compression and mobile acceleration framework for 3D CNNs, seamlessly integrating neural network weight pruning and compiler code generation techniques. We propose and investigate two structured sparsity schemes i.e., the vanilla structured sparsity and kernel group structured (KGS) sparsity that are mobile acceleration friendly. The vanilla sparsity removes whole kernel groups, while KGS sparsity is a more fine-grained structured sparsity that enjoys higher flexibility while exploiting full on-device parallelism. We propose a reweighted regularization pruning algorithm to achieve the proposed sparsity schemes. The inference time speedup due to sparsity is approaching the pruning rate of the whole model FLOPs (floating point operations). RT3D demonstrates up to 29.1x speedup in end-to-end inference time comparing with current mobile frameworks supporting 3D CNNs, with moderate 1%~1.5% accuracy loss. The end-to-end inference time for 16 video frames could be within 150 ms, when executing representative C3D and R(2+1)D models on a cellphone. For the first time, real-time execution of 3D CNNs is achieved on off-the-shelf mobiles. Wei Niu 0002, Mengshu Sun, Zhengang Li 0001, Jou-An Chen, Jiexiong Guan, Xipeng Shen, Yanzhi Wang 0001, Sijia Liu 0001, Xue Lin 0001, Bin Ren 0002 |
AAAI | 4 |
| 2019 | Finding Documents Related to Taiwan in the Veritable Records of Qing Using Relevance Feedback
Hsin-Hsuan Sung, Jou-An Chen, Jieh Hsiang |
TPDL | 2 |
| 2018 | Exploring the relationships between EFL learners' choices of multimedia and their approaches to learning English
Jou-An Chen, Ching-Fang Juan, Jyh-Chong Liang |
ICCE | 1 |