Bowen Duan 0003

dblp:155/8161-3 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0004-9085-5025ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 FractalCloud: A Fractal-Inspired Architecture for Efficient Large-Scale Point Cloud Processing
abstract
Three-dimensional (3D) point clouds are increasingly used in applications such as autonomous driving, robotics, and virtual reality (VR). Point-based neural networks (PNNs) have demonstrated strong performance in point cloud analysis, originally targeting small-scale inputs. However, as PNNs evolve to process large-scale point clouds with hundreds of thousands of points, all-to-all computation and global memory access in point cloud processing introduce substantial overhead, causing$O\left(n^{2}\right)$computational complexity and memory traffic where$n$is the number of points. Existing accelerators, primarily optimized for small-scale workloads, overlook this challenge and scale poorly due to inefficient partitioning and non-parallel architectures. To address these issues, we propose FractalCloud, a fractal-inspired hardware architecture for efficient large-scale 3D point cloud processing. FractalCloud introduces two key optimizations: (1) a co-designed Fractal method for shape-aware and hardware-friendly partitioning, and (2) block-parallel point operations that decompose and parallelize all point operations. A dedicated hardware design with on-chip fractal and flexible parallelism further enables fully parallel processing within limited memory resources. Implemented in 28 nm technology as a chip layout with a core area of$1.5 ~\text{mm}^{2}$, FractalCloud achieves$21.7 \times$speedup and$27 \times$energy reduction over state-of-the-art accelerators while maintaining network accuracy, demonstrating its scalability and efficiency for PNN inference. The code for FractalCloud is available at https://github.com/Yuzhe-Fu/FractalCloud.
Yuzhe Fu, Changchun Zhou 0001, Hancheng Ye, Bowen Duan 0003, Qiyu Huang, Chiyue Wei, Cong Guo 0003, Hai Li 0001, Yiran Chen 0001
HPCA4
2026 EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
Bowen Duan 0003, Cong Guo 0003, Chiyue Wei, Haoxuan Shan, Yuzhe Fu, Xinhua Chen, Changchun Zhou 0001, Hai Li 0001, Yiran Chen 0001
ISCA1
2025 Transitive Array: An Efficient GEMM Accelerator with Result Reuse
abstract
Deep Neural Networks (DNNs) and Large Language Models (LLMs) have revolutionized artificial intelligence, yet their deployment faces significant memory and computational challenges, especially in resource-constrained environments.Quantization techniques have mitigated some of these issues by reducing data precision, primarily focusing on General Matrix Multiplication (GEMM).This study introduces a novel sparsity paradigm, transitive sparsity, which leverages the reuse of previously computed results to substantially minimize computational overhead in GEMM operations.By representing transitive relations using a directed acyclic graph, we develop an efficient strategy for determining optimal execution orders, thereby overcoming inherent challenges related to execution dependencies and parallelism.Building on this foundation, we present the Transitive Array, a multiplication-free accelerator designed to exploit transitive sparsity in GEMM.Our architecture effectively balances computational workloads across multiple parallel lanes, ensuring high efficiency and optimal resource utilization.Comprehensive evaluations demonstrate that the Transitive Array achieves approximately 7.46× and 3.97× speedup and 2.31× and 1.65× energy reduction compared to state-of-the-art accelerators such as Olive and BitVert while maintaining comparable model accuracy on LLaMA models.
Cong Guo 0003, Chiyue Wei, Bowen Duan 0003, Song Han 0003, Hai Li 0001, Yiran Chen 0001
ISCA4
2025 Phi: Leveraging Pattern-based Hierarchical Sparsity for High-Efficiency Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) are gaining attention for their energy efficiency and biological plausibility, utilizing 0-1 activation sparsity through spike-driven computation.While existing SNN accelerators exploit this sparsity to skip zero computations, they often overlook the unique distribution patterns inherent in binary activations.In this work, we observe that particular patterns exist in spike activations, which we can utilize to reduce the substantial computation of SNN models.Based on these findings, we propose a novel pattern-based hierarchical sparsity framework, termed Phi, to optimize computation.Phi introduces a two-level sparsity hierarchy: Level 1 exhibits vector-wise sparsity by representing activations with pre-defined patterns, allowing for offline pre-computation with weights and significantly reducing most runtime computation.Level 2 features element-wise sparsity by complementing the Level 1 matrix, using a highly sparse matrix to further reduce computation while maintaining accuracy.We present an algorithm-hardware co-design approach.Algorithmically, we employ a k-means-based pattern selection method to identify representative patterns and introduce a pattern-aware fine-tuning technique to enhance Level 2 sparsity.Architecturally, we design Phi, a dedicated hardware architecture that efficiently processes the two levels of Phi sparsity on the fly.Extensive experiments demonstrate that Phi achieves a 3.45× speedup and a 4.93× improvement in energy efficiency compared to stateof-the-art SNN accelerators, showcasing the effectiveness of our framework in optimizing SNN computation.
Chiyue Wei, Bowen Duan 0003, Cong Guo 0003, Jingyang Zhang, Qingyue Song, Hai Li 0001, Yiran Chen 0001
ISCA2
2025 Bi-Dynamic Graph ODE for Opinion Evolution
abstract
Modeling opinion dynamics in social networks has been the focus of multiple disciplines in recent decades. Previous studies have often modeled the opinion dynamics as a discrete and homogeneous process, neglecting its continuous and complex nature. To fill this gap, we propose a Bi-Dynamics Graph Ordinary Differential Equation (BDG-ODE) framework, which models complex opinion dynamics as the result of two dynamical processes: the evolution of positive and negative opinions. The proposed model incorporates a dual opinion encoder that processes positive and negative opinions independently. Furthermore, the temporal opinion evolution is modeled through bidirectional graph ordinary differential equations, which allows the model to capture the changes in opinion in continuous time. We introduce an opinion synthesis decoder that effectively maps the evolved representations from the latent space back to the opinion space. Extensive experiments conducted on six datasets with varying characteristics highlight the superiority of BDG-ODE in forecasting opinion evolution within social networks. It achieved an average accuracy improvement of 23.16%, an average enhancement of 29.46% in the F1 score, and an average mean square error of difference improvement of 90. 30%, and an average correlation coefficient improvement of 45.93%, significantly outperforming eight state-of-the-art models. The code for reproduction is available: https://github.com/tsinghua-fib-lab/Bi-Dynamic-Graph-ODE-for-Opinion-Evolution.
Bowen Duan 0003, Henggang Deng, Jinghua Piao, Huandong Wang, Yue Wang 0007
KDD (1)1