EDBT 2026 Demo / reviewers in the wild / expert
Xizi Chen
dblp:209/9829
· DBLP profile ↗
11ranked-venue papers
6as first author
6since 2021 · last 2026
0000-0001-8155-6606ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 6 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploiting the Irregular Input Sparsity in Systolic Array-based DNN Accelerators via Local Soft PoolingabstractOne promising approach to mitigating the computational complexity of deep neural networks is to leverage the sparsity of input activations that results from the application of the ReLU function. However, the irregular distribution of zero-valued inputs poses a challenge for efficient implementation in existing regular architectures, such as systolic arrays. Previous works usually depend on specialized architectures to bypass the redundant computations during runtime. In contrast to these prior strategies, we propose a local soft pooling method to efficiently exploit the irregular input sparsity in systolic array-based architectures. Through local soft pooling, adjacent input rows can be safely merged at runtime, compressing the sparse input matrix into a compact format that is only 1/3 to 1/2 of its original size. The compact matrix can then be directly fed into the systolic array for computation. A computation saving of 67.78% is achieved across various networks on both CIFAR-10 and ImageNet with negligible accuracy loss. As a result, the throughput and energy efficiency are improved by 2.72 and 2.07 times, respectively. Desheng Fu, Jingbo Jiang, Jingyang Zhu, Xizi Chen, Chi-Ying Tsui |
ASP-DAC | 5 |
| 2025 | MA-HybridBTS: Modality-Aware Multi-Scale Hybrid 3D Conv-Transformer for Brain Tumor SegmentationabstractBrain tumor segmentation (BTS) on magnetic resonance imaging (MRI) is essential for diagnosis yet remains challenging due to substantial inter-patient variations and the need to integrate cross-modal interactions. In this work, we propose a modality-aware BTS framework, MA-HybridBTS, to address these challenges through three key components. First, we introduce a modality-aware multi-scale feature extraction (MAMS) strategy that employs 3D dilated convolutions with tailored dilation rates to each sequence, concurrently refining tumor-core details while expanding global context. Second, we present a modality-aware hierarchical fusion (MAHF) module that explicitly leverages clinically established inter-modal dependencies to guide progressive feature fusion. Finally, we propose an adaptive loss function (ALF) that dynamically up-weights sub-regions with larger losses, thereby improving segmentation accuracy for smaller and harder-to-delineate tumor compartments. The proposed methods achieve state-of-the-art performance on the BraTS2021 dataset, delivering an average Dice score of 0.8984 and an average 95% Hausdorff distance (HD95) of 4.14 mm. Ablation studies further validate the effectiveness of each proposed component. The code is publicly available on GitHub: https://github.com/Jia7888/code. Shiqi Miao, Xizi Chen, Lida Zhu |
BIBM | 4 |
| 2025 | MSU3D: Multi-Scale 3D Convolutional Neural Network for Lung Nodule SegmentationabstractPrecise segmentation of pulmonary nodules is essential for the early detection of lung cancer, yet remains challenging due to the substantial diversity in nodule size, shape, and density. To tackle this problem, we propose a lightweight multi-scale 3D U-shaped convolutional network named MSU3D. It embeds a pooling-enhanced multi-scale extraction block (MSE3D) to enlarge receptive fields without extra parameters, and a channelwise 3D fusion module (MSF3D) that interleaves and re-weights cross-scale features for rich yet compact representations. The performance of the proposed method is evaluated on LUNA16 using standard 5-fold cross-validations. Experimental results show that MSU3D delivers competitive results, raising DSC to$\mathbf{9 1. 3 4 \%}$, PPV to$\mathbf{9 2. 9 5 \%}$, and SEN to$\mathbf{9 2. 0 7 \%}$, respectively. Mengtong Wu, Xulei Zhao, Lida Zhu, Xizi Chen |
BIBM | 5 |
| 2023 | Late Breaking Results: Weight Decay is ALL You Need for Neural Network SparsificationabstractThe heuristic iterative pruning strategy has been widely used for neural network sparsification. However, it is challenging to identify the right connections to remove at each pruning iteration with only a one-shot evaluation of weight magnitude, especially at the early pruning stage. The erroneously removed connections, unfortunately, can hardly be recovered. In this work, we propose a weight decay strategy as a substitute for pruning, which let the "insignificant" weights moderately decay instead of being directly clamped to zero. At the end of the training, the vast majority of redundant weights will naturally become close to zero, making it easier to identify which connections could be removed safely. Experimental results show that the proposed weight decay method can achieve an ultra-high sparsity of 99%. Compared to the current pruning strategy, the model size is further reduced by 34%, improving the compression rate from 69× to 106× at the same accuracy. Xizi Chen, Fengshi Tian, Chi-Ying Tsui |
DAC | 1 |
| 2023 | Accelerating Large Kernel Convolutions with Nested Winograd TransformationabstractRecent literature has shown that convolutional neural networks (CNNs) with large kernels outperform vision transformers (ViTs) and CNNs with stacked small kernels in many computer vision tasks, such as object detection and image restoration. The Winograd transformation helps reduce the number of repetitive multiplications in convolution and is widely supported by many commercial AI processors. Researchers have proposed accelerating large kernel convolutions by linearly decomposing them into many small kernel convolutions and then sequentially accelerating each small kernel convolution with the Winograd algorithm. This work proposes a nested Winograd algorithm that iteratively decomposes a large kernel convolution into small kernel convolutions and proves it to be more effective than the linear decomposition Winograd transformation algorithm. Experiments show that compared to the linear decomposition Winograd algorithm, the proposed algorithm reduces the total number of multiplications by 1.4 to 10.5 times for computing 4×4 to 31×31 convolutions. Jingbo Jiang, Xizi Chen, Chi-Ying Tsui |
VLSI-SoC | 2 |
| 2023 | Tight Compression: Compressing CNN Through Fine-Grained Pruning and Weight Permutation for Efficient ImplementationabstractThe unstructured sparsity after pruning poses a challenge to the efficient implementation of deep learning models in existing regular architectures like systolic arrays. On the other hand, coarse-grained structured pruning is suitable for implementation in regular architectures but tends to have higher accuracy loss than unstructured pruning when the pruned models are of the same size. In this work, we propose a model compression method based on a novel weight permutation scheme to fully exploit the fine-grained weight sparsity in the hardware design. Through permutation, the optimal arrangement of the weight matrix is obtained, and the sparse weight matrix is further compressed to a small and dense format to make full use of the hardware resources. Two pruning granularities are explored. In addition to the unstructured weight pruning, we also propose a more fine-grained subword-level pruning to further improve the compression performance. Compared to the state-of-the-art works, the matrix compression rate is significantly improved from$5.88\times $to$14.13\times $. As a result, the throughput and energy efficiency are improved by 2.75 and 1.86 times, respectively. Xizi Chen, Jingyang Zhu, Jingbo Jiang, Chi-Ying Tsui |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2020 | Tight Compression: Compressing CNN Model Tightly Through Unstructured Pruning and Simulated Annealing Based PermutationabstractThe unstructured sparsity after pruning poses a challenge to the efficient implementation of deep learning models in existing regular architectures like systolic arrays. The coarse-grained structured pruning, on the other hand, tends to have higher accuracy loss than unstructured pruning when the pruned models are of the same size. In this work, we propose a compression method based on the unstructured pruning and a novel weight permutation scheme. Through permutation, the sparse weight matrix is further compressed to a small and dense format to make full use of the hardware resources. Compared to the state-of-the-art works, the matrix compression rate is effectively improved from 5.88x to 10.28x. As a result, the throughput and energy efficiency are improved by 2.12 and 1.57 times, respectively. Xizi Chen, Jingyang Zhu, Jingbo Jiang, Chi-Ying Tsui |
DAC | 1 |
| 2019 | CompRRAE: RRAM-based convolutional neural network accelerator with reduced computations through a runtime activation estimationabstractRecently Resistive-RAM (RRAM) crossbar has been used in the design of the accelerator of convolutional neural networks (CNNs) to solve the memory wall issue. However, the intensive multiply-accumulate computations (MACs) executed at the crossbars during the inference phase are still the bottleneck for the further improvement of energy efficiency and throughput. In this work, we explore several methods to reduce the computations for the RRAM-based CNN accelerators. First, the output sparsity resulting from the widely employed Rectified Linear Unit is exploited, and a significant portion of computations are bypassed through an early detection of the negative output activations. Second, an adaptive approximation is proposed to terminate the MAC early when the sum of the partial results of the remaining computations is considered to be within a certain range of the intermediate accumulated result and thus has an insignificant contribution to the inference. In order to determine these redundant computations, a novel runtime estimation on the maximum and minimum values of each output activation is developed and used during the MAC operation. Experimental results show that around 70% of the computations can be reduced during the inference with a negligible accuracy loss smaller than 0.2%. As a result, the energy efficiency and the throughput are improved by over 2.9 and 2.8 times, respectively, compared with the state-of-the-art RRAM-based accelerators. Xizi Chen, Jingyang Zhu, Jingbo Jiang, Chi-Ying Tsui |
ASP-DAC | 1 |
| 2019 | SubMac: Exploiting the subword-based computation in RRAM-based CNN accelerator for energy saving and speedup
Xizi Chen, Jingbo Jiang, Jingyang Zhu, Chi-Ying Tsui |
Integr. | 1 |
| 2018 | A high-throughput and energy-efficient RRAM-based convolutional neural network using data encoding and dynamic quantizationabstractTo solve the scaling, memory wall and high power density issues, recently RRAM-based accelerators, which show a better energy and area efficiency compared with the CMOS-based counterparts, have been proposed for convolutional neural networks. However, the RRAM-based architectures still face several design challenges, including the high energy and timing overhead at the analog/digital (A/D) conversion and interfacing circuits. To address these issues, we propose several novel optimization schemes in this work. First an encoding scheme for the synaptic weights and the input feature maps is proposed to reduce the energy of the in-situ computation and the bit-resolution of the A/D conversion. Then the resolution of the A/D conversion is further optimized for a lower energy consumption. Moreover, a dynamic quantization scheme for the multiply-accumulate operations (MACs) is proposed to improve the throughput and the energy efficiency by reducing the number of partial products. Experimental results show that the throughput, the energy efficiency and the area efficiency are improved by 2 to 4 times when compared with the state-of-the-art RRAM-based accelerators. Xizi Chen, Jingbo Jiang, Jingyang Zhu, Chi-Ying Tsui |
ASP-DAC | 1 |
| 2018 | SparseNN: An energy-efficient neural network accelerator exploiting input and output sparsityabstractContemporary Deep Neural Network (DNN) contains millions of synaptic connections with tens to hundreds of layers. The large computational complexity poses a challenge to the hardware design. In this work, we leverage the intrinsic activation sparsity of DNN to substantially reduce the execution cycles and the energy consumption. An end-to-end training algorithm is proposed to develop a lightweight (less than 5% overhead) run-time predictor for the output activation sparsity on the fly. Furthermore, an energy-efficient hardware architecture, SparseNN, is proposed to exploit both the input and output sparsity. SparseNN is a scalable architecture with distributed memories and processing elements connected through a dedicated on-chip network. Compared with the state-of-the-art accelerators which only exploit the input sparsity, SparseNN can achieve a 10%-70% improvement in throughput and a power reduction of around 50%. Jingyang Zhu, Jingbo Jiang, Xizi Chen, Chi-Ying Tsui |
DATE | 3 |