EDBT 2026 Demo / reviewers in the wild / expert
Xuanpeng Zhu
dblp:360/9884
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2026
0009-0000-4859-9604ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STP: Semantic-Triggered Prefetching for Event-Driven WorkloadsabstractHardware prefetchers such as SPP rely on address-delta history and perform poorly on event-driven workloads, where event-type switches invalidate recent patterns. On an HFT order-book engine, SPP achieves only 8.0% L2 prefetch accuracy with 85.9% late prefetches; on B+Tree under uniform-random access, it degrades IPC by 6.7–8.1%. We present STP (Semantic-Triggered Prefetcher), which exposes event type through one non-privileged x86 hint instruction, SETHINT imm8, and combines two lightweight mechanisms: an Event Footprint Table (EFT) that replays high-frequency missed lines at type transitions, and a PC-Localized Working Set (PLWS) that gates next-line prefetching by per-PC miss rate with asymmetric feedback throttling. In gem5 on one HFT and four B+Tree settings, STP achieves up to 10 × higher L2 prefetch accuracy with 3–10 × fewer requests (up to 90% less bandwidth). On HFT, STP improves IPC by 36.1% over NoPF and by 5.4% over SPP (0.543 vs. 0.515). Under low locality, STP limits impact to − 2.0% or +0.8%, where SPP drops 6.7–8.1%. Hardware cost is 5.4 KB, reducible to ~3.1 KB. Shichen Peng, Yupeng Gui, Han He, Zhengyang Cao, Ruiqi Tang, Xuanpeng Zhu, Xiaoyang Zeng, Yibo Fan |
ACM Great Lakes Symposium on VLSI | 7 |
| 2025 | BLOOM: Bit-Slice Framework for DNN Acceleration with Mixed-PrecisionabstractDeep neural networks (DNNs) have revolutionized numerous AI applications, but their vast model sizes and limited hardware resources present significant deployment challenges. Model quantization offers a promising solution to bridge the gap between DNN size and hardware capacity. While INT8 quantization has been widely used, recent research has pushed for even lower precision, such as INT4. However, the presence of outliers-values with unusually large magnitudes-limits the effectiveness of current quantization techniques. Previous compression-based acceleration methods that incorporate outlieraware encoding introduce complex logic. A critical issue we have identified is that serialization and deserialization dominate the encoding/decoding time in these compression workflows, leading to substantial performance penalties during workflow execution. To address this challenge, we introduce a novel computing approach and a compatible architecture design named “BLOOM”. BLOOM leverages the strengths of the “bit-slicing” method, effectively combining structured mixed-precision and bit-level sparsity with adaptive dataflow techniques. The key insight of BLOOM is that outliers require higher precision, while normal values can be processed at lower precision. By interleaving 4-bit values, we efficiently exploit the inherent sparsity in the highprecision components. As a result, the BLOOM-based accelerator outperforms the existing outlier-aware accelerators by an average $1.2 \sim 4.0 \times$ speedup and $24.6 \% \sim 71.3 \%$ energy reduction, respectively, without model accuracy loss. Fangxin Liu, Ning Yang 0012, Zongwu Wang, Xuanpeng Zhu, Haidong Yao, Xiankui Xiong, Li Jiang 0002, Haibing Guan |
DAC | 4 |
| 2025 | OPS: Outlier-Aware Precision-Slice Framework for LLM AccelerationabstractLarge language models (LLMs) have transformed numerous AI applications, with on-device deployment becoming increasingly important for reducing cloud computing costs and protecting user privacy. However, the astronomical model size and limited hardware resources pose significant deployment challenges. Model quantization is a promising approach to mitigate this gap, but the presence of outliers in LLMs reduces its effectiveness. Previous efforts addressed this issue by employing compression-based encoding for mixed-precision quantization. These approaches struggle to balance model accuracy with hard-ware efficiency due to their value-wise outlier granularity and complex encoding/decoding hardware logic. To address this, we propose OPS (Outlier-aware Precision-Slicing), an acceleration framework that exploits massive sparsity in the higher-order part of LLMs by splitting 16-bit values into a 4-bit/12-bit format. Crucially, OPS introduces an early bird mechanism that leverages the high-order 4-bit computation to predict the importance of the full calculation result. This mechanism enables efficient computational skips by continuing execution only for important computations and using preset values for less significant ones. This scheme can be efficiently integrated with existing hardware accelerators like systolic arrays without complex encoding/decoding. As a result, OPS outperforms state-of-the-art outlier-aware accelerators, achieving a 1.3 − 4.3× performance boost with minimal model accuracy loss. This approach enables more efficient on-device LLM deployment, effectively balancing computational efficiency and model accuracy. Fangxin Liu, Ning Yang 0012, Zongwu Wang, Xuanpeng Zhu, Haidong Yao, Xiankui Xiong, Li Jiang 0002 |
DATE | 4 |
| 2024 | Multi-Weather Degradation-Aware Transformer for Image RestorationabstractRestoring images under different adverse weather conditions with a single model is practical in many applications. Most existing weather restoration approaches are only able to handle a specific type of degradation, which is often insufficient in real-world scenarios where the weather type is unknown. In this paper, we propose a holistic solution to solve multiple weather degradations using a single model. Specifically, we build a weather-type aware Transformer, an efficient architecture that can restore images degraded by different adverse weathers with the same set of parameters. For model training, we first use contrastive loss to train an auxiliary hypernetwork capable of extracting content-independent, distortion-aware feature embeddings. Guided by these weather-dependent features, the image restoration Transformer can adaptively modulate its parameters using hypernetworks and feature-wise linear modulation blocks, conducting both local and global operations adaptively for images with different degradations. Qualitative and quantitative results on the multi-weather benchmark demonstrate that our model achieves significant improvements compared with previous state-of-the-arts, with even less computational cost. Ruoxi Zhu, Minfeng Wu, Xiankui Xiong, Xuanpeng Zhu, Yibo Fan |
ICASSP | 4 |
| 2024 | SFFTNet: Sparse Feature Fusion Transformer Network for Image DeblurringabstractThe U-Net structure, with its an encoder-decoder architecture, has been widely adopted by many deep learning methods for image deblurring. Most methods concentrate on the design of encoder and decoder block and use skip connections to connect them. However, this simple skip connection strategy is insufficient to fully exploit the correlation of multi-scale features, which can result in a potential loss of deblurring performance. To address this issue, we design an effective Sparse Feature Fusion Transformer capable of integrating multi-scale features to replace the skip connections in the U-Net framework. Specifically, We propose a cross-attention mechanism with a learnable top-k selection operator to adaptively preserve highly correlated cross-attention values for feature fusion. This approach ensures that the fused multi-scale features effectively integrate contextual information, resulting in high-quality image deblurring. Additionally, we introduce the position encoding generator scheme to ensure that our deblurring network can be applied to images of any size. Comprehensive experimental results demonstrate that our proposed method outperforms the the state-of-the-art methods. Faxing Lei, Ming-e Jing, Xiankui Xiong, Xuanpeng Zhu, Yibo Fan |
ISCAS | 6 |
| 2023 | Luminance-Preserving Visible and Near-Infrared Image Fusion Network with Edge GuidanceabstractNear-infrared (NIR) images and visible (VIS) images can provide mutually complementary information for each other, thus the fusion of the two modalities can create images of high quality even in adverse conditions. However, the luminance of NIR and VIS images may be inconsistent in some regions, resulting in color distortion and unrealistic appearance in the fused images. The existing methods perform poorly at luminance retention. Aiming at the problem and based on deep learning framework, we propose an edge-guided method which can be applied to the image fusion network. Edge maps are utilized as prior knowledge of images to boost the performance of the neural network. Additionally, we propose a luminance-preserving loss function combined with max-edge loss to further improve the image quality. Experimental results show the superiority of our method. Ruoxi Zhu, Yi Ling, Xiankui Xiong, Dong Xu 0015, Xuanpeng Zhu, Yibo Fan |
ICIP | 5 |