EDBT 2026 Demo / reviewers in the wild / expert
Zhaohui Xu
dblp:119/6058
· DBLP profile ↗
8ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | COMET: Towards Practical W4A4KV4 LLMs ServingabstractQuantization is a widely-used compression technology to reduce the overhead of serving large language models (LLMs) on terminal devices and in cloud data centers. However, prevalent quantization methods, such as 8-bit weight-activation or 4-bit weight-only quantization, achieve limited performance improvements due to poor support for low-precision (e.g., 4-bit) activation. This work, for the first time, realizes practical W4A4KV4 serving for LLMs, fully utilizing the INT4 tensor cores on modern GPUs and reducing the memory bottleneck caused by the KV cache. Specifically, we propose a novel fine-grained mixed-precision quantization algorithm (FMPQ) that compresses most activations into 4-bit with negligible accuracy loss. To support mixed-precision matrix multiplication for W4A4 and W4A8, we develop a highly optimized W4Ax kernel. Our approach introduces a novel mixed-precision data layout to facilitate access and fast dequantization for activation and weight tensors, utilizing the GPU's software pipeline to hide the overhead of data loading and conversion. Additionally, we propose fine-grained streaming multiprocessor (SM) scheduling to achieve load balance across different SMs. We integrate the optimized W4Ax kernel into our inference framework, COMET, and provide efficient management to support popular LLMs such as LLaMA-3-70B. Extensive evaluations demonstrate that, when running LLaMA family models on a single A100-80G-SMX4, COMET achieves a kernel-level speedup of 2.88x over cuBLAS and a 2.02x throughput improvement compared to TensorRT-LLM from an end-to-end framework perspective. Long Cheng 0003, Haimeng Ren, Zhaohui Xu, Yudong Pan, Mengdi Wang 0004, Xiaowei Li 0001, Yinhe Han 0001, Ying Wang 0001 |
ASPLOS (2) | 4 |
| 2025 | Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMMabstractThe billion-scale Large Language Models (LLMs) necessitate deployment on expensive server-grade GPUs with large-storage HBMs and abundant computation capability. As LLM-assisted services become popular, achieving cost-effective LLM inference on budget-friendly hardware becomes the current trend. This has sparked extensive research into relocating LLM parameters from expensive GPUs to external host memory. However, the restricted bandwidth between the host and GPU memory limits the inference performance of existing solutions. This work introduces Hermes, a budget-friendly system that leverages the near-data processing units (NDP) within commodity DRAM DIMMs to enhance the performance of a single consumer-grade GPU, achieving efficient LLM inference. We recognize that the inherent activation sparsity in LLMs naturally divides weight parameters into two categories, termed “hot” and “cold” neurons, respectively. Hot neurons, which consist of only approximately 20% of all weight parameters, account for 80% of the total computational load. In contrast, cold neurons make up the other 80% of parameters but are responsible for just 20% of the computational workload. Leveraging this observation, we propose a heterogeneous computing strategy: mapping hot neurons to a single computation-efficient GPU without large-capacity HBMs, while offloading cold neurons to NDP-DIMMs, which offer large memory size but limited computation capabilities. In addition, the dynamic nature of activation sparsity necessitates a real-time partition of hot and cold neurons and adaptive remapping of cold neurons across multiple NDP-DIMM modules. To tackle these issues, we introduce a lightweight predictor that ensures optimal real-time neuron partition and adjustment between GPU and NDP-DIMMs. Furthermore, we utilize a window-based online scheduling mechanism to maintain load balance among multiple NDP-DIMM modules. In summary, Hermes facilitates the deployment of LLaMA2-70B on consumer-grade hardware at a rate of 13.75 tokens/s and realizes an average 75.24 × speedup over the state-of-the-art offloading-based inference system on popular LLMs. Bing Li 0017, Haimeng Ren, Zhaohui Xu, Mengdi Wang 0004, Xiaowei Li 0001, Yinhe Han 0001, Ying Wang 0001 |
HPCA | 5 |
| 2024 | Drift: Leveraging Distribution-based Dynamic Precision Quantization for Efficient Deep Neural Network AccelerationabstractQuantization is one of the most hardware-efficient ways to reduce inference costs for deep neural network (DNN) models. Nevertheless, with the continuous increase of DNN model sizes (240× in two years) and the emergence of large language models, existing static quantization methods fail to utilize the sparsity and redundancy of models sufficiently. Motivated by the pervasive dynamism in data tensors across DNN models, we propose a dynamic precision quantization algorithm to further reduce computational costs beyond statically quantized DNN models. Furthermore, we find that existing precision-flexible accelerators cannot support the DNN models with dynamic precision. To this end, we design a novel accelerator, Drift, and achieve online scheduling to efficiently support dynamic precision execution. We conduct experiments with various DNN models, including CNN-based and Transformer-based models. Evaluation results show that Drift achieves 2.85× speedup and 3.12× energy saving compared to existing precision-flexible accelerators with statically quantized models. Zhaohui Xu, Yintao He, Ying Wang 0001, Huawei Li 0001, Xiaowei Li 0001, Yinhe Han 0001 |
DAC | 2 |
| 2024 | SMOTE-kTLNN: A hybrid re-sampling method based on SMOTE and a two-layer nearest neighbor classifier
Liyan Jia, Zhaohui Xu |
Expert Syst. Appl. | 4 |
| 2024 | Seismic Facies Classification Using Label-Integrated and VMD-Augmented TransformerabstractSeismic facies classification is crucial in interpreting subsurface geological structures for oil and gas exploration. Traditional methods for seismic facies classification usually rely on handcrafted features and heuristic rules, limiting their ability to capture complex geological patterns. We suggest a label-integrated and VMD-augmented transformer (LIVAT) to address these issues, which refers to transformers’ embedding stage infused with seismic facies labels and uses variational mode decomposition (VMD) to augment training data. Four advanced time-series transformer models are selected to verify the effectiveness of our proposed label-integrated embedding and VMD-augmentation on F3 Netherlands and New Zealand Parihaka datasets. Moreover, we evaluate the performance of LIVAT by comparing it with convolutional neural networks (CNNs) and bidirectional long short-term networks (BiLSTM) in few-shot learning. Experimental results demonstrate that LIVAT achieves superior classification accuracy and outperforms existing deep learning methods, showcasing its potential as a powerful tool for automatic seismic facies interpretation. Jinlong Huo, Naihao Liu, Zhaohui Xu, Xinguang Wang, Jinghuai Gao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | TDMO: Dynamic multi-dimensional oversampling for exploring data distribution based on extreme gradient boosting learning
Liyan Jia, Zhaohui Xu |
Inf. Sci. | 4 |
| 2018 | Yixue Adaptive Learning System and Its Promise on Improving Student Learning
Zhaohui Xu, Zhenyue Zhu, Mingyu Feng |
CSEDU (2) | 3 |
| 2017 | Quantification of microbial species in solid state fermentation samples using signature genomic sequencesabstractSolid state fermentation processes are mediated by the collective metabolism of specialized microbial communities. Monitoring the relative abundance of dominating species is a critical task in quality control, which is traditionally done by wet lab techniques, such as quantitative PCR (qPCR). In this study, we developed a computational method to quantify microbial species in metagenomes based on their signature genomic sequences, i.e., unique k-mers. Bacterial species found in fermentation starters of a Chinese liquor producer were used as examples to demonstrate the development and application of the method. A database was constructed, comprising 562 complete genome sequences of 93 bacterial species that had been found in relevant fermentation samples. K-mers in length of 12 were extracted from each species and compared against each other to identify the ones that were unique to each species. The quantity of a species was determined by the average frequencies of unique k-mers encountered in the metagenome. Six dominating bacterial species were chosen as reporter species to test the quantification method. Four metagenome datasets were simulated, which contained various portions of sequence reads generated from the genomes of the reporter species. The amount of reads sampled from each reporter species followed a pre-determined ratio, i.e., a known relationship in relative abundance. For each simulated dataset, the cell number of each reporter species was computed based on the unique k-mers found in the metagenome. In all datasets, the computed quantities of the reporter species reflected the expected relative abundance by displaying a linear relationship with the pre-determined ratio. This demonstrates that quantification based on a set of unique k-mers is a reliable way to detect relative abundance among species. Besides industrial fermentation, this method may also be applied to areas such as wastewater treatment, microbiota analysis, etc. Zhaohui Xu, Sankardas Roy |
BIBM | 1 |