VLDB 2026 Research / reviewers in the wild / expert
Wenxi Zhu
dblp:144/3549
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Leveraging Negative Sequential Patterns for Advanced Genomic Analysis and ClassificationabstractViruses represent the most abundant biological entities and play pivotal roles in microbial ecosystems. As prominent human pathogens, they are closely linked to human morbidity and mortality. Accurate identification of viral sequences from viral genome sequences is essential, yet existing genome-based mod-els-mainly relying on composition or frequency features-often show limited interpretability and lower accuracy, especially on complex or imbalanced datasets. To address these limitations, we propose GeneNSPCla (Genomic Negative Sequential Patternbased Classification), a novel viral classification framework based on negative sequential patterns (NSPs) that extracts discriminative absence-based features from nucleotide sequences of RNA viral genomes. By transforming these NSPs into numerical feature vectors and integrating them into multiple supervised classifiers, GeneNSPCla effectively captures both presence and absence signals in viral sequences. Furthermore, we propose a negative pattern mining algorithm, GONPM, for processing genomic data, which can discover longer and more biologically meaningful negative sequential patterns. The experimental results demonstrate that the average accuracy of GONPM in 8 classifiers has improved by 5.25 % compared to the original negative pattern mining algorithm and by 19.97 % compared to the positive pattern mining algorithm. These findings highlight the effectiveness of incorporating absence-based sequential information, providing a new and complementary perspective for viral genome analysis and classification. Wenxi Zhu, Wensheng Gan, Zhenlian Qi |
BIBM | 1 |
| 2025 | Large Language Models for Bioinformatics: Applications and Challenges
Wenxi Zhu, Wensheng Gan, Zhenlian Qi, Philip S. Yu |
IEEE Big Data | 1 |
| 2025 | NM-SpMM: Accelerating Matrix Multiplication Using N: M Sparsity with GPGPUabstractDeep learning demonstrates effectiveness across a wide range of tasks. However, the dense and over-parameterized nature of these models results in significant resource consumption during deployment. In response to this issue, weight pruning, particularly through$N: M$sparsity matrix multiplication, offers an efficient solution by transforming dense operations into semisparse ones.$N: M$sparsity provides an option for balancing performance and model accuracy, but introduces more complex programming and optimization challenges. To address these issues, we design a systematic top-down performance analysis model for$N: M$sparsity. Meanwhile, NM-SpMM is proposed as an efficient general$N: M$sparsity implementation. Based on our performance analysis, NM-SpMM employs a hierarchical blocking mechanism as a general optimization to enhance data locality, while memory access optimization and pipeline design are introduced as sparsity-aware optimization, allowing it to achieve close-to-theoretical peak performance across different sparsity levels. Experimental results show that NM-SpMM is 2.1x faster than nmSPARSE (the state-of-the-art for general$N: M$sparsity) and 1.4× to 6.3× faster than cuBLAS's dense GEMM operations, closely approaching the theoretical maximum speedup resulting from the reduction in computation due to sparsity. NM-SpMM is open source and publicly available at https://github.com/M-H482/NM-SpMM. Du Wu, Zhelang Deng, Jintao Meng 0001, Wenxi Zhu, Bingqiang Wang, Amelie Chi Zhou, Peng Chen 0035, Minwen Deng, Yanjie Wei, Shengzhong Feng, Yi Pan 0001 |
IPDPS | 7 |
| 2025 | A Sample-Free Compilation Framework for Efficient Dynamic Tensor ComputationabstractDynamic-shape tensor computation poses challenges for shape-specific compilation due to variable input dimensions. Existing compilers rely on shape samples, incurring high tuning costs and performance degradation on unseen inputs. We present Helix, a dynamic tensor compilation framework with sample-free compilation and architecture-guided optimization to achieve both compilation efficiency and shape-general performance. To avoid shape sampling, Helix constructs shape-agnostic compilation by decomposing computations across architectural layers. A bidirectional strategy combines top-down abstraction to align tensor computations with architectural hierarchies, and bottom-up kernel construction to build efficient execution strategies from reusable, architecture-aligned micro-kernels. A hybrid analyzer ensures accuracy through profiling at lower architectural levels, and achieves scalability through architecture-informed modeling at higher levels and runtime. This hierarchical design eliminates shape-specific tuning and enables shape-adaptive execution. Evaluations conducted on x86 CPUs, ARM CPUs, and NVIDIA GPUs demonstrate that Helix reduces compilation time by 174 × over the existing compilers and delivers 2.26 × and 3.29 × execution speedups over vendor libraries and dynamic-shape compilers, respectively. Yangjie Zhou 0001, Weihao Cui, Zihan Liu 0002, Peng Chen 0035, Mohamed Wahib, Cong Guo 0003, Siyuan Feng 0007, Jintao Meng 0001, Haidong Lan, Jingwen Leng, Yun Lin 0001, Jin Song Dong 0001, Wenxi Zhu, Minwen Deng |
SC | 15 |
| 2024 | Large Model Fine-tuning for Suicide Risk Detection Using Iterative Dual-LLM Few-Shot Learning with Adaptive Prompt Refinement for Dataset ExpansionabstractAn approach to detecting suicide risk in social media posts is presented, addressing the challenges of limited and imbalanced datasets. The proposed workflow combines large language models (LLMs), few-shot learning, and expert-supervised prompt optimization. To address class imbalance, a novel Iterative Dual-LLM Few-Shot Learning with Adaptive Prompt Refinement (IDFL-APR) method for dataset expansion is introduced. This method employs two LLMs: one for classifying posts using few-shot learning, and another for dynamically optimizing prompts, with expert oversight to prevent overfitting. The optimized LLM then identifies high-confidence samples from unclassified data, focusing on underrepresented categories. Back-translation techniques are further applied to enhance textual diversity and achieve dataset balance across all categories. Subsequently, the bloomz-3b model is fine-tuned using optimized hyperparameters, implementing an active learning strategy to iteratively augment the training set. The proposed approach significantly enhances suicide risk detection accuracy. The best-performing model achieved a weighted F1-score of 0.7154 on the test set provided by the Suicide Ideation Detection on Social Media Challenge, a track of the IEEE Big Data 2024 BigData Cup. These results demonstrate a robust solution to the inherent challenges in suicide detection tasks. Jingyun Bi, Wenxi Zhu, Jingyun He, Xinshen Zhang, Chong Xian |
IEEE Big Data | 2 |
| 2024 | autoGEMM: Pushing the Limits of Irregular Matrix Multiplication on Arm ArchitecturesabstractThis paper presents an open-source library that pushes the limits of performance portability for irregular General Matrix Multiplication (GEMM) on the widely-used Arm architectures. Our library, autoGEMM, is designed to support a wide range of Arm processors: from edge devices to HPCgrade CPUs. autoGEMM generates optimized kernels for various hardware configurations by auto-combining fragments of autogenerated micro-kernels that employ hand-written optimizations to maximize computational efficiency. We optimize the kernel pipeline by tuning the register reuse and the data load/store overlapping. In addition, we use a dynamic tiling scheme to generate balanced tile shapes. Finally, we position autoGEMM on top of the TVM framework where our dynamic tiling scheme prunes the search space for TVM to identify the optimal combination of parameters for code optimization. Evaluations on five different classes of Arm chips demonstrate the advantages of autoGEMM. For small matrices, autoGEMM achieves 98% of peak and up to 2.0x speedup over state-of-the-art libraries such as LIBXSMM and LibShalom. For irregular matrices (i.e. tall skinny and long rectangles), autoGEMM is 1.3-2.0x faster than widely-used libraries such as OpenBLAS and Eigen. autoGEMM is publicly available at: https://github.com/wudu98/autoGEMM. Du Wu, Jintao Meng 0001, Wenxi Zhu, Minwen Deng, Xiao Wang 0004, Tao Luo 0014, Mohamed Wahib, Yanjie Wei |
SC | 3 |
| 2022 | Efficient Phase-Functioned Real-time Character Control in Mobile Games: A TVM Enabled ApproachabstractIn this paper, we propose a highly efficient computing method for game character control with phase-functioned neural networks (PFNN). The primary challenge to accelerate PFNN on mobile platforms is that PFNN dynamically produces weight matrices with an argument, phase, which is individual to each game character. Therefore existing libraries that generally assume frozen weight matrices are inefficient to accelerate PFNN. The situation becomes even worse when multiple characters are present. To address the challenges, we reformulate the equations and leverage the deep learning compiler stack TVM to build a cross-platform, high-performance implementation. Evaluations reveal that our solutions deliver close-to-peak performance on various platforms, from high-performance servers to energy-efficient mobile platforms. This work is publicly available at https://github.com/turbo0628/pfnn_tvm. Haidong Lan, Wenxi Zhu, Du Wu, Xinghui Fu, Liu Wei, Jintao Meng 0001, Minwen Deng |
ICPP | 2 |