VLDB 2026 Research / reviewers in the wild / expert
Bingqiang Wang
dblp:41/3365
· DBLP profile ↗
22ranked-venue papers
0as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5Databases, data management, data science and information retrieval · 3Artificial intelligence and machine learning · 1Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Synchronized Dual-Ring: A Synergistic Algorithm for Bandwidth-Efficient Collective Communication Leveraging Die-to-Die Direct Interconnects
Dongxiang Zhang, Erkun Zhang, Shixun Zhang, Bingqiang Wang, Fangjiong Chen |
IPDPS | 5 |
| 2025 | Using Analytical Performance/Power Model and Fine-Grained DVFS to Enhance AI Accelerator Energy EfficiencyabstractRecent advancements in deep learning have significantly increased AI processors' energy consumption, which is becoming a critical factor limiting AI development. Dynamic Voltage and Frequency Scaling (DVFS) stands as a key method in power optimization. However, due to the latency of DVFS control in AI processors, previous works typically apply DVFS control at the granularity of a program's entire duration or sub-phases, rather than at the level of AI operators. Yijia Zhang 0002, Fuchun Wei, Bingqiang Wang, Yanlin Liu, Zhiheng Hu, Xiaoxin Xu, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen, Chen Tian 0001 |
ASPLOS (1) | 4 |
| 2025 | Improving the Energy Efficiency of AI Clusters Through Variability-Aware Frequency Scaling and Task Allocation
Dongxiang Zhang, Bingqiang Wang, Qiang Wang 0060, Shixun Zhang |
ICA3PP (6) | 3 |
| 2025 | AUE: A Normalized Energy Efficiency Metric for AI Servers Under LLM WorkloadsabstractUnder the rapid advancement of large model-driven artificial intelligence, the surging energy consumption of AI training and inference tasks has created an urgent need for precise and comparable energy efficiency metrics to guide the design and deployment of green computing systems. While existing metrics such as PUE and Green500 metrics focus on infrastructure or traditional numerical computations, they cannot reflect the characteristics of AI workloads. Although applicationoriented metrics like J/response and J/token are designed for LLMs, they remain susceptible to biases induced by model scale and output strategies, lacking cross-model comparability. This paper proposes a novel AI energy efficiency metric, AUE, defined as the energy consumed per thousand tokens per billion activated parameters. By normalizing model size effects, AUE accurately reflects the energy efficiency of underlying computational resources. We theoretically justify the validity of AUE and conduct experiments on a server equipped with$4 \times$Ascend NPU 910C accelerators, evaluating dense Transformer and MoE architectures across both training and inference workloads. Experimental results demonstrate that traditional J/token metrics disproportionately favor smaller models, whereas AUE reveals true energy utilization efficiency. For instance, while Qwen3 0.6B shows superior J/token values compared to Qwen3 14B, the 14B model achieves a significantly better AUE of 12.88 J/(KToken GParam) versus 24.35 J/(KToken GParam) for the 0.6 B model, consistent with measured FLOPs where the 14B model outperforms its smaller counterpart. With advantages including simple measurement procedures and compatibility across platforms and model architectures, AUE provides a viable pathway toward establishing a normalized and standardized AI energy efficiency evaluation framework. Dongxiang Zhang, Qiang Wang 0060, Bingqiang Wang, Shixun Zhang, Yonghong Tian 0001 |
ICPADS | 5 |
| 2025 | NM-SpMM: Accelerating Matrix Multiplication Using N: M Sparsity with GPGPUabstractDeep learning demonstrates effectiveness across a wide range of tasks. However, the dense and over-parameterized nature of these models results in significant resource consumption during deployment. In response to this issue, weight pruning, particularly through$N: M$sparsity matrix multiplication, offers an efficient solution by transforming dense operations into semisparse ones.$N: M$sparsity provides an option for balancing performance and model accuracy, but introduces more complex programming and optimization challenges. To address these issues, we design a systematic top-down performance analysis model for$N: M$sparsity. Meanwhile, NM-SpMM is proposed as an efficient general$N: M$sparsity implementation. Based on our performance analysis, NM-SpMM employs a hierarchical blocking mechanism as a general optimization to enhance data locality, while memory access optimization and pipeline design are introduced as sparsity-aware optimization, allowing it to achieve close-to-theoretical peak performance across different sparsity levels. Experimental results show that NM-SpMM is 2.1x faster than nmSPARSE (the state-of-the-art for general$N: M$sparsity) and 1.4× to 6.3× faster than cuBLAS's dense GEMM operations, closely approaching the theoretical maximum speedup resulting from the reduction in computation due to sparsity. NM-SpMM is open source and publicly available at https://github.com/M-H482/NM-SpMM. Du Wu, Zhelang Deng, Jintao Meng 0001, Wenxi Zhu, Bingqiang Wang, Amelie Chi Zhou, Peng Chen 0035, Minwen Deng, Yanjie Wei, Shengzhong Feng, Yi Pan 0001 |
IPDPS | 8 |
| 2025 | Accelerating Model Training on Ascend Chips: An Industrial System for Profiling, Analysis and Optimization
Zhibin Wang 0002, Ruyi Zhang 0005, Chen Tian 0001, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen, Bingqiang Wang, Yonghong Tian 0001, Yan Zhang 0002, Hui Wang 0030, Fuchun Wei, Boquan Sun, Bin She, Teng Su, Yaoyuan Wang, Guyue Liu |
USENIX ATC | 9 |
| 2024 | Improving GPU Energy Efficiency through an Application-transparent Frequency Scaling Policy with Performance AssuranceabstractPower consumption is one of the top limiting factors in high-performance computing systems and data centers, and dynamic voltage and frequency scaling (DVFS) is an important mechanism to control power. Existing works using DVFS to improve GPU energy efficiency suffer from the limitation that their policies either impact performance too much or require offline application profiling or code modification, which severely limits their applicability on large clusters. To address this issue, we propose a novel GPU DVFS policy, GEEPAFS, which improves the energy efficiency of GPUs while providing performance assurance. GEEPAFS is application-transparent as it does not require any offline profiling or code modification on user applications. To achieve this, GEEPAFS models application performance online based on our quantitative analysis of a correlation between performance and GPU memory bandwidth utilization. Based on their relationship, GEEPAFS builds a fold-line frequency-performance model for applications being executed, and it applies the model to guide the setting of GPU frequency to maximize energy efficiency while ensuring the performance loss is bounded. Through experiments on NVIDIA V100 and A100 GPUs, we show that GEEPAFS is able to improve the energy efficiency by 26.7% and 20.2% on average. While achieving this improvement, the average performance loss is only 5.8%, and the worst-case performance loss is 12.5% among all 33 tested applications. Pengxiang Xu, Bingqiang Wang |
EuroSys | 5 |
| 2024 | DSO: A GPU Energy Efficiency Optimizer by Fusing Dynamic and Static InformationabstractIncreased reliance on graphics processing units (GPUs) for high-intensity computing tasks raises challenges regarding energy consumption. To address this issue, dynamic voltage and frequency scaling (DVFS) has emerged as a promising technique for conserving energy while maintaining the quality of service (QoS) of GPU applications. However, existing solutions using DVFS are hindered by inefficiency or inaccuracy as they depend either on dynamic or static information respectively, which prevents them from being adopted to practical power management schemes. To this end, we propose a novel energy efficiency optimizer, called DSO, to explore a light weight solution that leverages both dynamic and static information to model and optimize the GPU energy efficiency. DSO firstly proposes a novel theoretical energy efficiency model which reflects the DVFS roofline phenomenon and considers the tradeoff between performance and energy. Then it applies machine learning techniques to predict the parameters of the above model with both GPU kernel runtime metrics and static code features. Experiments on modern DVFS-enabled GPUs indicate that DSO can enhance energy efficiency by 19% whilst maintaining performance within a 5% loss margin. Laiyi Li, Weile Luo, Bingqiang Wang |
IWQoS | 5 |
| 2018 | mSNP: A Massively Parallel Algorithm for Large-Scale SNP DetectionabstractSingle Nucleotide Polymorphism (SNP) detection is a fundamental procedure of whole genome analysis. SOAPsnp, a classic tool for detection, would take more than one week to analyze one typical human genome, which limits the efficiency of downstream analyses. In this paper, we present mSNP, an optimized version of SOAPsnp, which leverages Intel Xeon Phi coprocessors for large-scale SNP detection. Firstly, we redesigned the essential data structures of SOAPsnp, which significantly reduces memory footprint and improves computing efficiency. Then we developed a coordinated parallel framework for a higher hardware utilization of both CPU and Xeon Phi. Also, we tailored the data structures and operations to utilize the wide VPU of Xeon Phi to improve data throughput. Last but not the least, we proposed a read-based window division strategy to improve throughput and obtain better load balance. mSNP is the first SNP detection tool empowered by Xeon Phi. We achieved a 38x single thread speedup on CPU, without any loss in precision. Moreover, mSNP successfully scaled to 4,096 nodes on Tianhe-2. Our experiments demonstrate that mSNP is efficient and scalable for large-scale human genome SNP detection. Yingbo Cui 0001, Shaoliang Peng, Yutong Lu, Xiaoqian Zhu, Bingqiang Wang, Chengkun Wu, Xiangke Liao |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2017 | Scalable Assembly for Massive Genomic GraphsabstractScientists increasingly want to assemble large genomes, metagenomes, and large numbers of individual genomes. In order to meet the demand for processing these huge datasets, parallel genome assembly is a vital step. Among all the parallel genome assemblers, de Bruijn graph based ones are most popular. However, the size of de Bruijn graph is determined by the number of distinct kmers used in the algorithm, thus redundant kmers in the genome datasets donot contribute to the graph size. The scalability of genome assemblers is influenced directly by the distinct kmers in the dataset or de Bruijn graph size, rather than the input dataset size. In order to assembly large genomes, we have artificially created 16 datasets of 4 Terabytes in total from the human reference genome. The human reference genome is firstly mutated with a 5% mutation rate, and then subjected to a genome sequencing data simulator ART. The simulated datasets have linearly increasing number of distinct kmers as the size/number of the combined datasets increases. We then evaluate all five time-consuming steps of the SWAP-Assembler 2.0 (SWAP2) using these 16 simulated datasets. Compared with our previous experiment on 1000 human dataset with fixed de Bruijn graph size, the weak-scaling test shows that SWAP2 can scale well from 1024 cores using one dataset to 16,384 cores. The percentage of time usage for all five steps of SWAP2 is fixed, and total time usage is also constant. The result showed that the time usage of graph simplification occupied almost 75% of the total time usage, which will be subject to further optimization for future work. Jintao Meng 0001, Jianqiu Ge, Yanjie Wei, Pavan Balaji, Bingqiang Wang |
CCGrid | 6 |
| 2017 | Bloomfish: A Highly Scalable Distributed K-mer Counting FrameworkabstractK-mer counting is a fundamental operation in DNA research and genome analytics; its application includes estimating genome assembly, understanding similarities in genomic samples, and merging a newly processed genome with a reference genome. As the genome dataset becomes larger and larger, designing a highly optimized distributed-memory implementation becomes more and more important. Current distributed-memory solutions have two limitations: they have a high memory footprint, and they do not provide advanced optimizations for loading enormous genome datasets into memory. Based on these observations, we present Bloomfish, a distributed, memory-efficient, scalable solution to the limits of current work. To keep a low memory footprint, Bloomfish leverages the compact hash array design of the single-node Jellyfish system and the optimized workflow of the high-performance MapReduce framework Mimir. We have also codesigned Mimir's I/O to efficiently load enormous datasets. We ran Bloomfish on the Tianhe-2 supercomputer with large sequence datasets (up to 24 TB). Our results show that Bloomfish achieves unprecedented scalability in genome analytics. Yanfei Guo, Yanjie Wei, Bingqiang Wang, Yutong Lu, Pietro Cicotti, Pavan Balaji, Michela Taufer |
ICPADS | 4 |
| 2016 | SWAP-Assembler 2: Optimization of De Novo Genome Assembler at Extreme ScaleabstractIn this paper, we analyze and optimize the most time-consuming steps of the SWAP-Assembler, a parallel genome assembler, so that it can scale to a large number of cores for huge genomes with sequencing data ranging from terabyes to petabytes. Performance analysis results show that the most time-consuming steps are input parallelization, k-mer graph construction, and graph simplification (edge merging). For the input parallelization, the input data is divided into virtual fragments with nearly equal size, and the start position and end position of each fragment are automatically separated at the beginning of the reads. In k-mer graph construction, in order to improve the communication efficiency, the message size is kept constant between any two processes by proportionally increasing the number of nucleotides to the number of processes in the input parallelization step for each round. The memory usage is also decreased because only a small part of the input data is processed in each round. With graph simplification, the communication protocol reduces the number of communication loops from four to two loops and decreases the idle communication time. The optimized assembler is denoted SWAP-Assembler 2 (SWAP2). In our experiments using a 1000 Genomes project dataset of 4 terabytes (the largest dataset ever used for assembling) on the supercomputer Mira, the results show that SWAP2 scales to 131,072 cores with an efficiency of 40%. We also compared our work with both the HipMer assembler and the SWAP-Assembler. On the Yanhuang dataset of 300 gigabytes, SWAP2 shows a 3X speedup and 4X better scalability compared with the HipMer assembler and is 45 times faster than the SWAP-Assembler. The SWAP2 software is available at https://sourceforge.net/projects/swapassembler. Jintao Meng 0001, Pavan Balaji, Yanjie Wei, Bingqiang Wang, Shengzhong Feng |
ICPP | 5 |
| 2015 | Accelerating large-scale biological database search on Xeon Phi-based neo-heterogeneous architecturesabstractIn this paper we present new parallelization techniques for searching large-scale biological sequence databases with the Smith-Waterman algorithm on Xeon Phi-based neoheterogenous architectures. In order to make full use of the compute power of both the multi-core CPU and the many-core Xeon Phi hardware, we use a collaborative computing scheme as well as hybrid parallelism. At the CPU side, we employ SSE intrinsics and multi-threading to implement SIMD parallelism. At the Xeon Phi side, we use Knights Corner vector instructions to gain more data parallelism. We have presented two dynamic task distribution schemes (thread level and device level) in order to achieve better load balancing. Furthermore, a multi-threaded asynchronous scheme is used to overlap communication and computation between CPUs and Xeon Phis. Evaluations on real protein sequence databases show that our method achieves a peak overall performance up to 220 GCUPS on a neo-heterogeneous platform consisting of two Intel E5-2620 CPUs and two Intel Xeon Phi 7110P cards. It also exhibits good scalability in terms of database size and query length. Our implementation is available at: http://turbo0628.github.io/LSBDS/. Haidong Lan, Bertil Schmidt, Bingqiang Wang |
BIBM | 4 |
| 2015 | The Challenge of Scaling Genome Big Data Analysis Software on TH-2 SupercomputerabstractWhole genome re-sequencing plays a crucial role in biomedical studies. The emergence of genomic big data calls for an enormous amount of computing power. However, current computational methods are inefficient in utilizing available computational resources. In this paper, we address this challenge by optimizing the utilization of the fastest supercomputer in the world - TH-2 supercomputer. TH-2 is featured by its neo-heterogeneous architecture, in which each compute node is equipped with 2 Intel Xeon CPUs and 3 Intel Xeon Phi coprocessors. The heterogeneity and the massive amount of data to be processed pose great challenges for the deployment of the genome analysis software pipeline on TH-2. Runtime profiling shows that SOAP3-dp and SOAPsnp are the most time-consuming components (up to 70% of total runtime) in a typical genome-analyzing pipeline. To optimize the whole pipeline, we first devise a number of parallel and optimization strategies for SOAP3-dp and SOAPsnp, respectively targeting each node to fully utilize all sorts of hardware resources provided both by CPU and MIC. We also employ a few scaling methods to reduce communication between different nodes. We then scaled up our method on TH-2. With 8192 nodes, the whole analyzing procedure took 8.37 hours to finish the analysis of a 300 TB dataset of whole genome sequences from 2,000 human beings, which can take as long as 8 months on a commodity server. The speedup is about 700x. Shaoliang Peng, Xiangke Liao, Canqun Yang, Yutong Lu, Jie Liu 0002, Yingbo Cui 0001, Chengkun Wu, Bingqiang Wang |
CCGRID | 9 |
| 2014 | SWAP-Assembler: scalable and efficient genome assembly towards thousands of coresabstractBACKGROUND: There is a widening gap between the throughput of massive parallel sequencing machines and the ability to analyze these sequencing data. Traditional assembly methods requiring long execution time and large amount of memory on a single workstation limit their use on these massive data. RESULTS: This paper presents a highly scalable assembler named as SWAP-Assembler for processing massive sequencing data using thousands of cores, where SWAP is an acronym for Small World Asynchronous Parallel model. In the paper, a mathematical description of multi-step bi-directed graph (MSG) is provided to resolve the computational interdependence on merging edges, and a highly scalable computational framework for SWAP is developed to automatically preform the parallel computation of all operations. Graph cleaning and contig extension are also included for generating contigs with high quality. Experimental results show that SWAP-Assembler scales up to 2048 cores on Yanhuang dataset using only 26 minutes, which is better than several other parallel assemblers, such as ABySS, Ray, and PASHA. Results also show that SWAP-Assembler can generate high quality contigs with good N50 size and low error rate, especially it generated the longest N50 contig sizes for Fish and Yanhuang datasets. CONCLUSIONS: In this paper, we presented a highly scalable and efficient genome assembly software, SWAP-Assembler. Compared with several other assemblers, it showed very good performance in terms of scalability and contig quality. This software is available at: https://sourceforge.net/projects/swapassembler. Jintao Meng 0001, Bingqiang Wang, Yanjie Wei, Shengzhong Feng, Pavan Balaji |
BMC Bioinform. | 2 |
| 2013 | GPU-Accelerated Bidirected De Bruijn Graph Construction for Genome Assembly
Mian Lu, Qiong Luo 0001, Bingqiang Wang, Junkai Wu, Jiuxin Zhao |
APWeb | 3 |
| 2013 | Improved Parallel Processing of Massive De Bruijn Graph for Genome Assembly
Jiefeng Cheng, Jintao Meng 0001, Bingqiang Wang, Shengzhong Feng |
APWeb | 4 |
| 2013 | GPU-accelerated adaptive compression framework for genomics dataabstractGenomics data is being produced at an unprecedented rate, especially in the context of clinical applications and grand challenge questions. There are various types of data in genomics research, most of which are stored as plain text tables. A data compression framework tailored to this file type is introduced in this paper, featuring a combination of generic compression algorithms, GPU acceleration, and column-major storage. This approach is the first to achieve both compression and decompression rates of around 100MB/s on commodity hardware without compromising compression ratio. By selecting appropriate compression schemes for each column of data, this framework efficiently exploits data redundancy while remaining applicable to a wide range of formats. The GPU-accelerated implementation also properly exploits the parallelism of compression algorithms. Finally, this paper presents a novel first-order Markov model based transformation, with evidence that it is at least as effective as Burrows-Wheeler and Move-To-Front in some contexts. GuiXin Guo, Zhiqiang Ye, Bingqiang Wang, Mian Lu, Simon See, Rui Mao 0001 |
IEEE BigData | 4 |
| 2012 | Improving Data Processing Time with Access Sequence PredictionabstractGenomic research nowadays often faces the problem of big data. The data size from genome sequencing process can grow very quickly and continuously creating the problem with storage and processing. BGI, one of the renowned genomic research institutes in China also faces the similar problem. The research at BGI depends on several sequencing machines. One machine pipeline may generate temporary data of around 1.4 terabytes. In addition, multiple read and write operations occur continuously during processing time. The I/O bottleneck thus degrades research throughput tremendously. Using a high performance computing system alone is not sufficiently effective in experimental results processing. In order to hide the I/O latency, an effective big data management framework is needed at BGI. In this paper, we proposed the hybrid prediction model for data access pattern. The goal is to predict the next pieces of data needed in the processor and preload them into the memory in order to improve the overall processing time. From the results obtained from the initial experiments, the proposed model can deliver high prediction accuracy in linear-time. Moreover, the error rate is low at 1.85%, which is better than the common methods used, such as Prediction Graph, ANN and ARMA. We believe that with some further fine-tuning, the model can be used as a part of the big data management framework deployed at BGI in the near future. Prasitchai Boonserm, Bingqiang Wang, Simon Chong Wee See, Tiranee Achalakul |
ICPADS | 2 |
| 2012 | SOAP3: ultra-fast GPU-based parallel alignment tool for short readsabstractAbstract Summary: SOAP3 is the first short read alignment tool that leverages the multi-processors in a graphic processing unit (GPU) to achieve a drastic improvement in speed. We adapted the compressed full-text index (BWT) used by SOAP2 in view of the advantages and disadvantages of GPU. When tested with millions of Illumina Hiseq 2000 length-100 bp reads, SOAP3 takes < 30 s to align a million read pairs onto the human reference genome and is at least 7.5 and 20 times faster than BWA and Bowtie, respectively. For aligning reads with up to four mismatches, SOAP3 aligns slightly more reads than BWA and Bowtie; this is because SOAP3, unlike BWA and Bowtie, is not heuristic-based and always reports all answers. Availability: SOAP3 is available at: http://www.cs.hku.hk/2bwt-tools/soap3; http://soap.genomics.org.cn/soap3.html. Contact: [email protected], [email protected] Chi-Man Liu, Thomas K. F. Wong, Edward Wu, Ruibang Luo, Siu-Ming Yiu, Yingrui Li, Bingqiang Wang, Xiaowen Chu 0001, Kaiyong Zhao, Ruiqiang Li, Tak Wah Lam |
Bioinform. | 7 |
| 2012 | Gene set analysis in the cloudabstractUNLABELLED: Cloud computing offers low cost and highly flexible opportunities in bioinformatics. Its potential has already been demonstrated in high-throughput sequence data analysis. Pathway-based or gene set analysis of expression data has received relatively less attention. We developed a gene set analysis algorithm for biomarker identification in the cloud. The resulting tool, YunBe, is ready to use on Amazon Web Services. Moreover, here we compare its performance to those obtained with desktop and computing cluster solutions. AVAILABILITY AND IMPLEMENTATION: YunBe is open-source and freely accessible within the Amazon Elastic MapReduce service at s3n://lrcv-crp-sante/app/yunbe.jar. Source code and user's guidelines can be downloaded from http://tinyurl.com/yunbedownload. Lu Zhang 0026, Shengchang Gu, Bingqiang Wang, Francisco Azuaje |
Bioinform. | 4 |
| 2011 | GSNP: A DNA Single-Nucleotide Polymorphism Detection System with GPU AccelerationabstractWe have developed GSNP, a software package with GPU acceleration, for single-nucleotide polymorphism detection on DNA sequences generated from second-generation sequencing equipment. Compared with SOAPsnp, a popular, high-performance CPU-based SNP detection tool, GSNP has several distinguishing features: First, we design a sparse data representation format to reduce memory access as well as branch divergence. Second, we develop a multipass sorting network to efficiently sort a large number of small arrays on the GPU. Third, we compute a table of frequently used scores once to avoid repeated, expensive computation and to reduce random memory access. Fourth, we apply customized compression schemes to the output data to improve the I/O performance. As a result, on a server equipped with an Intel Xeon E5630 2.53 GHZ CPU and an NVIDIA Tesla M2050 GPU, it took GSNP about two hours to analyze a whole human genome dataset whereas the CPU-based, single-threaded SOAPsnp took three days for the same task on the same machine. Mian Lu, Jiuxin Zhao, Qiong Luo 0001, Bingqiang Wang, Shaohua Fu, Zhe Lin 0004 |
ICPP | 4 |