VLDB 2026 Research / reviewers in the wild / expert
Qiang Wang 0006
dblp:64/5630-6
· DBLP profile ↗
20ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-7078-7545ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Energy-Efficient 0.56-pJ/cycle AVFS System Based on a Fast Transient Response Digital LDO and a Self-Calibrating Elastic Clock
Jiliang Liu, Zhengbin Pang, Fangxu Lv, Shijie Li 0002, Qiang Wang 0006, Lizhou Wu, Chengzhuo Zhao |
ISCAS | 6 |
| 2026 | Model compression-driven instruction data mining framework with integrated adversarial attack strategies
Qiang Wang 0006, Bo Ding 0001, Huaimin Wang 0001 |
Expert Syst. Appl. | 1 |
| 2026 | Pay more attention to the robustness of LLMs on adversarial prompt for instruction data mining
Qiang Wang 0006, Bo Ding 0001, Huaimin Wang 0001 |
Neural Networks | 1 |
| 2025 | PIAR: Path-Improved Adaptive Routing for Dragonfly NetworksabstractFor the next-generation exascale supercomputing communication systems, Dragonfly topology offers strong scalability, low latency, and cost efficiency. Dragonfly networks have already been implemented in current supercomputers and will continue to expand in future systems. Adaptive routing in Dragonfly topologies is critical for network performance. The traditional UGAL routing algorithm, which uses the valiant mechanism to select non-minimal paths, does not adequately consider the impact of high hops in non-minimal paths, often unnecessarily increasing the average path length, thereby increasing network latency and load. Furthermore, UGAL inaccurately estimates the congestion of the entire routing path based on local information, leading to suboptimal routing decisions that limit the algorithm's performance. In this paper, we propose PIAR, a novel pathimproved adaptive routing algorithm. PIAR dynamically selects paths based on the status of local and global channels, prioritizing non-minimal paths with fewer hops to reduce network latency and load, thereby improving network performance. Additionally, we present the microarchitecture of the routing computation unit. Our evaluation results demonstrate that, compared with advanced algorithms such as PAR$_{\text {PH }}$, TPR, and UGAL LE, PIAR achieves an average throughput improvement of 19.2 % and reduces latency by up to$\mathbf{1 3. 4 \%}$under the single synthetic traffic. Under mixed traffic, PIAR achieves an average throughput improvement of$\mathbf{2 3. 6 \%}$and reduces the latency by up to$\mathbf{3 3. 8 \%}$. For application workloads, PIAR achieves an average reduction of 24.0 % in packet latency. Qiang Wang 0006, Jinbo Xu, Guo Chen 0001 |
CLUSTER | 2 |
| 2025 | PSCA: A FPGA-based Protein Structure Comparison Accelerator with Symmetric Simplified Matrix
Hui Su, Xingyun Qi, Qiang Wang 0006, Puguang Liu, Haoyu Liao |
ICA3PP (6) | 4 |
| 2025 | Enhancing Transformer Inference Efficiency on FPGA Through Fully Fusion and Integer-Only Quantization TechniquesabstractThe Transformer architecture has revolutionized the field of natural language processing (NLP) through its selfattention mechanism. However, its high computational complexity and memory requirement present significant deployment challenges on resource-constrained edge devices. While existing research predominantly focuses on accelerating linear operations via model compression and approximation techniques, the inefficiencies and high deployment costs of nonlinear operations (e.g., Softmax and LayerNorm) remain critically understudied. Although some studies have attempted to mitigate these challenges through techniques such as kernel fusion and integer-only quantization, these approaches still suffer from partial fusion and inefficient quantization with retained division operations, leaving significant efficiency gains unexploited. To bridge these gaps, we propose a fully fused Transformer accelerator that co-optimizes both linear and nonlinear operations while minimizing memory bottlenecks. For linear computations, our design incorporates a deeply optimized compute engine featuring double buffering, an output-stationary tiling strategy, and DSP-packing technology to maximize throughput. For nonlinear operations, we introduce a delayed computation strategy for vector-wise operators, effectively reducing memory bandwidth pressure and dependency stalls. Furthermore, we propose a hardware-efficient, divisionfree integer-only quantization scheme, leveraging$\log 2$quantization for Softmax and a polynomial-enhanced approximation for LayerNorm to eliminate costly floating-point units, thereby significantly reducing latency and resource overhead. Through systematic design space exploration, our solution, deployed on the Zynq Z-7100 platform, achieves 1.376 TOPS for BERT inference, demonstrating a$\mathbf{1. 6 5 - 2. 5 1} \boldsymbol{\times}$higher computational efficiency compared to prior works. Zhenqi Li, Puguang Liu, Qiang Wang 0006, Yankang Zhao, Hanyuan Li, Xingyun Qi |
ICCD | 5 |
| 2025 | Bandwidth Optimized Scalable Designs with Inter-Layer Overlapping for MPI BroadcastabstractMPI (Massage Passing Interface) has been the dominant programming model for developing large-scale parallel applications. Existing work mainly focuses on the vast parallelism of modern multi-/many-cores architectures to parallelize MPI collectives. However, the abundant bandwidth provided by modern interconnects is either underutilized when processing small messages or overwhelmed by large messages. MPI_Bcast is one of the most widely used collective primitives in MPI, which broadcasts data from one process to all processes in the communication domain. In this paper, we address the issue of load imbalance arising from traditional tree-based designs in order to strike a better balance between bandwidth and latency of MPI Broadcast. By evenly distributing the broadcast load across lower layers, we effectively leverage the available resources at upper-layer nodes that recursively execute broadcast across different layers in a sequential and contention-free manner. This approach improves the scalability and performance of large-scale message broadcasts. Additionally, a generic inter-layer overlapped scheme is proposed to reduce overall broadcast latency by fully overlapping inter-layer and intra-layer data transmission. We further implement an online adaptive scheme for tree degree tuning to achieve optimal design at various message and system sizes. Extensive experiments are conducted to evaluate the performance of this Bandwidth-optimized Inter-layer Overlapping (BIO) design at both the microbenchmark and application levels. BIO-based MPI_Bcast designs demonstrate performance speedups of up to 2.71x compared to state-of-the-art MPI libraries. For application-level evaluation, BIO provides up to 165% acceleration for the initialization of the distributed deep learning model of Horovod with the PyTorch application. Qiang Wang 0006, Bo Yang 0023, Dongsheng Li 0001 |
ICDCS | 4 |
| 2025 | Max-Informative Unlabeled Sample Replay for Semi-Supervised Class-Incremental Learning in Audio ClassificationabstractAudio classification constitutes a critical task aimed at assigning meaningful labels to audio recordings. Despite commendable efforts in this field, prior endeavors have encountered notable challenges. Firstly, the majority of existing solutions operate under the assumption of a fixed vocabulary for classification, neglecting the crucial need for systems to adapt continuously to dynamic data streams. Secondly, these approaches often assume that input data is fully annotated, a presumption that frequently diverges from practical scenarios. In response to these challenges, we present a novel semi-supervised class-incremental learning framework that utilizes memory replay with unlabeled data and introduces temporal consistency regularization to better alleviate catastrophic forgetting. Additional, we employ Iterative Projection and Matching to select the most informative samples for storage, addressing the issues of reservoir sampling while maintaining a balanced buffer. We showcase the framework's efficacy through a series of comprehensive experiments conducted on both the ESC-50 and Google Speech Commands datasets. The results demonstrate the remarkable performance of our proposed approach, particularly in few-shot scenarios. Qiang Wang 0006, Bo Ding 0001, Huaimin Wang 0001 |
JCC | 1 |
| 2025 | Memory replay with unlabeled data for semi-supervised class-incremental learning via temporal consistency
Qiang Wang 0006, Kele Xu, Bo Ding 0001, Huaimin Wang 0001 |
Frontiers Comput. Sci. | 1 |
| 2025 | Accurate and Efficient Fine-Tuning of Quantized Large Language Models Through Optimal Balance in AdaptationabstractAbstract Large Language Models (LLMs) have demonstrated impressive performance across various domains. However, the enormous number of model parameters makes fine-tuning challenging, significantly limiting their application and deployment. Existing solutions combine parameter quantization with Low-Rank Adaptation (LoRA), reducing memory usage but causing performance degradation. Additionally, converting fine-tuned models to low-precision representations further degrades performance. In this paper, we identify an imbalance in fine-tuning quantized LLMs with LoRA: overly complex adapter inputs and outputs versus low effective trainability of the adapter, leading to underfitting during fine-tuning. Thus, we propose Quantized LLMs fine-tuning with Balanced Low-Rank Adaptation (Q-BLoRA), which simplifies the adapter’s inputs and outputs while increasing the adapter’s rank to alleviate underfitting during fine-tuning. For low-precision deployment, we propose Quantization-Aware fine-tuning with Balanced Low-Rank Adaptation (QA-BLoRA), which aligns with the block-wise quantization and facilitates quantization-aware fine-tuning of low-rank adaptation based on the parameter merging of Q-BLoRA. Both Q-BLoRA and QA-BLoRA are easily implemented and offer the following optimizations: (i) Q-BLoRA consistently achieves state-of-the-art accuracy compared to baselines and other variants; (ii) QA-BLoRA enables the direct generation of low-precision inference models, which exhibit significant performance improvements over other low-precision models. We validate the effectiveness of Q-BLoRA and QA-BLoRA across various models and scenarios. Code has been made available at https://github.com/xiaocaigou/qbaraqahira. Zhiquan Lai, Qiang Wang 0006, Xionglve Li, Dongsheng Li 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2025 | An Adaptive 56-Gb/s Duo-PAM4 Detector Using Reduced Branch Maximum Likelihood Sequence Detection in a 28-nm CMOS Wireline ReceiverabstractThis paper describes an adaptive duo-binary four-level pulse amplitude modulation (Duo-PAM4) detector that significantly reduces the bit error rate (BER) of the conventional wireline transceivers under high insertion loss (IL) channels. The parallel maximum likelihood sequence detection (MLSD) combined with parallel feed-forward equalization (FFE) is proposed to generate, equalize, and detect Duo-PAM4 signals, thus reducing BER compared to conventional decision feedback equalizer (DFE) and slicers. The proposed reduced branch MLSD reduces power consumption compared to MLSD. An improved delay zero-forcing algorithm for Duo-PAM4 is proposed to achieve fast convergence of the FFE tap coefficients, reducing convergence time by up to 72.5% compared to conventional ZF algorithms for Duo-PAM4. Both the proposed and conventional detectors are implemented in a 28-nm CMOS process at 56 Gb/s and 38-dB insertion loss. The FFE+MLSD and FFE+RB-MLSD reduce the BER by two orders of magnitude compared to conventional FFE+DFE+slicer. The RB-MLSD reduces power consumption by 21.1% compared to conventional MLSD. The detector can be easily migrated to 112Gb/s or 224Gb/s transceivers. Chaolong Xu, Fangxu Lv, Qiang Wang 0006, Xiaoyue Hu, Cewen Liu, Zhouhao Yang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | A Novel High-Speed Adaptive Duobinary Digital Detector Based on the Feed-Forward Equalizer and the Maximum Likelihood Sequence Detector for Wireline TransceiversabstractTo solve the high bit error rate (BER) problem of conventional 56-Gb/s nonreturn-to-zero (NRZ) transceivers under high-insertion loss (IL) channels, this study proposes a high-speed adaptive duobinary (DB) digital detector based on the feed-forward equalizer (FFE) and the maximum likelihood sequence detector (MLSD). In this detector, adaptive FFE is combined with channel characteristics to generate DB signals and complete equalization, thus extending the transmission bandwidth and eye height and allowing a larger sampling phase offset. The parallel MLSD is used to complete the detection and decoding of DB signals to reduce the BER. An adaptive algorithm is proposed to avoid the long convergence time of the conventional zero-forcing (ZF) algorithm applied to the DB detector, so that it can be applied to various bit rates and IL channels. In this study, the verification of this DB detector is accomplished at 56 Gb/s. The platform based on a 56-Gb/s analog front-end chip (AFEC) and field-programmable gate array (FPGA) proves that the detector can work well in 12–56 Gb/s and multiple IL channels. The BER was less than 2e-8 at 56 Gb/s on −42-dB channel loss at 28 GHz. The structure can be well used for higher rate transceivers, such as 112 Gb/s. Chaolong Xu, Fangxu Lv, Xingyun Qi, Qiang Wang 0006, Zhang Luo, Shijie Li 0002, Geng Zhang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | Automatic Implementation of Large-Scale CNNs on FPGA Cluster Based on HLS4MLabstractConvolutional Neural Networks (CNNs) have demonstrated remarkable performance across various computer vision tasks. Due to the computational and data-intensive nature of CNNs, Field-Programmable Gate Arrays (FPGAs) are exceptionally well-suited for accelerating the CNN computation process. However, large-scale CNNs such as ResNet-84 contain an enormous number of parameters that exceed the capacity of a single FPGA, rendering the deployment on a single device impractical. In this paper, we develop an automated end-to-end design flow for mapping large-scale CNNs across multiple FPGAs, on the basis of the HLS4ML dataflow architecture. We propose a graph optimization method to streamline the CNN structure and reduce resource consumption. We also summarize a resource allocation algorithm that automatically determines the specific hardware resources necessary for each CNN layer. Furthermore, we introduce a partitioning methodology capable of effectively segmenting CNNs into multiple FPGAs, and providing each subgraph with specific interfaces to support communication between different FPGAs. To validate the methodology, we construct a multi-FPGA platform interconnected via LVDS. We select two typical networks, ResNet-8 and ResNet-84, as the benchmarks for evaluation. The experimental results demonstrate that our approach significantly outperforms existing solutions. It attains an 18.6-fold increase in speed over Vitis AI, a 2.2-fold improvement over FINN, and a 3.4-fold enhancement over the original single-FPGA HLS4ML implementation for ResNet-8. For ResNet-84, our method achieves a remarkable 33.6-fold speedup over Vitis AI. Additionally, when compared to other non-automated multi-FPGA solutions, our methodology still exhibits significant performance improvements. Xingyun Qi, Yankang Zhao, Zhenqi Li, Hanyuan Li, Qiang Wang 0006 |
ISPA | 8 |
| 2018 | mmCNN: A Novel Method for Large Convolutional Neural Network on Memory-Limited DevicesabstractDeep learning recently has been widely used in many interactive application fields including but not limited to object recognition, speech recognition, natural language processing and so on. At the same time more and more attractive interactive applications (face recognition and augmented reality) are available on wearable and mobile devices. However, traditional deep learning methods such as CNN cost a lot of memory resources. This challenge makes it difficult to apply the powerful deep learning method on mobile memory limited platforms. In this paper we present a novel memory management strategy called mmCNN to solve this problem. This method helps us deploy a trained large size CNN on an any memory size platform including GPU, FPGA and memory-limited mobile devices. In our experiments, we run a feed-forward CNN process in an extremely small memory size (as low as 5MB) on a GPU platform. The result shows that our method saves more than 98% memory compared to a traditional CNN algorithm and further saves more than 90% compared to the sate-of-the-art related work "vDNN". Our work improve the computing scalability of interaction applications and break the memory bottleneck of using deep learning method on a memory-limited devices. Shijie Li 0002, Yong Dou, Jinwei Xu, Qiang Wang 0006, Xin Niu 0002 |
COMPSAC (1) | 4 |
| 2018 | Deep Image Clustering Using Convolutional Autoencoder Embedding with Inception-Like BlockabstractImage clustering is one of the challenging tasks in machine learning, and has been extensively used in various applications. Recently, various deep clustering methods has been proposed. These methods take a two-stage approach, feature learning and clustering, sequentially or jointly. We observe that these works usually focus on the combination of reconstruction loss and clustering loss, relatively little work has focused on improving the learning representation of the neural network for clustering. In this paper, we propose a deep convolutional embedded clustering algorithm with inception-like block (DCECI). Specifically, an inception-like block with different type of convolution filters are introduced in the symmetric deep convolutional network to preserve the local structure of convolution layers. We simultaneously minimize the reconstruction loss of the convolutional autoencoders with inception-like block and the clustering loss. Experimental results on multiple image datasets exhibit the promising performance of our proposed algorithm compared with other competitive methods. Qiang Wang 0006, Rongchun Li, Peng Qiao, Ke Yang 0004, Shijie Li 0002, Yong Dou |
ICIP | 1 |
| 2018 | Temporal Pyramid Relation Network for Video-Based Gesture RecognitionabstractGesture recognition in video is an important application of computer vision. However, there are few works talked about the temporal order or relation of the frames in video, which is important for model gestures. In this paper, we propose Temporal Pyramid Relation Network (TPRN) which can model the temporal relation of video frames effectively and efficiently. First, we use Temporal Pyramid Pooling (TPP) layer to get temporal feature sequences of multiple scale pyramids. Then, a Temporal Relation Network (TRN) is stacked on the feature sequence of each scale respectively to model the temporal relations of video frames at multiple scales. At last, representations of all scales are aggregated to get the final prediction. TPRN can take video clips of various length as input and is scalable for video length. We evaluate TPRN on a recently released very large video-based gesture recognition dataset - 20BN-Jester dataset v1, and TPRN achieves competitive performance. Ke Yang 0004, Rongchun Li, Peng Qiao, Qiang Wang 0006, Dongsheng Li 0001, Yong Dou |
ICIP | 4 |
| 2018 | Local kernel alignment based multi-view clustering using extreme learning machine
Qiang Wang 0006, Yong Dou, Xinwang Liu 0002, Fei Xia 0003, Ke Yang 0004 |
Neurocomputing | 1 |
| 2017 | An FPGA-based processor for training convolutional neural networksabstractConvolutional neural networks (CNNs) have gained great success in various computer vision applications. However, training a CNN model is computation-intensive and time-consuming. Hence training is mainly processed on large clusters of high-performance processors like server CPUs and GPUs. In this paper, we propose an FPGA-based processor design to accelerate the training process of CNNs. We first analyze the operations in all types of CNN layers in the training process. A uniform computation engine design is proposed to efficiently carry out all kinds of operations based on the analysis. Then a scalable accelerator framework is presented that exploits the parallelism further by unrolling the loops in two levels. The proposed accelerator design is demonstrated by implementing a processor on the Xilinx ZU19EG FPGA working at 200 MHz. The evaluation results on a group of CNN models show that our processor is 5.7 to 10.7-fold faster than the software implementations on the Intel Core i5-4440 CPU(@3.10GHz). Yong Dou, Jingfei Jiang, Qiang Wang 0006, Paul Chow |
FPT | 4 |
| 2017 | A fast and memory saved GPU acceleration algorithm of convolutional neural networks for target detection
Shijie Li 0002, Yong Dou, Xin Niu 0002, Qiang Wang 0006 |
Neurocomputing | 5 |
| 2016 | Multi-view clustering with extreme learning machine
Qiang Wang 0006, Yong Dou, Xinwang Liu 0002, Shijie Li 0002 |
Neurocomputing | 1 |