Wang Wang

dblp:26/1377 · DBLP profile ↗
← Back
19ranked-venue papers
4as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Understanding Data Influence With Differential Approximation
abstract
Data plays a pivotal role in the groundbreaking advancements in artificial intelligence. The quantitative analysis of data significantly contributes to model training, enhancing both the efficiency and quality of data utilization. However, existing data analysis tools often lag in accuracy. For instance, many of these tools even assume that the loss function of neural networks is convex. These limitations make it challenging to implement current methods effectively. In this paper, we introduce a new formulation to approximate a sample's influence by accumulating the differences in influence between consecutive learning steps, which we term Diff-In. Specifically, we formulate the sample-wise influence as the cumulative sum of its changes/differences across successive training iterations. By employing second-order approximations, we approximate these difference terms with high accuracy while eliminating the need for model convexity required by existing methods. Despite being a second-order method, Diff-In maintains computational complexity comparable to that of first-order methods and remains scalable. This efficiency is achieved by computing the product of the Hessian and gradient, which can be efficiently approximated using finite differences of first-order gradients. We assess the approximation accuracy of Diff-In both theoretically and empirically. Our theoretical analysis demonstrates that Diff-In achieves significantly lower approximation error compared to existing influence estimators. Extensive experiments further confirm its superior performance across multiple benchmark datasets in three data-centric tasks: data cleaning, data deletion, and coreset selection. Notably, our experiments on data pruning for large-scale vision-language pre-training show that Diff-In can scale to millions of data points and outperforms strong baselines.
Haoru Tan, Sitong Wu, Xiuzhe Wu, Wang Wang, Zeke Xie, Gui-Song Xia, Xiaojuan Qi 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Linearization of Quadrature Digital Power Amplifiers by Neural Network of ULR_LSTM: Unsupervised Learning Residual LSTM
abstract
For the first time, this paper presents an unsupervised learning residual long short-term memory (ULR_LSTM) neural network to develop a digital predistortion (DPD) method for the linearization of digital power amplifiers (DPAs). Our method eliminates the need for iterative learning control (ILC) to obtain the ideal input of the DPA required by state-of-the-arts (SOTAs), which leads to high computational complexity and extensive training time. We perform behavioral modeling of the DPA using the R_LSTM network. After determining the optimal behavioral model architecture, the corresponding DPD model is obtained through an inverse training process. A 15-bit transformer-based quadrature DPA chip incorporating Class-G and IQ-cell-sharing techniques was implemented in a 28nm CMOS process to validate our proposed method. Experimental results demonstrate outstanding linearization performance comparing to prior arts, achieving an error vector magnitude (EVM) of -40.4dB for the 802.11ax 40MHz 64QAM signal.
Luyi Guo, Yicheng Li 0002, Wang Wang, Manni Li, Zijian Huang 0017, Yinyin Lin, Yun Yin, Hongtao Xu
DATE4
2025 Digital Predistortion for Quadrature Digital Power Amplifiers Using Deep Neural Network of AT_LSTM: Attention LSTM
Wending Zhao, Yicheng Li 0002, Wang Wang, Manni Li, Zijian Huang 0017, Yinyin Lin, Yun Yin, Hongtao Xu
ACM Great Lakes Symposium on VLSI4
2025 CPSnB: Compressing and Processing Spatial Similarity near Memory Bank for DNNs
abstract
Near memory bank processing (NMBP) architecture only benefits memory-bound operations of DNNs in terms of energy consumption. Drawing on the insight that data compression can reduce the compute density of operators, transforming compute-bound operations into memory-bound operations, We propose CPSnB, a NMBP architecture combined with preserving numerical jump-spatial similarity compression (PNJ-SSC) method. CPSnB provides a tiling strategy for optimizing operators of different DNN models. Compared to the systolic host-side accelerator and existing dense and sparse NMBP, CPSnB significantly reduces energy consumption. Analysis of the experimental results indicates that a 60% compression ratio of activation can enhance the versatility of CPSnB in processing DNN operators to 22.3 times.
Wang Wang, Manni Li, Zijian Huang 0017, Yinyin Lin, Chengchen Wang, Xiankui Xiong
ISCAS1
2025 APCPU: Adaptive-Pooling Compression Processing Unit for Energy-Efficient DNNs Processing
abstract
Integrating compression in the multiply-and-accumulate (MAC) path can significantly improve the energy efficiency of DNN operators. However, existing unstructured sparse compression (USSC) methods struggle to effectively compress activations with low sparsity. Computing core processing USSC face challenges such as load imbalance and complex index control circuit design. Based on insights into local spatial correlation, a block-wise adaptive-pooling compression (APC) method is proposed to achieve a high compression ratio for activations. Furthermore, this paper proposes an APCPU to integrate APC into the MAC path with minimal overhead, facilitating highly energy-efficient sparse processing of DNN operators. Leveraging a hybrid data flow design to achieve load balancing results in speedups of 1.25× to 1.33×. The experiment results show that the APCPU achieves energy savings of 1.35× and 1.27× compared to JPZ-PU, and 2.63× and 2.71× compared to CSC-PU when evaluated on AlexNet and Bert.
Wang Wang, Wending Zhao, Manni Li, Zijian Huang 0017, Yinyin Lin, Chengchen Wang, Xiankui Xiong
ISCAS1
2025 GPOS: A General and Precise Offloading Strategy for High Generality of DNN Acceleration by OCP and NDP Co-Optimizing
abstract
The arithmetic intensity (ArI) of different DNNs can be opposite. This challenges the generality of single acceleration architectures, including both dedicated on-chip processing (OCP) and near-data processing (NDP). Neither architecture can simultaneously achieve optimal energy efficiency and performance for operators with opposite ArI. It is relatively straightforward to think of combining the respective advantages of OCP and NDP. However, few publications have addressed their real-time co-optimization, primarily due to the lack of a quantifiable offloading method. Here, we propose GPOS, a general and precise offloading strategy that supports high generality of DNN acceleration. GPOS comprehensively considers the complex interactions between OCP and NDP, including hardware configurations, dataflow (DF), DNN model, and interdie data movements (DMs). Three quantifiable indicators—ArI, execution cost (Ex-cost), and DM-cost—are employed to precisely evaluate the impacts of these interactions on energy and latency. GPOS adopts a four-step flow with progressive refinement: each of the first three steps focuses on a single indicator at the operator level, while the final step performs context-based calibration to address operator interdependencies and avoid offsetting NDP benefits. Narrowing down offloading candidates in step 1 and step 3 significantly accelerates real-time quantitative analysis. Optimized mapping techniques and NDP-input stationary DF are proposed to reduce Ex-cost and extend operator types supported by NDP. Next, for the first time, sparsity—one of the most popular methods for energy optimization that can alter data reuse or ArI—is quantitatively investigated for its impacts on offloading using GPOS. Our evaluations include representative DNNs, including GPT-2, Bert, RNN, CNN, and MLP. GPOS achieves the minimum energy and latency for each benchmark, with geometric mean speedups of 49.0% and 94.1%, and geometric mean energy savings of 45.8% and 89.2% over All-OCP and All-NDP, respectively. GPOS also reduces offloading analysis latency by a geometric mean of 92.7% compared to the evaluation that traverses each operator and its relative combinations. On average, sparsity further improves performance and energy efficiency by increasing the number of operators offloaded to NDP. However, for DNNs where all operators exhibit either very high or very low ArI, the number of offloaded operators remains unchanged, even after sparsity is applied.
Wang Wang, Manni Li, Zijian Huang 0017, Yinyin Lin, Chengchen Wang, Xiankui Xiong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 INF-PCA: Implicit Neural Field-Based Interactive Point Cloud Semantic Annotation
abstract
Point cloud semantic segmentation helps Intelligent Transportation Systems understand traffic scenes by assigning semantic label to each point in the point cloud, and it relies on large amounts of annotated training data. Nevertheless, manually annotating large-scale datasets of complex traffic scenes is quite time-consuming and tedious. This paper proposes INF-PCA, an interactive point cloud semantic annotation method based on implicit neural field, which allows users to achieve high-quality, large-scene and fast-response semantic annotation with only a few dozen mouse clicks. Firstly, the appearance, geometry and semantics of the point clouds are jointly represented by an implicit neural field, which maps a 3D spatial coordinate to its corresponding attributes. Secondly, an uncertainty-based semantic entropy loss and a supervoxel-based local consistency loss are designed to force the network to produce deterministic predictions with local consistency, thus generating smoother and more accurate boundaries. Furthermore, an active learning-based strategy for click-free annotation is proposed and analyzed to further reduce annotation pressure. Comprehensive experiments on multiple datasets including the road scene dataset Toronto3D revealed that INF-PCA can achieve more accurate annotations with faster response speed and only half of the clicks employed by the state-of-the-art methods, and that INF-PCA can be directly applied to intelligent transportation applications such as interactive segmentation of road scenes, inventory of transportation infrastructure assets, and production of high-definition map.
Chen Long, Wang Wang, Zhen Dong 0005, Bisheng Yang
IEEE Trans. Intell. Transp. Syst.5
2025 Bi-Constraints Diffusion: A Conditional Diffusion Model With Degradation Guidance for Metal Artifact Reduction
abstract
In recent years, score-based diffusion models have emerged as effective tools for estimating score functions from empirical data distributions, particularly in integrating implicit priors with inverse problems like CT reconstruction. However, score-based diffusion models are rarely explored in challenging tasks such as metal artifact reduction (MAR). In this paper, we introduce a Bi-Constraints Diffusion Model for Metal Artifact Reduction (BCDMAR), an innovative approach that enhances iterative reconstruction with a conditional diffusion model for MAR. This method employs a metal artifact degradation operator in place of the traditional metal-excluded projection operator in the data-fidelity term, thereby preserving structure details around metal regions. However, score-based diffusion models tend to be susceptible to grayscale shifts and unreliable structures, making it challenging to reach an optimal solution. To address this, we utilize a pre-corrected image as a prior constraint, guiding the generation of the score-based diffusion model. By iteratively applying the score-based diffusion model and the data-fidelity step in each sampling iteration, BCDMAR effectively maintains reliable tissue representation around metal regions and produces highly consistent structures in non-metal regions. Through extensive experiments focused on metal artifact reduction tasks, BCDMAR demonstrates superior performance over other state-of-the-art unsupervised and supervised methods, both quantitatively and qualitatively.
Mengting Luo, Tao Wang 0167, Linchao He, Wang Wang, Hu Chen 0002, Peixi Liao, Yi Zhang 0018
IEEE Trans. Medical Imaging5
2024 Dual-Stream Heterogeneous Graph Neural Network Based on Zero-Shot Embeddings for Predicting miRNA-Drug Sensitivity
abstract
MicroRNAs (miRNAs) are a class of non-coding RNA molecules that have been shown to be closely associated with the sensitivity of chemotherapeutic drugs in cancer treatment. Given the high cost and extended duration of traditional biological experiments, there is an urgent need to develop computational models to predict the sensitivity scores between miRNAs and drugs. In this study, we proposed a dual-stream graph neural network method based on Zero-Shot Embeddings, named DSHGZS, to explore the potential sensitivity scores between miRNAs and drugs. DSHGZS first constructs two heterogeneous graphs with different isomorphic subgraphs based on zero-shot embeddings obtained from large language models (LLMs) and known miRNA-drug association data. It then utilized the enhanced LLM-derived node feature representations, embedding them into the layer feature learning process of the two heterogeneous graphs to generate high-quality vector representations of miRNAs and drugs. The learned high-quality feature embeddings are subsequently used in a segmented inner product decoder to evaluate the sensitivity association scores between miRNAs and drugs. To address the model’s excessive reliance on high-quality feature representations, we employed PCA to extract the core representations of the LLM-derived node features for data augmentation. Case studies demonstrated that DSHGZS is an effective tool for predicting potential sensitivity scores between miRNAs and drugs.
Wang Wang, Wenhui Xiao, Xiangzheng Fu
BIBM2
2024 LauWS: Local Adaptive Unstructured Weight Sparsity of Load Balance for DNN in Near-Data Processing
abstract
Memory wall issue has become the overwhelming bottleneck of future systems due to the explosive parameter growth and low computing density large language model (LLM). Near-data processing (NDP) could alleviate data traffic and energy consumption, but the storage demand of LLM is still enormous. Weight sparsity is helpful for reducing data capacity. Unstructured sparsity sacrifices less accuracy compared to structured one, but the random non-zero values distribution in NDP leads to load imbalance among parallel processing units. Here we propose LauWS which is seamlessly combined into various prior arts of sparsity. LauWS follows the local characteristics of feature distribution in weight matrix for various models, preserving even tiny features and discarding non-feature values as far as possible region by region. That is the key for LauWS achieving a trade-off between high prune ratio (PR) and less accuracy loss (AL). Evaluations are carried out based on a GDDR6-based bank-NDP system. The typical optimization compared to the no-prune includes 38% speedup at 0.8PR with no AL for MLP, 22.7% speedup at 0.5PR with no AL for GPT-2, 23.6% speedup at 0.5PR with the lowest perplexity for OPT-125m.
Wang Wang, Manni Li, Yinyin Lin, Guhyun Kim, Yosub Song, Chengchen Wang, Xiankui Xiong
ISCAS2
2024 AMAD: Active learning-based multivariate time series anomaly detection for large-scale IT systems
abstract
Multivariate time series anomaly detection on key performance indicators helps mitigate the impact of large-scale IT system anomalies. Due to the large volume and the abstract nature of multivariate time series, previous works have tended to make overly strict or optimistic hypotheses on labeling costs and resulted in unsatisfactory results. Thus, it remains a challenge to make an appropriate trade-off between labeling costs and model performance. This research proposes AMAD, an active learning-based approach that works to address this problem. Its core idea is to provide the learner model high-value label queries via an ensemble query strategy, which dynamically adapts to the estimated anomaly ratio and the ever-changing model performance. Moreover, it is the first to discuss in detail the margin effect of query strategies on model performance, to our knowledge, this has not been investigated in previous works on time series anomaly detection. Extensive experiments on five public datasets demonstrate that AMAD works well and robustly on various real-world scenarios, which outperforms state-of-the-art baseline methods by 16% and 11% in terms of recall and F1, with only 3% of data being labeled.
Rongwei Yu, Wang Wang
Comput. Secur.3
2023 A Multi-source Domain Adaption Approach to Minority Disk Failure Prediction
Wang Wang, Xuehai Tang, Biyu Zhou, Yangchen Dong, Yuanhang Feng, Jizhong Han, Songlin Hu 0001
ICA3PP (2)1
2022 Improving disk failure detection accuracy via data augmentation
abstract
Frequently happening of disk failures seriously affects the dependability and service quality of cloud data centers. Recently, machine learning (ML) based methods are popularly adopted to proactively predict forthcoming disk failures via supervised learning. However, the high imbalance of failure samples and healthy samples is a huge obstacle for existing detection methods to establish high performance detection model. This paper presents a data augmentation method MSGMD, which can efficiently generate high quality failure samples to alleviate the data imbalance of the training set, so as to effectively improve the performance of any supervised failure detection models. First, MSGMD converts failure samples (multivariate time series) into multiple univariate time series via decomposing the spatial relations among features. Then it learns the temporal correlation of each feature via a policy-based reinforcement learning model trained in an adversarial way. After that, it generates failure samples by combining feature series sampled from learned distribution. Finally, it filters out low quality generated samples with a confidence-based method. Experimental results on real-world datasets show that, through data augmentation, MSGMD can improve the FDR and F1-Score of the state-of-the-art disk failure detection model by 31.59% and 30.74% respectively on average.
Wang Wang, Xuehai Tang, Biyu Zhou, Wenjie Xiao, Jizhong Han, Songlin Hu 0001
IWQoS1
2022 Statistical Observations of Three Co-Existing NBTI Behaviors in 28 nm HKMG by On-Chip Monitor With Less Recovery Impact
abstract
An on-chip digital sensor has been demonstrated in 28nm High-k Metal Gate (HKMG) for bias temperature instability (BTI) statistical characterization with the benefits: fast statistical measurement, less recovery impact (Toff-stress@around 15ns, Fast Period Sampling (FPS) @around 300ns), and high resolution (0.1mV of$\Delta $Vth). As far as we know, it is the first time to statistically observe the very early stage of trap recovery of individual device in practical scenario, e.g., static random-access memory (SRAM). We find that three Negative BTI (NBTI) recovery behaviors, 2/3/4-step with clear transition slope, co-exist in HKMG devices. Our further analysis ascribes the phenomena to co-existing of four types of defects in 28nm HKMG Devices Under Test (DUTs). Three types are recoverable and one unrecoverable. The transition slope instead of steep drop between steps is the aggregative effects of one certain type of recoverable defect contained across DUTs. More types of defects lead to more Vth shift. But the contribution percentage of unrecoverable defect remains quite close, while recoverable defects dominate the Vth degradation. Only when the Toff-stress is less than the starting point of 1st recover step (within 1$\mu \text{s}$in our case), accurate and consistent Vth degradation data can be achieved.
Yarong Fu, Wang Wang, Manni Li, Yinyin Lin
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 Fed-Tra: Improving Accuracy of Deep Learning Model on Non-iid in Federated Learning
Wenjie Xiao, Xuehai Tang, Biyu Zhou, Wang Wang, Yangchen Dong, Liangjun Zang, Jizhong Han, Songlin Hu 0001
ICA3PP (1)4
2020 Tail: An Automated and Lightweight Gradient Compression Framework for Distributed Deep Learning
abstract
Existing gradient compression schemes fail to automatically determine the compression ratio or are accompanied by high compression overhead. To address this, we present Tail, an automated and lightweight gradient compression framework stacked by three modules, quantization, sparsification, and encoding. Without any hand-tuned effort, quantization module automatically adjusts the compression ratio along training iterations to retain accuracy first. Then, sparsification and encoding modules are successively applied to the quantized gradient to further improve compression ratio. Moreover, Tail reduces the compression overhead by approximate computing in the automated decision-making process. Experiments validate that Tail can reduce communication traffic by an order of magnitude while retaining or even improving model accuracy.
Jinrong Guo, Songlin Hu 0001, Wang Wang, Chunrong Yao, Jizhong Han, Ruixuan Li 0001
DAC3
2020 Accelerating Distributed Deep Learning By Adaptive Gradient Quantization
abstract
To accelerate distributed deep learning, gradient quantization technique is widely used to reduce the communication cost. However, the existing quantization schemes suffer from either model accuracy degradation or low compression ratio (arisen from a redundant setting of quantization level or high overhead in determining the level). In this work, we propose a novel adaptive quantization scheme (AdaQS) to explore the balance between model accuracy and quantization level. AdaQS determines the quantization level automatically according to gradient's mean to standard deviation ratio (MSDR). Then, to reduce the quantization overhead, we employ a computationally-friendly way of moment estimation to calculate the MSDR. Finally, theoretical analysis of AdaQS's convergence is conducted for non-convex objectives. Experiments demonstrate that AdaQS performs excellently on very deep model GoogleNet with 2.55% accuracy improvement relative to vanilla SGD and achieves 1.8x end-to-end speedup on AlexNet in a distributed cluster with 4*4 GPUs.
Jinrong Guo, Wantao Liu, Wang Wang, Jizhong Han, Ruixuan Li 0001, Songlin Hu 0001
ICASSP3
2019 AccUDNN: A GPU Memory Efficient Accelerator for Training Ultra-Deep Neural Networks
abstract
With the implementation of mainstream DL frameworks, scarce GPU memory resource is the primary bottleneck that hinders the trainability and training efficiency of ultra-deep neural networks (UDNN). Prior memory optimization works focus on removing the trainability restriction but leave the training efficiency out of consideration. To fill the gap, we present "AccUDNN", an accelerator that aims to make full use of finite GPU memory resource to speed up the training process of UDNN in this paper. AccUDNN mainly includes two modules: memory optimizer and hyperparameter tuner. Memory optimizer develops a novel performance-model guided dynamic swap out/in strategy to meet trainability first and further remedy the efficiency degradation in other swapping strategies. Then, a hyperparameter tuner is designed to explore the efficiency-optimal minibatch size and the matched learning rate after applying the dynamic swapping strategy. Evaluations demonstrate that AccUDNN cuts down the GPU memory requirement of ResNet-152 from more than 24GB to 8GB. In turn, given 12GB GPU memory budget, the efficiency-optimal minibatch size can reach 4.2x larger than Caffe and finally improve the scaling efficiency (speedup) of 8 GPUs' cluster by 1.9x.
Jinrong Guo, Wantao Liu, Wang Wang, Chunrong Yao, Jizhong Han, Ruixuan Li 0001, Songlin Hu 0001
ICCD3
2019 A GPU memory efficient speed-up scheme for training ultra-deep neural networks: poster
abstract
Ultra-deep neural network(UDNN) tends to yield higher-quality model but its training process is often difficult to handle. Scarce GPU DRAM capacity is the primary bottleneck that limits the depth of neural network and the range of trainable minibatch size. In this paper, we present a scheme that dedicates to make the utmost use of finite GPU memory resource to speed up the training process for UDNN. Firstly, a performance-model guided dynamic swap out/in strategy between GPU and host memory is carefully orchestrated to tackle the out-of-memory problem without introducing performance penalty. Then, a hyperparameter (minibatch size, learning rate) tuning policy is designed to explore the optimal configuration after applying the swap strategy from the perspectives of training time and final accuracy simultaneously. Finally, we verify the effectiveness of our scheme in both single and distributed GPU mode.
Jinrong Guo, Wantao Liu, Wang Wang, Qu Lu, Songlin Hu 0001, Jizhong Han, Ruixuan Li 0001
PPoPP3