Weiguo Gao

dblp:07/2810 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Theory of computation · 3 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Brief Announcement: An I/O-Efficient Parallel FFT for Heterogeneous Architectures via a Single Global Exchange
abstract
Fast Fourier transforms are increasingly constrained by data movement: on modern many-core CPUs and accelerators, the cost of redistributing layouts can dominate the nominal O(N log N) arithmetic. This paper presents PACO, a Processor-Aware Communication-Optimal FFT framework that fuses the transpose-like and digit-reversal layout changes of a slab-distributed FFT into a single global permutation. PACO follows a fixed three-stage pipeline, LocalFFT-OneGlobalPermutation-LocalFFT: the local stages use an SCO (semi-cache-oblivious) cube FFT with lazy transpositions, while the middle stage is implemented by PARA-CO-DRP, a load-balanced parallel cache-oblivious digit-reversal permutation.
Shina Guo, Weiguo Gao
SPAA2
2026 Exploiting intra-core heterogeneity for high-performance dense linear algebra on Huawei Ascend 910 NPUs
Shina Guo, Weiguo Gao, Xinghui Tian
CCF Trans. High Perform. Comput.2
2026 Toward theoretical insights into diffusion trajectory distillation via operator merging
Weiguo Gao, Ming Li 0074
Neural Networks1
2025 A Mixed Precision Jacobi SVD Algorithm
abstract
We propose a mixed precision Jacobi algorithm for computing the singular value decomposition (SVD) of a dense matrix. After appropriate preconditioning, the proposed algorithm computes the SVD in a lower precision as an initial guess and then performs one-sided Jacobi rotations in the working precision as iterative refinement. By carefully transforming a lower precision solution to a higher precision one, our algorithm achieves about \(2\times\) speedup on the x86-64 architecture compared to the usual one-sided Jacobi SVD algorithm in LAPACK, without sacrificing the accuracy.
Weiguo Gao, Yuxin Ma 0002, Meiyue Shao
ACM Trans. Math. Softw.1
2024 Decentralized Natural Policy Gradient with Variance Reduction for Collaborative Multi-Agent Reinforcement Learning
abstract
This paper studies a policy optimization problem arising from collaborative multi-agent reinforcement learning in a decentralized setting where agents communicate with their neighbors over an undirected graph to maximize the sum of their cumulative rewards. A novel decentralized natural policy gradient method, dubbed Momentum-based Decentralized Natural Policy Gradient (MDNPG), is proposed, which incorporates natural gradient, momentum-based variance reduction, and gradient tracking into the decentralized stochastic gradient ascent framework. The $\mathcal{O}(n^{-1}\epsilon^{-3})$ sample complexity for MDNPG to converge to an $\epsilon$-stationary point has been established under standard assumptions, where $n$ is the number of agents. It indicates that MDNPG can achieve the optimal convergence rate for decentralized policy gradient methods and possesses a linear speedup in contrast to centralized optimization methods. Moreover, superior empirical performance of MDNPG over other state-of-the-art algorithms has been demonstrated by extensive numerical experiments.
Jinchi Chen, Weiguo Gao, Ke Wei 0001
J. Mach. Learn. Res.3
2024 Serving DNN Inference With Fine-Grained Spatio-Temporal Sharing of GPU Servers
abstract
Deep Neural Networks(DNNs) are commonly deployed as online inference services. To meet interactive latency requirements of requests, DNN services require the use ofGraphics Processing Unit(GPU) to improve their responsiveness. The unique characteristics of inference workloads pose new challenges to manage GPU resources. First, the GPU scheduler needs to carefully manage requests to meet their latency targets. Second, a single inference task often underutilizes GPU resources. Third, the fluctuating patterns of inference workloads pose difficulties in determining the resources allocated to each DNN model. Therefore, it is critical for the GPU scheduler to maximize GPU utilization by collocating multiple DNN models without violating the latencyService-Level Objectives(SLOs) of requests. However, we find that existing works are not adequate for achieving this goal among latency-sensitive inference tasks. Hence, we propose FineST, a scheduling framework for serving DNNs with fine-grained spatio-temporal sharing of GPU inference servers. To maximize GPU utilization, FineST allocates intra-GPU computing resources from both spatial and temporal dimensions across DNNs in a cost-effective way, while predicting interference overheads under diverse consolidated executions for controlling SLO violation rates. Compared to a state-of-the-art work, FineST improves the peak throughput of serving heterogeneous DNNsby up to 64.7% under SLO constraints.
Yaqiong Peng, Weiguo Gao, Haocheng Peng
IEEE Trans. Serv. Comput.2
2023 Controllable Music Inpainting with Mixed-Level and Disentangled Representation
abstract
Music inpainting, which is to complete the missing part of a piece given some context, is an important task of automated music generation. In this study, we contribute a controllable inpainting model by combining the high expressivity of mixed-level, disentangled music representations and the strong predictive power of masked language modeling. The model enables flexible user controls over both time scope (inpainted length and location) and semantic features that composers often consider during composition, say rhythm pattern and chords. The key model design is to simultaneously predict disentangled representations of different time ranges. Such design aims to mirror the thought process of a professional composer who can take into account of the music flow of various semantic features at different hierarchies in parallel. Results show that our model produces much higher quality music compared to the baseline, and the subjective evaluation shows that our model generates much better results than the baseline and can generate melodies that are similar to human composition.
Shiqi Wei, Ziyu Wang 0008, Weiguo Gao, Gus Xia
ICASSP3
2023 A New Framework of Swarm Learning Consolidating Knowledge From Multi-Center Non-IID Data for Medical Image Segmentation
abstract
Large training datasets are important for deep learning-based methods. For medical image segmentation, it could be however difficult to obtain large number of labeled training images solely from one center. Distributed learning, such as swarm learning, has the potential to use multi-center data without breaching data privacy. However, data distributions across centers can vary a lot due to the diverse imaging protocols and vendors (known as feature skew). Also, the regions of interest to be segmented could be different, leading to inhomogeneous label distributions (referred to as label skew). With such non-independently and identically distributed (Non-IID) data, the distributed learning could result in degraded models. In this work, we propose a novel swarm learning approach, which assembles local knowledge from each center while at the same time overcomes forgetting of global knowledge during local training. Specifically, the approach first leverages a label skew-awared loss to preserve the global label knowledge, and then aligns local feature distributions to consolidate global knowledge against local feature skew. We validated our method in three Non-IID scenarios using four public datasets, including the Multi-Centre, Multi-Vendor and Multi-Disease Cardiac Segmentation (M&Ms) dataset, the Federated Tumor Segmentation (FeTS) dataset, the Multi-Modality Whole Heart Segmentation (MMWHS) dataset and the Multi-Site Prostate T2-weighted MRI segmentation (MSProsMRI) dataset. Results show that our method could achieve superior performance over existing methods. Code will be released via https://zmiclab.github.io/projects.html once the paper gets accepted.
Zheyao Gao, Fuping Wu, Weiguo Gao, Xiahai Zhuang
IEEE Trans. Medical Imaging3
2022 Music Phrase Inpainting Using Long-Term Representation and Contrastive Loss
abstract
Deep generative modeling has already become the leading technique for music automation. However, long-term generation remains a challenging task as most methods fall short in preserving a natural structure and the overall musicality when the generation scope exceeds several beats. In this study, we tackle the problem of long-term, phrase-level symbolic melody inpainting by equipping a sequence prediction model with phrase-level representation (as an extra condition) and contrastive loss (as an extra optimization term). The underlying ideas are twofold. First, to predict phrase-level music, we need phrase-level representations as a better context. Second, we should predict notes and their high-level representations simultaneously, while contrastive loss serves as a better target for abstract representations. Experimental results show that our method significantly outperforms the baselines. In particular, contrastive loss plays a critical role in the generation quality, and the phase-level representation further enhances the structure of long-term generation.1
Shiqi Wei, Gus Xia, Yixiao Zhang 0002, Weiguo Gao
ICASSP5
2022 Vectorized Hankel Lift: A Convex Approach for Blind Super-Resolution of Point Sources
abstract
We consider the problem of resolving$r$point sources from$n$samples at the low end of the spectrum when point spread functions (PSFs) are not known. Assuming that the spectrum samples of the PSFs lie in low dimensional subspace (let$s$denote the dimension), we can formulate it as a matrix recovery problem, followed by location estimation. By exploiting the low rank structure of the vectorized Hankel matrix associated with the target matrix, a convex approach called Vectorized Hankel Lift is proposed for the matrix recovery. It is shown that$n\gtrsim rs\log ^{4} n$samples are sufficient for Vectorized Hankel Lift to achieve the exact recovery. For the location retrieval from the matrix, applying the single snapshot MUSIC method within the vectorized Hankel lift framework corresponds to the spatial smoothing technique proposed to improve the performance of the MMV MUSIC for the direction-of-arrival (DOA) estimation.
Jinchi Chen, Weiguo Gao, Sihan Mao, Ke Wei 0001
IEEE Trans. Inf. Theory2
2021 A Diversity-Enhanced and Constraints-Relaxed Augmentation for Low-Resource Classification
Guang Liu 0007, Hailong Huang 0003, Yuzhao Mao, Weiguo Gao, Jianping Shen
DASFAA (2)4
2021 Adversarial Mixing Policy for Relaxing Locally Linear Constraints in Mixup
abstract
Mixup is a recent regularizer for current deep classification networks.Through training a neural network on convex combinations of pairs of examples and their labels, it imposes locally linear constraints on the model's input space.However, such strict linear constraints often lead to under-fitting which degrades the effects of regularization.Noticeably, this issue is getting more serious when the resource is extremely limited.To address these issues, we propose the Adversarial Mixing Policy (AMP), organized in a "min-max-rand" formulation, to relax the Locally Linear Constraints in Mixup.Specifically, AMP adds a small adversarial perturbation to the mixing coefficients rather than the examples.Thus, slight non-linearity is injected in-between the synthetic examples and synthetic labels.By training on these data, the deep networks are further regularized, and thus achieve a lower predictive error rate.Experiments on five text classification benchmarks and five backbone models have empirically shown that our methods reduce the error rate over Mixup variants in a significant margin (up to 31.3%), especially in low-resource conditions (up to 17.5%).
Guang Liu 0007, Yuzhao Mao, Hailong Huang 0003, Weiguo Gao
EMNLP (1)4
2021 Processor-Aware Cache-Oblivious Algorithms✱
abstract
Frigo et al. proposed an ideal cache model and a recursive technique to design sequential cache-efficient algorithms in a cache-oblivious fashion. Ballard et al. pointed out that it is a fundamental open problem to extend the technique to an arbitrary architecture. Ballard et al. raised another open question on how to parallelize Strassen’s algorithm exactly and efficiently on an arbitrary number of processors.
Weiguo Gao
ICPP2
2021 SOFT: Softmax-free Transformer with Linear Complexity
abstract
Vision transformers (ViTs) have pushed the state-of-the-art for various visual recognition tasks by patch-wise image tokenization followed by self-attention. However, the employment of self-attention modules results in a quadratic complexity in both computation and memory usage. Various attempts on approximating the self-attention computation with linear complexity have been made in Natural Language Processing. However, an in-depth analysis in this work shows that they are either theoretically flawed or empirically ineffective for visual recognition. We further identify that their limitations are rooted in keeping the softmax self-attention during approximations. Specifically, conventional self-attention is computed by normalizing the scaled dot-product between token feature vectors. Keeping this softmax operation challenges any subsequent linearization efforts. Based on this insight, for the first time, a softmax-free transformer or SOFT is proposed. To remove softmax in self-attention, Gaussian kernel function is used to replace the dot-product similarity without further normalization. This enables a full self-attention matrix to be approximated via a low-rank matrix decomposition. The robustness of the approximation is achieved by calculating its Moore-Penrose inverse using a Newton-Raphson method. Extensive experiments on ImageNet show that our SOFT significantly improves the computational efficiency of existing ViT variants. Crucially, with a linear complexity, much longer token sequences are permitted in SOFT, resulting in superior trade-off between accuracy and complexity.
Jinghan Yao, Junge Zhang, Xiatian Zhu, Hang Xu 0004, Weiguo Gao, Chunjing Xu, Tao Xiang 0002, Li Zhang 0040
NeurIPS6
2021 Penalized -regression-based bicluster localization
Hanjia Gao, Zheng-Jian Bai, Weiguo Gao, Shuqin Zhang
Pattern Recognit.3
2020 Multi-scale Two-way Deep Neural Network for Stock Trend Prediction
abstract
Stock Trend Prediction(STP) has drawn wide attention from various fields, especially Artificial Intelligence. Most previous studies are single-scale oriented which results in information loss from a multi-scale perspective. In fact, multi-scale behavior is vital for making intelligent investment decisions. A mature investor will thoroughly investigate the state of a stock market at various time scales. To automatically learn the multi-scale information in stock data, we propose a Multi-scale Two-way Deep Neural Network. It learns multi-scale patterns from two types of scale-information, wavelet-based and downsampling-based, by eXtreme Gradient Boosting and Recurrent Convolutional Neural Network, respectively. After combining the learned patterns from the two-way, our model achieves state-of-the-art performance on FI-2010 and CSI-2016, where the latter is our published long-range stock dataset to help future studies for STP task. Extensive experimental results on the two datasets indicate that multi-scale information can significantly improve the STP performance and our model is superior in capturing such information.
Guang Liu 0007, Yuzhao Mao, Hailong Huang 0003, Weiguo Gao, Jianping Shen, Ruifan Li, Xiaojie Wang 0006
IJCAI5
2011 Large scale plane wave pseudopotential density functional theory calculations on GPU clusters
abstract
In this work, we present our implementation of the density functional theory (DFT) plane wave pseudopotential (PWP) calculations on GPU clusters. This GPU version is developed based on a CPU DFT-PWP code: PEtot, which can calculate up to a thousand atoms on thousands of processors. Our test indicates that the GPU version can have a ~10 times speed-up over the CPU version. A detail analysis of the speed-up and the scaling on the number of CPU/GPU computing units (up to 256) are presented. The success of our speed-up relies on the adoption a hybrid reciprocal-space and band-index parallelization scheme. As far as we know, this is the first GPU DFT-PWP code scalable to large number of CPU/GPU computing units. We also outlined the future work, and what is needed to further increase the computational speed by another factor of 10.
Long Wang 0008, Weile Jia, Weiguo Gao, Xuebin Chi, Lin-Wang Wang
SC4
2008 An Implementation and Evaluation of the AMLS Method for Sparse Eigenvalue Problems
abstract
We describe an efficient implementation and present a performance study of an automated multi-level substructuring (AMLS) method for sparse eigenvalue problems. We assess the time and memory requirements associated with the key steps of the algorithm, and compare it with the shift-and-invert Lanczos algorithm. Our eigenvalue problems come from two very different application areas: accelerator cavity design and normal-mode vibrational analysis of polyethylene particles. We show that the AMLS method, when implemented carefully, outperforms the traditional method in broad application areas when large numbers of eigenvalues are sought, with relatively low accuracy.
Weiguo Gao, Xiaoye S. Li, Chao Yang 0001, Zhaojun Bai
ACM Trans. Math. Softw.1