VLDB 2026 Research / reviewers in the wild / expert
Wei Tao 0002
dblp:17/6159-2
· DBLP profile ↗
14ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-8273-6649ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Computer networks · 1Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Incorporating prior knowledge into style embedding for unsupervised text style transfer
Yahao Hu, Wei Tao 0002, Yifei Xie 0001, Zhisong Pan 0003 |
Comput. Speech Lang. | 2 |
| 2026 | Optimizing the Adversarial Perturbation With a Momentum-Based Adaptive MatrixabstractGenerating adversarial examples (AEs) can be formulated as an optimization problem. Among various optimization-based attacks, the gradient-based PGD and the momentum-based MI-FGSM have garnered considerable interest. However, all these attacks use the sign function to scale their perturbations, which raises several theoretical concerns from the point of view of optimization. In this paper, we first reveal that PGD is actually a specific reformulation of the projected gradient method using only the current gradient to determine its step-size. Further, we show that when we utilize a conventional adaptive matrix with the accumulated gradients to scale the perturbation, PGD becomes AdaGrad. Motivated by this analysis, we present a novel momentum-based attack AdaMI, in which the perturbation is optimized with an interesting momentum-based adaptive matrix. AdaMI is proved to attain optimal convergence for convex problems, indicating that it addresses the non-convergence issue of MI-FGSM, thereby ensuring stability of the optimization process. The experiments demonstrate that the proposed momentum-based adaptive matrix can serve as a general and effective technique to boost adversarial transferability over the state-of-the-art methods across different networks while maintaining better stability and imperceptibility. Wei Tao 0002, Xin Liu 0042, Wei Li 0116, Qing Tao 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2024 | On the Convergence of an Adaptive Momentum Method for Adversarial AttacksabstractAdversarial examples are commonly created by solving a constrained optimization problem, typically using sign-based methods like Fast Gradient Sign Method (FGSM). These attacks can benefit from momentum with a constant parameter, such as Momentum Iterative FGSM (MI-FGSM), to enhance black-box transferability. However, the monotonic time-varying momentum parameter is required to guarantee convergence in theory, creating a theory-practice gap. Additionally, recent work shows that sign-based methods fail to converge to the optimum in several convex settings, exacerbating the issue. To address these concerns, we propose a novel method which incorporates both an innovative adaptive momentum parameter without monotonicity assumptions and an adaptive step-size scheme that replaces the sign operation. Furthermore, we derive a regret upper bound for general convex functions. Experiments on multiple models demonstrate the efficacy of our method in generating adversarial examples with human-imperceptible noise while achieving high attack success rates, indicating its superiority over previous adversarial example generation methods. Wei Tao 0002, Shuohao Li |
AAAI | 2 |
| 2024 | Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
Xin Liu 0042, Wei Tao 0002, Wei Li 0116, Dazhi Zhan, Zhisong Pan 0003 |
IJCAI | 2 |
| 2024 | Location and time embedded feature representation for spatiotemporal traffic prediction
Wei Li 0116, Xin Liu 0042, Wei Tao 0002, Lei Zhang 0126, Junhua Zou, Zhisong Pan 0002 |
Expert Syst. Appl. | 3 |
| 2023 | Token-level disentanglement for unsupervised text style transfer
Yahao Hu, Wei Tao 0002, Yifei Xie 0001, Zhisong Pan 0003 |
Neurocomputing | 2 |
| 2022 | A convergence analysis of Nesterov's accelerated gradient method in training deep linear neural networks
Xin Liu 0042, Wei Tao 0002 |
Inf. Sci. | 2 |
| 2022 | Provable convergence of Nesterov's accelerated gradient method for over-parameterized neural networks
Xin Liu 0042, Wei Tao 0002 |
Knowl. Based Syst. | 3 |
| 2022 | Momentum Acceleration in the Individual Convergence of Nonsmooth Convex Optimization With ConstraintsabstractMomentum technique has recently emerged as an effective strategy in accelerating convergence of gradient descent (GD) methods and exhibits improved performance in deep learning as well as regularized learning. Typical momentum examples include Nesterov's accelerated gradient (NAG) and heavy-ball (HB) methods. However, so far, almost all the acceleration analyses are only limited to NAG, and a few investigations about the acceleration of HB are reported. In this article, we address the convergence about the last iterate of HB in nonsmooth optimizations with constraints, which we name individual convergence. This question is significant in machine learning, where the constraints are required to impose on the learning structure and the individual output is needed to effectively guarantee this structure while keeping an optimal rate of convergence. Specifically, we prove that HB achieves an individual convergence rate of O(1/√t) , where t is the number of iterations. This indicates that both of the two momentum methods can accelerate the individual convergence of basic GD to be optimal. Even for the convergence of averaged iterates, our result avoids the disadvantages of the previous work in restricting the optimization problem to be unconstrained as well as limiting the performed number of iterations to be predefined. The novelty of convergence analysis presented in this article provides a clear understanding of how the HB momentum can accelerate the individual convergence and reveals more insights about the similarities and differences in getting the averaging and individual convergence rates. The derived optimal individual convergence is extended to regularized and stochastic settings, in which an individual solution can be produced by the projection-based operation. In contrast to the averaged output, the sparsity can be reduced remarkably without sacrificing the theoretical optimal rates. Several real experiments demonstrate the performance of HB momentum strategy. Wei Tao 0002, Qing Tao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Gradient Descent Averaging and Primal-dual Averaging for Strongly Convex OptimizationabstractAveraging scheme has attracted extensive attention in deep learning as well as traditional machine learning. It achieves theoretically optimal convergence and also improves the empirical model performance. However, there is still a lack of sufficient convergence analysis for strongly convex optimization. Typically, the convergence about the last iterate of gradient descent methods, which is referred to as individual convergence, fails to attain its optimality due to the existence of logarithmic factor. In order to remove this factor, we first develop gradient descent averaging (GDA), which is a general projection-based dual averaging algorithm in the strongly convex setting. We further present primal-dual averaging for strongly convex cases (SC-PDA), where primal and dual averaging schemes are simultaneously utilized. We prove that GDA yields the optimal convergence rate in terms of output averaging, while SC-PDA derives the optimal individual convergence. Several experiments on SVMs and deep learning models validate the correctness of theoretical analysis and effectiveness of algorithms. Wei Tao 0002, Wei Li 0116, Zhisong Pan 0003, Qing Tao 0001 |
AAAI | 1 |
| 2021 | The Role of Momentum Parameters in the Optimal Convergence of Adaptive Polyak's Heavy-ball Methods
Wei Tao 0002, Qing Tao 0001 |
ICLR | 1 |
| 2020 | Regularized shapelet learning for scalable time series classification
Huiyun Zhao, Zhisong Pan 0003, Wei Tao 0002 |
Comput. Networks | 3 |
| 2020 | Primal Averaging: A New Gradient Evaluation Step to Attain the Optimal Individual ConvergenceabstractMany well-known first-order gradient methods have been extended to cope with large-scale composite problems, which often arise as a regularized empirical risk minimization in machine learning. However, their optimal convergence is attained only in terms of the weighted average of past iterative solutions. How to make the individual convergence of stochastic gradient descent (SGD) optimal, especially for strongly convex problems has now become a challenging problem in the machine learning community. On the other hand, Nesterov's recent weighted averaging strategy succeeds in achieving the optimal individual convergence of dual averaging (DA) but it fails in the basic mirror descent (MD). In this paper, a new primal averaging (PA) gradient operation step is presented, in which the gradient evaluation is imposed on the weighted average of all past iterative solutions. We prove that simply modifying the gradient operation step in MD by PA strategy suffices to recover the optimal individual rate for general convex problems. Along this line, the optimal individual rate of convergence for strongly convex problems can also be achieved by imposing the strong convexity on the gradient operation step. Furthermore, we extend PA-MD to solve regularized nonsmooth learning problems in the stochastic setting, which reveals that PA strategy is a simple yet effective extra step toward the optimal individual convergence of SGD. Several real experiments on sparse learning and SVM problems verify the correctness of our theoretical analysis. Wei Tao 0002, Zhisong Pan 0003, Qing Tao 0001 |
IEEE Trans. Cybern. | 1 |
| 2020 | The Strength of Nesterov's Extrapolation in the Individual Convergence of Nonsmooth OptimizationabstractThe extrapolation strategy raised by Nesterov, which can accelerate the convergence rate of gradient descent methods by orders of magnitude when dealing with smooth convex objective, has led to tremendous success in training machine learning tasks. In this article, the convergence of individual iterates of projected subgradient (PSG) methods for nonsmooth convex optimization problems is theoretically studied based on Nesterov's extrapolation, which we name individual convergence. We prove that Nesterov's extrapolation has the strength to make the individual convergence of PSG optimal for nonsmooth problems. In light of this consideration, a direct modification of the subgradient evaluation suffices to achieve optimal individual convergence for strongly convex problems, which can be regarded as making an interesting step toward the open question about stochastic gradient descent (SGD) posed by Shamir. Furthermore, we give an extension of the derived algorithms to solve regularized learning tasks with nonsmooth losses in stochastic settings. Compared with other state-of-the-art nonsmooth methods, the derived algorithms can serve as an alternative to the basic SGD especially in coping with machine learning problems, where an individual output is needed to guarantee the regularization structure while keeping an optimal rate of convergence. Typically, our method is applicable as an efficient tool for solving large-scale l1-regularized hinge-loss learning problems. Several comparison experiments demonstrate that our individual output not only achieves an optimal convergence rate but also guarantees better sparsity than the averaged solution. Wei Tao 0002, Zhisong Pan 0003, Qing Tao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |