EDBT 2026 Demo / reviewers in the wild / expert
Tao Li 0054
dblp:75/4601-54
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-8010-1447ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Trustworthy machine learning · 54% Deep learning architectures and training · 15% Efficient and distributed learning · 14% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
machine unlearning |
0.9 | 1 | 2025 | Towards Natural Machine Unlearning · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Machine learning › Trustworthy machine learning
privacy and data protection |
0.9 | 1 | 2025 | Towards Natural Machine Unlearning · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection |
0.8 | 1 | 2024 | Low-Dimensional Gradient Helps Out-of-Distribution Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Machine learning › Trustworthy machine learning
robustness |
0.8 | 1 | 2024 | Low-Dimensional Gradient Helps Out-of-Distribution Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Machine learning › Optimization for machine learning › gradient-based optimization
sharpness-aware minimization |
0.8 | 1 | 2024 | Friendly Sharpness-Aware Minimization · CVPR 2024 |
Machine learning › Deep learning architectures and training
training optimization |
0.8 | 1 | 2024 | Friendly Sharpness-Aware Minimization · CVPR 2024 |
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction |
0.7 | 1 | 2023 | Low Dimensional Trajectory Hypothesis is True: DNNs Can Be Trained in Tiny Subspaces · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Machine learning › Efficient and distributed learning
efficient training |
0.7 | 1 | 2023 | Trainable Weight Averaging: Efficient Training by Optimizing Historical Solutions · ICLR 2023 |
Machine learning › Efficient and distributed learning
model compression |
0.7 | 1 | 2023 | Low Dimensional Trajectory Hypothesis is True: DNNs Can Be Trained in Tiny Subspaces · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Machine learning › Deep learning architectures and training
weight averaging |
0.7 | 1 | 2023 | Trainable Weight Averaging: Efficient Training by Optimizing Historical Solutions · ICLR 2023 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.6 | 1 | 2022 | Subspace Adversarial Training · CVPR 2022 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training |
0.6 | 1 | 2022 | Subspace Adversarial Training · CVPR 2022 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness › adversarial training › fast adversarial training
catastrophic overfitting |
0.6 | 1 | 2022 | Subspace Adversarial Training · CVPR 2022 |
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
robustness to label noise |
0.2 | 1 | 2023 | Low Dimensional Trajectory Hypothesis is True: DNNs Can Be Trained in Tiny Subspaces · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Machine learning › Optimization for machine learning
constrained optimization |
0.2 | 1 | 2022 | Subspace Adversarial Training · CVPR 2022 |
Machine learning › Optimization for machine learning › optimization
subspace optimization |
0.2 | 1 | 2022 | Subspace Adversarial Training · CVPR 2022 |
Methods — techniques the papers use, named apart from their topics
relabeling · 0.9fine-tuning · 0.9stochastic gradient noise · 0.8principal component analysis · 0.8linear dimension reduction · 0.8exponential moving average · 0.8convergence analysis · 0.8weight averaging · 0.7stochastic optimization · 0.7dynamic linear dimensionality reduction · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-head ensemble of smoothed classifiers for certified robustness
Kun Fang 0004, Qinghua Tao, Yingwen Wu, Tao Li 0054, Xiaolin Huang, Jie Yang 0002 |
Neural Networks | 4 |
| 2025 | Towards Natural Machine UnlearningabstractMachine unlearning (MU) aims to eliminate information that has been learned from specific training data, namely forgetting data, from a pretrained model. Currently, the mainstream of relabeling-based MU methods involves modifying the forgetting data with incorrect labels and subsequently fine-tuning the model. While learning such incorrect information can indeed remove knowledge, the process is quite unnatural as the unlearning process undesirably reinforces the incorrect information and leads to over-forgetting. Towards more natural machine unlearning, we inject correct information from the remaining data to the forgetting samples when changing their labels. Through pairing these adjusted samples with their labels, the model tends to use the injected correct information and naturally suppresses the information meant to be forgotten. Albeit straightforward, such a first step towards natural machine unlearning can significantly outperform current state-of-the-art approaches. In particular, our method substantially reduces the over-forgetting problem and leads to strong robustness across different unlearning tasks, making it a promising candidate for practical machine unlearning. Zhengbao He, Tao Li 0054, Xinwen Cheng, Zhehao Huang, Xiaolin Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Friendly Sharpness-Aware MinimizationabstractSharpness-Aware Minimization (SAM) has been instrumental in improving deep neural network training by minimizing both training loss and loss sharpness. Despite the practical success, the mechanisms behind SAM's generalization enhancements remain elusive, limiting its progress in deep learning optimization. In this work, we investigate SAM's core components for generalization improvement and introduce “Friendly-SAM” (F-SAM) to further enhance SAM's generalization. Our investigation reveals the key role of batch-specific stochastic gradient noise within the adversarial perturbation, i.e., the current minibatcli gradient, which significantly influences SAM's generalization performance. By decomposing the adversarial perturbation in SAM into full gradient and stochastic gradient noise components, we discover that relying solely on the full gradient component degrades generalization while excluding it leads to improved performance. The possible reason lies in the full gradient component's increase in sharpness loss for the entire dataset, creating inconsistencies with the subsequent sharpness minimization step solely on the current minibatcli data. Inspired by these insights, F-SAM aims to mitigate the negative effects of the full gradient component. It removes the full gradient estimated by an exponentially moving average (EMA) of historical stochastic gradients, and then leverages stochastic gradient noise for improved generalization. Moreover, we provide theoretical validation for the EMA approximation and prove the convergence of F-SAM on non-convex problems. Extensive experiments demonstrate the superior generalization performance and robustness of F-SAM over vanilla SAM. Code is available at https://github.com/nblt/F-SAM. Tao Li 0054, Pan Zhou 0002, Zhengbao He, Xinwen Cheng, Xiaolin Huang |
CVPR | 1 |
| 2024 | Low-Dimensional Gradient Helps Out-of-Distribution DetectionabstractDetecting out-of-distribution (OOD) samples is essential for ensuring the reliability of deep neural networks (DNNs) in real-world scenarios. While previous research has predominantly investigated the disparity between in-distribution (ID) and OOD data through forward information analysis, the discrepancy in parameter gradients during the backward process of DNNs has received insufficient attention. Existing studies on gradient disparities mainly focus on the utilization of gradient norms, neglecting the wealth of information embedded in gradient directions. To bridge this gap, in this paper, we conduct a comprehensive investigation into leveraging the entirety of gradient information for OOD detection. The primary challenge arises from the high dimensionality of gradients due to the large number of network parameters. To solve this problem, we propose performing linear dimension reduction on the gradient using a designated subspace that comprises principal components. This innovative technique enables us to obtain a low-dimensional representation of the gradient with minimal information loss. Subsequently, by integrating the reduced gradient with various existing detection score functions, our approach demonstrates superior performance across a wide range of detection tasks. For instance, on the ImageNet benchmark with ResNet50 model, our method achieves an average reduction of 11.15 % in the false positive rate at 95 % recall (FPR95) compared to the current state-of-the-art approach. Yingwen Wu, Tao Li 0054, Xinwen Cheng, Jie Yang 0002, Xiaolin Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Towards robust neural networks via orthogonal diversity
Kun Fang 0004, Qinghua Tao, Yingwen Wu, Tao Li 0054, Feipeng Cai, Xiaolin Huang, Jie Yang 0002 |
Pattern Recognit. | 4 |
| 2023 | Trainable Weight Averaging: Efficient Training by Optimizing Historical Solutions
Tao Li 0054, Zhehao Huang, Qinghua Tao, Yingwen Wu, Xiaolin Huang |
ICLR | 1 |
| 2023 | Low Dimensional Trajectory Hypothesis is True: DNNs Can Be Trained in Tiny SubspacesabstractDeep neural networks (DNNs) usually contain massive parameters, but there is redundancy such that it is guessed that they could be trained in low-dimensional subspaces. In this paper, we propose a Dynamic Linear Dimensionality Reduction (DLDR) based on the low-dimensional properties of the training trajectory. The reduction method is efficient, supported by comprehensive experiments: optimizing DNNs in 40-dimensional spaces can achieve comparable performance as regular training over thousands or even millions of parameters. Since there are only a few variables to optimize, we develop an efficient quasi-Newton-based algorithm, obtain robustness to label noise, and improve the performance of well-trained models, which are three follow-up experiments that can show the advantages of finding such low-dimensional subspaces. The code is released (Pytorch: https://github.com/nblt/DLDR and Mindspore: https://gitee.com/mindspore/docs/tree/r1.6/docs/sample_code/dimension_reduce_training). Tao Li 0054, Zhehao Huang, Qinghua Tao, Yipeng Liu 0001, Xiaolin Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Subspace Adversarial TrainingabstractSingle-step adversarial training (AT) has received wide attention as it proved to be both efficient and robust. However, a serious problem of catastrophic overfitting exists, i.e., the robust accuracy against projected gradient descent (PGD) attack suddenly drops to 0% during the training. In this paper, we approach this problem from a novel perspective of optimization and firstly reveal the close link between the fast-growing gradient of each sample and overfitting, which can also be applied to understand robust overfitting in multi-step AT. To control the growth of the gradient, we propose a new AT method, Subspace Adversarial Training (Sub-AT), which constrains AT in a carefully extracted subspace. It successfully resolves both kinds of overfitting and significantly boosts the robustness. In subspace, we also allow single-step AT with larger steps and larger radius, further improving the robustness performance. As a result, we achieve state-of-the-art single-step AT performance. Without any regularization term, our single-step AT can reach over 51 % robust accuracy against strong PGD-50 attack of radius 8/255 on CIFAR-10, reaching a competitive performance against standard multi-step PGD-10 AT with huge computational advantages. The code is released at https://github.com/nblt/Sub-AT. Tao Li 0054, Yingwen Wu, Sizhe Chen, Kun Fang 0004, Xiaolin Huang |
CVPR | 1 |