EDBT 2026 Demo / reviewers in the wild / expert
Qinghua Tao
dblp:182/9643
· DBLP profile ↗
21ranked-venue papers
6as first author
19since 2021 · last 2025
0000-0001-9705-7748ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 3 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-head ensemble of smoothed classifiers for certified robustness
Kun Fang 0004, Qinghua Tao, Yingwen Wu, Tao Li 0054, Xiaolin Huang, Jie Yang 0002 |
Neural Networks | 2 |
| 2024 | Feature Learning using Multi-view Kernel Partial Least SquaresabstractThe multi-view learning deals with data of multiple views, aiming to explore the underlying relations between different views and use them for various tasks.In this paper, we derive a multi-view extension of kernel partial least squares for unsupervised feature learning.We establish the optimization objective in the primal as the pairwise covariance between the projection scores and derive that this model can be trained in the dual form by solving an eigenvalue problem.Experiments are also conducted to verify the effectiveness of the method with real-life multi-view datasets, where the proposed method is adopted as a feature extractor and then the clustering task is conducted for performance comparisons. Xinjie Zeng, Qinghua Tao, Johan A. K. Suykens |
ESANN | 2 |
| 2024 | Self-Attention through Kernel-Eigen Pair Sparse Variational Gaussian ProcessesabstractWhile the great capability of Transformers significantly boosts prediction accuracy, it could also yield overconfident predictions and require calibrated uncertainty estimation, which can be commonly tackled by Gaussian processes (GPs). Existing works apply GPs with symmetric kernels under variational inference to the attention kernel; however, omitting the fact that attention kernels are in essence asymmetric. Moreover, the complexity of deriving the GP posteriors remains high for large-scale data. In this work, we propose Kernel-Eigen Pair Sparse Variational Gaussian Processes (KEP-SVGP) for building uncertainty-aware self-attention where the asymmetry of attention kernels is tackled by Kernel SVD (KSVD) and a reduced complexity is acquired. Through KEP-SVGP, i) the SVGP pair induced by the two sets of singular vectors from KSVD w.r.t. the attention kernel fully characterizes the asymmetry; ii) using only a small set of adjoint eigenfunctions from KSVD, the derivation of SVGP posteriors can be based on the inversion of a diagonal matrix containing singular values, contributing to a reduction in time complexity; iii) an evidence lower bound is derived so that variational parameters and network weights can be optimized with it. Experiments verify our excellent performances and efficiency on in-distribution, distribution-shift and out-of-distribution benchmarks. Yingyi Chen, Qinghua Tao, Francesco Tonin, Johan A. K. Suykens |
ICML | 2 |
| 2024 | Learning in Feature Spaces via Coupled Covariances: Asymmetric Kernel SVD and Nyström methodabstractIn contrast with Mercer kernel-based approaches as used e.g. in Kernel Principal Component Analysis (KPCA), it was previously shown that Singular Value Decomposition (SVD) inherently relates to asymmetric kernels and Asymmetric Kernel Singular Value Decomposition (KSVD) has been proposed. However, the existing formulation to KSVD cannot work with infinite-dimensional feature mappings, the variational objective can be unbounded, and needs further numerical evaluation and exploration towards machine learning. In this work, i) we introduce a new asymmetric learning paradigm based on coupled covariance eigenproblem (CCE) through covariance operators, allowing infinite-dimensional feature maps. The solution to CCE is ultimately obtained from the SVD of the induced asymmetric kernel matrix, providing links to KSVD. ii) Starting from the integral equations corresponding to a pair of coupled adjoint eigenfunctions, we formalize the asymmetric Nyström method through a finite sample approximation to speed up training. iii) We provide the first empirical evaluations verifying the practical utility and benefits of KSVD and compare with methods resorting to symmetrization or linear SVD across multiple tasks. Qinghua Tao, Francesco Tonin, Alex Lambert, Yingyi Chen, Panagiotis Patrinos, Johan A. K. Suykens |
ICML | 1 |
| 2024 | Kernel PCA for Out-of-Distribution DetectionabstractOut-of-Distribution (OoD) detection is vital for the reliability of Deep Neural Networks (DNNs).
Existing works have shown the insufficiency of Principal Component Analysis (PCA) straightforwardly applied on the features of DNNs in detecting OoD data from In-Distribution (InD) data.
The failure of PCA suggests that the network features residing in OoD and InD are not well separated by simply proceeding in a linear subspace, which instead can be resolved through proper non-linear mappings.
In this work, we leverage the framework of Kernel PCA (KPCA) for OoD detection, and seek suitable non-linear kernels that advocate the separability between InD and OoD data in the subspace spanned by the principal components.
Besides, explicit feature mappings induced from the devoted task-specific kernels are adopted so that the KPCA reconstruction error for new test samples can be efficiently obtained with large-scale data.
Extensive theoretical and empirical results on multiple OoD data sets and network structures verify the superiority of our KPCA detector in efficiency and efficacy with state-of-the-art detection performance. Kun Fang 0004, Qinghua Tao, Kexin Lv, Mingzhen He, Xiaolin Huang, Jie Yang 0002 |
NeurIPS | 2 |
| 2024 | Revisiting Deep Ensemble for Out-of-Distribution Detection: A Loss Landscape Perspective
Kun Fang 0004, Qinghua Tao, Xiaolin Huang, Jie Yang 0002 |
Int. J. Comput. Vis. | 2 |
| 2024 | Deep Kernel Principal Component Analysis for multi-level feature learningabstractPrincipal Component Analysis (PCA) and its nonlinear extension Kernel PCA (KPCA) are widely used across science and industry for data analysis and dimensionality reduction. Modern deep learning tools have achieved great empirical success, but a framework for deep principal component analysis is still lacking. Here we develop a deep kernel PCA methodology (DKPCA) to extract multiple levels of the most informative components of the data. Our scheme can effectively identify new hierarchical variables, called deep principal components, capturing the main characteristics of high-dimensional data through a simple and interpretable numerical optimization. We couple the principal components of multiple KPCA levels, theoretically showing that DKPCA creates both forward and backward dependency across levels, which has not been explored in kernel methods and yet is crucial to extract more informative features. Various experimental evaluations on multiple data types show that DKPCA finds more efficient and disentangled representations with higher explained variance in fewer principal components, compared to the shallow KPCA. We demonstrate that our method allows for effective hierarchical data exploration, with the ability to separate the key generative factors of the input data both for large datasets and when few training samples are available. Overall, DKPCA can facilitate the extraction of useful patterns from high-dimensional data by learning more informative features organized in different levels, giving diversified aspects to explore the variation factors in the data, while maintaining a simple mathematical formulation. Francesco Tonin, Qinghua Tao, Panagiotis Patrinos, Johan A. K. Suykens |
Neural Networks | 2 |
| 2024 | Towards robust neural networks via orthogonal diversity
Kun Fang 0004, Qinghua Tao, Yingwen Wu, Tao Li 0054, Feipeng Cai, Xiaolin Huang, Jie Yang 0002 |
Pattern Recognit. | 2 |
| 2023 | Measuring the Transferability of ℓ∞ Attacks by the ℓ2 NormabstractDeep neural networks could be fooled by adversarial examples with trivial differences to original samples. To keep the difference imperceptible in human eyes, researchers bound the adversarial perturbations by the ℓ∞norm, which is now commonly served as the standard to align the strength of different attacks for a fair comparison. However, we propose that using the ℓ∞norm alone is not sufficient in measuring the attack strength, because even with a fixed ℓ∞distance, the ℓ2distance also greatly affects the attack transferability between models. Through the discovery, we reach more in-depth understandings towards the attack mechanism, i.e., several existing methods attack black-box models better partly because they craft perturbations with 70% to 130% larger ℓ2distances. Since larger perturbations naturally lead to better transferability, we thereby advocate that the strength of attacks should be simultaneously measured by both the ℓ∞and ℓ2norm. Our proposal is firmly supported by extensive experiments on ImageNet dataset from 7 attacks, 4 white-box models, and 9 black-box models. Sizhe Chen, Qinghua Tao, Zhixing Ye, Xiaolin Huang |
ICASSP | 2 |
| 2023 | Tensorized LSSVMS For Multitask RegressionabstractMultitask learning (MTL) can utilize the relatedness between multiple tasks for performance improvement. The advent of multimodal data allows tasks to be referenced by multiple indices. High-order tensors are capable of providing efficient representations for such tasks, while preserving structural task-relations. In this paper, a new MTL method is proposed by leveraging low-rank tensor analysis and constructing tensorized Least Squares Support Vector Machines, namely the tLSSVM-MTL, where multilinear modelling and its nonlinear extensions can be flexibly exerted. We employ a high-order tensor for all the weights with each mode relating to an index and factorize it with CP decomposition, assigning a shared factor for all tasks and retaining task-specific latent factors along each index. Then an alternating algorithm is derived for the nonconvex optimization, where each resulting subproblem is solved by a linear system. Experimental results demonstrate promising performances of our tLSSVM-MTL. Jiani Liu 0002, Qinghua Tao, Ce Zhu, Yipeng Liu 0001, Johan A. K. Suykens |
ICASSP | 2 |
| 2023 | Trainable Weight Averaging: Efficient Training by Optimizing Historical Solutions
Tao Li 0054, Zhehao Huang, Qinghua Tao, Yingwen Wu, Xiaolin Huang |
ICLR | 3 |
| 2023 | Primal-Attention: Self-attention through Asymmetric Kernel SVD in Primal RepresentationabstractRecently, a new line of works has emerged to understand and improve self-attention in Transformers by treating it as a kernel machine. However, existing works apply the methods for symmetric kernels to the asymmetric self-attention, resulting in a nontrivial gap between the analytical understanding and numerical implementation. In this paper, we provide a new perspective to represent and optimize self-attention through asymmetric Kernel Singular Value Decomposition (KSVD), which is also motivated by the low-rank property of self-attention normally observed in deep layers. Through asymmetric KSVD, i) a primal-dual representation of self-attention is formulated, where the optimization objective is cast to maximize the projection variances in the attention outputs; ii) a novel attention mechanism, i.e., Primal-Attention, is proposed via the primal representation of KSVD, avoiding explicit computation of the kernel matrix in the dual; iii) with KKT conditions, we prove that the stationary solution to the KSVD optimization in Primal-Attention yields a zero-value objective. In this manner, KSVD optimization can be implemented by simply minimizing a regularization loss, so that low-rank property is promoted without extra decomposition. Numerical experiments show state-of-the-art performance of our Primal-Attention with improved efficiency. Moreover, we demonstrate that the deployed KSVD optimization regularizes Primal-Attention with a sharper singular value decay than that of the canonical self-attention, further verifying the great potential of our method. To the best of our knowledge, this is the first work that provides a primal-dual representation for the asymmetric kernel in self-attention and successfully applies it to modelling and optimization. Yingyi Chen, Qinghua Tao, Francesco Tonin, Johan A. K. Suykens |
NeurIPS | 2 |
| 2023 | Low Dimensional Trajectory Hypothesis is True: DNNs Can Be Trained in Tiny SubspacesabstractDeep neural networks (DNNs) usually contain massive parameters, but there is redundancy such that it is guessed that they could be trained in low-dimensional subspaces. In this paper, we propose a Dynamic Linear Dimensionality Reduction (DLDR) based on the low-dimensional properties of the training trajectory. The reduction method is efficient, supported by comprehensive experiments: optimizing DNNs in 40-dimensional spaces can achieve comparable performance as regular training over thousands or even millions of parameters. Since there are only a few variables to optimize, we develop an efficient quasi-Newton-based algorithm, obtain robustness to label noise, and improve the performance of well-trained models, which are three follow-up experiments that can show the advantages of finding such low-dimensional subspaces. The code is released (Pytorch: https://github.com/nblt/DLDR and Mindspore: https://gitee.com/mindspore/docs/tree/r1.6/docs/sample_code/dimension_reduce_training). Tao Li 0054, Zhehao Huang, Qinghua Tao, Yipeng Liu 0001, Xiaolin Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Jigsaw-ViT: Learning jigsaw puzzles in vision transformerabstractThe success of Vision Transformer (ViT) in various computer vision tasks has promoted the ever-increasing prevalence of this convolution-free network. The fact that ViT works on image patches makes it potentially relevant to the problem of jigsaw puzzle solving, which is a classical self-supervised task aiming at reordering shuffled sequential image patches back to their original form. Solving jigsaw puzzle has been demonstrated to be helpful for diverse tasks using Convolutional Neural Networks (CNNs), such as feature representation learning, domain generalization and fine-grained classification. In this paper, we explore solving jigsaw puzzle as a self-supervised auxiliary loss in ViT for image classification, named Jigsaw-ViT. We show two modifications that can make Jigsaw-ViT superior to standard ViT: discarding positional embeddings and masking patches randomly. Yet simple, we find that the proposed Jigsaw-ViT is able to improve on both generalization and robustness over the standard ViT, which is usually rather a trade-off. Numerical experiments verify that adding the jigsaw puzzle branch provides better generalization to ViT on large-scale image classification on ImageNet. Moreover, such auxiliary loss also improves robustness against noisy labels on Animal-10N, Food-101N, and Clothing1M, as well as adversarial examples. Our implementation is available at https://yingyichen-cyy.github.io/Jigsaw-ViT. Yingyi Chen, Xi Shen 0001, Qinghua Tao, Johan A. K. Suykens |
Pattern Recognit. Lett. | 4 |
| 2022 | Adversarial Attack on Attackers: Post-Process to Mitigate Black-Box Score-Based Query AttacksabstractThe score-based query attacks (SQAs) pose practical threats to deep neural networks by crafting adversarial perturbations within dozens of queries, only using the model's output scores. Nonetheless, we note that if the loss trend of the outputs is slightly perturbed, SQAs could be easily misled and thereby become much less effective. Following this idea, we propose a novel defense, namely Adversarial Attack on Attackers (AAA), to confound SQAs towards incorrect attack directions by slightly modifying the output logits. In this way, (1) SQAs are prevented regardless of the model's worst-case robustness; (2) the original model predictions are hardly changed, i.e., no degradation on clean accuracy; (3) the calibration of confidence scores can be improved simultaneously. Extensive experiments are provided to verify the above advantages. For example, by setting $\ell_\infty=8/255$ on CIFAR-10, our proposed AAA helps WideResNet-28 secure 80.59% accuracy under Square attack (2500 queries), while the best prior defense (i.e., adversarial training) only attains 67.44%. Since AAA attacks SQA's general greedy strategy, such advantages of AAA over 8 defenses can be consistently observed on 8 CIFAR-10/ImageNet models under 6 SQAs, using different attack targets, bounds, norms, losses, and strategies. Moreover, AAA calibrates better without hurting the accuracy. Our code is available at https://github.com/Sizhe-Chen/AAA. Sizhe Chen, Zhehao Huang, Qinghua Tao, Yingwen Wu, Cihang Xie, Xiaolin Huang |
NeurIPS | 3 |
| 2022 | Short-Term Traffic Flow Prediction Based on the Efficient Hinging Hyperplanes Neural NetworkabstractTraffic flow (TF) prediction is an important and yet a challenging task in transportation systems, since the TF involves high nonlinearities and is affected by many elements. Recently, neural networks have attracted much attention for TF prediction, but they are commonly black boxes with complex architectures and difficult to be interpreted, e.g., the contributions of specific traffic elements are not explicit, hardly providing informative guidance. In this paper, we aim at addressing more interpretable short-term TF prediction with joint consideration to high accuracy, and thus introduces a pragmatic method by applying the efficient hinging hyperplanes neural network (EHHNN) simply built upon sparse neuron connections. In the proposed method, different traffic factors are incorporated into the inputs, including their spatial-temporal information. Besides the pursuit of accuracy, we further extend the ANOVA decomposition of EHHNNs to the interpretation analysis with specifications to traffic data, in which the contributions concerning specific traffic variables are detected quantitatively. As such, the proposed method firstly applies the EHHNN to filter out more important traffic variables for dimensionality reduction while maintaining accurate prediction. Then, variable interpretation analysis is performed from different perspectives, e.g. to quantitatively investigate the influence of traffic factors and also their spatial-temporal impacts. Therefore, a predictor and an analyzing tool can both be attained for the TF by exerting the flexibility and extending the interpretability of EHHNNs, which is promising to provide informative guidance to future traffic control. Numerical experiments verify the effectiveness and potential of the proposed method in TF prediction and analysis. Qinghua Tao, Zhen Li 0032, Jun Xu 0008, Shu Lin 0002, Bart De Schutter, Johan A. K. Suykens |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Center-Aware Adversarial Autoencoder for Anomaly DetectionabstractAnomaly detection based on subspace learning has attracted much attention, in which the compactness of subspace is commonly considered as the core concern. Most related studies directly optimize the distance from the subspace representation to the fixed center, and the influence of the anomaly level of each normal sample is not considered to adjust the normal concentrated areas. In such cases, it is difficult to isolate the normal areas from the anomaly ones by making the subspace compact. To this end, we propose a center-aware adversarial autoencoder (CA-AAE) method, which detects anomaly samples by acquiring more compact and discriminative subspace representations. To fully exploit the subspace information to improve the compactness, anomaly-level description and feature learning are novelly integrated herein by dividing the output space of the encoder into presubspace and postsubspace. In presubspace, the toward-center prior distribution is imposed by the adversarial learning mechanism, and the anomaly level of normal samples can be described from a probabilistic perspective. In postsubspace, a novel center-aware strategy is established to enhance the compactness of the postsubspace, which achieves adaptive adjustment of the normal areas by constructing a weighted center based on the anomaly level. Then, a flexible anomaly score function is constructed in the testing stage, in which both the toward-center loss and the reconstruction loss are combined to balance the information in the learned subspace and the original space. Compared to other related methods, the proposed CA-AAE shows the effectiveness and advantages in numerical experiments. Daoming Li, Qinghua Tao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Toward Deep Adaptive Hinging HyperplanesabstractThe adaptive hinging hyperplane (AHH) model is a popular piecewise linear representation with a generalized tree structure and has been successfully applied in dynamic system identification. In this article, we aim to construct the deep AHH (DAHH) model to extend and generalize the networking of AHH model for high-dimensional problems. The network structure of DAHH is determined through a forward growth, in which the activity ratio is introduced to select effective neurons and no connecting weights are involved between the layers. Then, all neurons in the DAHH network can be flexibly connected to the output in a skip-layer format, and only the corresponding weights are the parameters to optimize. With such a network framework, the backpropagation algorithm can be implemented in DAHH to efficiently tackle large-scale problems and the gradient vanishing problem is not encountered in the training of DAHH. In fact, the optimization problem of DAHH can maintain convexity with convex loss in the output layer, which brings natural advantages in optimization. Different from the existing neural networks, DAHH is easier to interpret, where neurons are connected sparsely and analysis of variance (ANOVA) decomposition can be applied, facilitating to revealing the interactions between variables. A theoretical analysis toward universal approximation ability and explicit domain partitions are also derived. Numerical experiments verify the effectiveness of the proposed DAHH. Qinghua Tao, Jun Xu 0008, Zhen Li 0032, Na Xie, Shuning Wang, Xiaoli Li 0011, Johan A. K. Suykens |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Learning with continuous piecewise linear decision trees
Qinghua Tao, Zhen Li 0032, Jun Xu 0008, Na Xie, Shuning Wang, Johan A. K. Suykens |
Expert Syst. Appl. | 1 |
| 2017 | Adaptive block coordinate DIRECT algorithm
Qinghua Tao, Xiaolin Huang, Shuning Wang, Li Li 0013 |
J. Glob. Optim. | 1 |
| 2016 | Multiple Gaussian graphical estimation with jointly sparse penalty
Qinghua Tao, Xiaolin Huang, Shuning Wang, Xiangming Xi, Li Li 0013 |
Signal Process. | 1 |