Xiaolin Huang

dblp:61/2227 · DBLP profile ↗
← Back
152ranked-venue papers
18as first author
86since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 94 · 10 first-author · 61 since 2021Graphics, computer vision, multimedia, augmented reality and games · 53 · 6 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1
YearPublicationVenuePosition
2026 Correction to: MUSO: achieving exact machine unlearning in over‑parameterized regimes
Ruikai Yang, Mingzhen He, Zhengbao He, Youmei Qiu, Xiaolin Huang
Mach. Learn.5
2026 Data imputation by pursuing better classification: A supervised kernel-based method
Ruikai Yang, Mingzhen He, Xiaolin Huang
Pattern Recognit.5
2025 ParseCaps: An Interpretable Parsing Capsule Network for Medical Image Diagnosis
abstract
Deep learning has excelled in medical image classification, but its clinical application is limited by poor interpretability. Capsule networks, known for encoding hierarchical relationships and spatial features, show potential in addressing this issue. Nevertheless, traditional capsule networks often underperform due to their shallow structures, and deeper variants lack hierarchical architectures, thereby compromising interpretability. This paper introduces a novel capsule network, ParseCaps, which utilizes the sparse axial attention routing and parse convolutional capsule layer to form a parse-tree-like structure, enhancing both depth and interpretability. Firstly, sparse axial attention routing optimizes connections between child and parent capsules, as well as emphasizes the weight distribution across instantiation parameters of parent capsules. Secondly, the parse convolutional capsule layer generates capsule predictions aligning with the parse tree. Finally, based on the loss design that is effective whether concept ground truth exists or not, ParseCaps advances interpretability by associating each dimension of the global capsule with a comprehensible concept, thereby facilitating clinician trust and understanding of the model's classification results. Experimental results on three medical datasets show that ParseCaps not only outperforms other capsule network variants in classification accuracy and robustness, but also provides interpretable explanations, regardless of the availability of concept labels.
Xinyu Geng, Xiaolin Huang, Fanglin Chen 0001, Jun Xu 0008
AAAI3
2025 Simulating Training Dynamics to Reconstruct Training Data from Deep Neural Networks
abstract
Whether deep neural networks (DNNs) memorize the training data is a fundamental open question in understanding deep learning. A direct way to verify the memorization of DNNs is to reconstruct training data from DNNs’ parameters. Since parameters are gradually determined by data throughout training, characterizing training dynamics is important for reconstruction. Pioneering works rely on the linear training dynamics of shallow NNs with large widths, but cannot be extended to more practical DNNs which have non-linear dynamics. We propose Simulation of training Dynamics (SimuDy) to reconstruct training data from DNNs. Specifically, we simulate the training dynamics by training the model from the initial parameters with a dummy dataset, then optimize this dummy dataset so that the simulated dynamics reach the same final parameters as the true dynamics. By incorporating dummy parameters in the simulated dynamics, SimuDy effectively describes non-linear training dynamics. Experiments demonstrate that SimuDy significantly outperforms previous approaches when handling non-linear training dynamics, and for the first time, most training samples can be reconstructed from a trained ResNet’s parameters.
Hanling Tian, Yuhang Liu 0003, Mingzhen He, Zhengbao He, Zhehao Huang, Ruikai Yang, Xiaolin Huang
ICLR7
2025 Pursuing Feature Separation based on Neural Collapse for Out-of-Distribution Detection
abstract
In the open world, detecting out-of-distribution (OOD) data, whose labels are disjoint with those of in-distribution (ID) samples, is important for reliable deep neural networks (DNNs). To achieve better detection performance, one type of approach proposes to fine-tune the model with auxiliary OOD datasets to amplify the difference between ID and OOD data through a separation loss defined on model outputs. However, none of these studies consider enlarging the feature disparity, which should be more effective compared to outputs. The main difficulty lies in the diversity of OOD samples, which makes it hard to describe their feature distribution, let alone design losses to separate them from ID features. In this paper, we neatly fence off the problem based on an aggregation property of ID features named Neural Collapse (NC). NC means that the penultimate features of ID samples within a class are nearly identical to the last layer weight of the corresponding class. Based on this property, we propose a simple but effective loss called Separation Loss, which binds the features of OOD data in a subspace orthogonal to the principal subspace of ID features formed by NC. In this way, the features of ID and OOD samples are separated by different dimensions. By optimizing the feature separation loss rather than purely enlarging output differences, our detection achieves SOTA performance on CIFAR10, CIFAR100 and ImageNet benchmarks without any additional data augmentation or sampling, demonstrating the importance of feature separation in OOD detection. Code is available at https://github.com/Wuyingwen/Pursuing-Feature-Separation-for-OOD-Detection.
Yingwen Wu, Ruiji Yu, Xinwen Cheng, Zhengbao He, Xiaolin Huang
ICLR5
2025 Primphormer: Efficient Graph Transformers with Primal Representations
abstract
Graph Transformers (GTs) have emerged as a promising approach for graph representation learning. Despite their successes, the quadratic complexity of GTs limits scalability on large graphs due to their pair-wise computations. To fundamentally reduce the computational burden of GTs, we propose a primal-dual framework that interprets the self-attention mechanism on graphs as a dual representation. Based on this framework, we develop Primphormer, an efficient GT that leverages a primal representation with linear complexity. Theoretical analysis reveals that Primphormer serves as a universal approximator for functions on both sequences and graphs, while also retaining its expressive power for distinguishing non-isomorphic graphs. Extensive experiments on various graph benchmarks demonstrate that Primphormer achieves competitive empirical results while maintaining a more user-friendly memory and computational costs.
Mingzhen He, Ruikai Yang, Hanling Tian, Youmei Qiu, Xiaolin Huang
ICML5
2025 Flat-LoRA: Low-Rank Adaptation over a Flat Loss Landscape
abstract
Fine-tuning large-scale pre-trained models is prohibitively expensive in terms of computation and memory costs. Low-Rank Adaptation (LoRA), a popular Parameter-Efficient Fine-Tuning (PEFT) method, offers an efficient solution by optimizing only low-rank matrices. Despite recent progress in improving LoRA’s performance, the relationship between the LoRA optimization space and the full parameter space is often overlooked. A solution that appears flat in the loss landscape of the LoRA space may still exhibit sharp directions in the full parameter space, potentially compromising generalization. We introduce Flat-LoRA, which aims to identify a low-rank adaptation situated in a flat region of the full parameter space. Instead of adopting the well-established sharpness-aware minimization approach, which incurs significant computation and memory overheads, we employ a Bayesian expectation loss objective to preserve training efficiency. Further, we design a refined strategy for generating random perturbations to enhance performance and carefully manage memory overhead using random seeds. Experiments across diverse tasks—including mathematical reasoning, coding abilities, dialogue generation, instruction following, and text-to-image generation—demonstrate that Flat-LoRA improves both in-domain and out-of-domain generalization. Code is available at https://github.com/nblt/Flat-LoRA.
Zhengbao He, Yasheng Wang, Lifeng Shang, Xiaolin Huang
ICML6
2025 SEB-Naver: A SE(2)-based Local Navigation Framework for Car-like Robots on Uneven Terrain
abstract
Autonomous navigation of car-like robots on uneven terrain poses unique challenges compared to flat terrain, particularly in traversability assessment and terrain-associated kinematic modelling for motion planning. This paper introduces SEB-Naver, a novel SE(2)-based local navigation framework designed to overcome these challenges. First, we propose an efficient traversability assessment method for SE(2) grids, leveraging GPU parallel computing to enable real-time updates and maintenance of local maps. Second, inspired by differential flatness, we present an optimization-based trajectory planning method that integrates terrain-associated kinematic models, significantly improving both planning efficiency and trajectory quality. Finally, we unify these components into SEB-Naver, achieving real-time terrain assessment and trajectory optimization. Extensive simulations and real-world experiments demonstrate the effectiveness and efficiency of our approach. The code is at https://github.com/ZJU-FAST-Lab/seb_naver.
Long Xu 0002, Xiaolin Huang, Donglai Xue, Zhichao Han 0002, Chao Xu 0001, Yanjun Cao, Fei Gao 0011
IROS3
2025 Stimulating Catastrophic Forgetting in Class-Wise Unlearning via UAP
Wenxing Zhou, Xinwen Cheng, Yingwen Wu, Ruikai Yang, Xiaolin Huang
ECML/PKDD (5)5
2025 MUSO: achieving exact machine unlearning in over-parameterized regimes
Ruikai Yang, Mingzhen He, Zhenghao He, Youmei Qiu, Xiaolin Huang
Mach. Learn.5
2025 Multi-head ensemble of smoothed classifiers for certified robustness
Kun Fang 0004, Qinghua Tao, Yingwen Wu, Tao Li 0054, Xiaolin Huang, Jie Yang 0002
Neural Networks5
2025 Towards Natural Machine Unlearning
abstract
Machine unlearning (MU) aims to eliminate information that has been learned from specific training data, namely forgetting data, from a pretrained model. Currently, the mainstream of relabeling-based MU methods involves modifying the forgetting data with incorrect labels and subsequently fine-tuning the model. While learning such incorrect information can indeed remove knowledge, the process is quite unnatural as the unlearning process undesirably reinforces the incorrect information and leads to over-forgetting. Towards more natural machine unlearning, we inject correct information from the remaining data to the forgetting samples when changing their labels. Through pairing these adjusted samples with their labels, the model tends to use the injected correct information and naturally suppresses the information meant to be forgotten. Albeit straightforward, such a first step towards natural machine unlearning can significantly outperform current state-of-the-art approaches. In particular, our method substantially reduces the over-forgetting problem and leads to strong robustness across different unlearning tasks, making it a promising candidate for practical machine unlearning.
Zhengbao He, Tao Li 0054, Xinwen Cheng, Zhehao Huang, Xiaolin Huang
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 A Decentralized Framework for Kernel PCA With Projection Consensus Constraints
abstract
This paper studies kernel PCA in a decentralized setting, where data are distributively observed with full features in local nodes, and a fusion center is prohibited. Compared with linear PCA, the use of kernel brings challenges to the design of decentralized consensus optimization: the local projection directions are data-dependent. As a result, the consensus constraint in distributed linear PCA is no longer valid. To overcome this problem, we propose a projection consensus constraint and obtain an effective decentralized consensus framework, where local solutions are expected to be the projection of the global solution on the column space of the local dataset. We also derive a fully non-parametric, fast, and convergent algorithm based on the alternative direction method of multiplier, of which each iteration is analytic and communication-efficient. Experiments on a truly parallel architecture are conducted on real-world data, showing that the proposed decentralized algorithm is effective in utilizing information from other nodes and takes great advantages in running time over the central kernel PCA.
Ruikai Yang, Lei Shi 0010, Xiaolin Huang
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 MOOD: Leveraging Out-of-Distribution Data to Enhance Imbalanced Semi-Supervised Learning
abstract
The imbalanced semi-supervised learning (SSL) has emerged as a critical research area due to the prevalence of class imbalanced and partially labeled data in real-world scenarios. As the requirement for data volume increases, naturally collected datasets inevitably contain out-of-distribution (OOD) samples. However, the performance of existing imbalanced SSL methods experiences a marked deterioration with OOD data. In this article, we propose an imbalanced SSL method called mixup-OOD (MOOD) to address this issue. The core idea is to "turn waste into treasure," exploring the potential of leveraging seemingly detrimental OOD data to expand the feature space, particularly for tail classes. Specifically, we first filter OOD data from unlabeled data, and then fuse it with labeled data to boost feature diversity for the tail classes. To avoid feature overlapping with OOD data, we develop a push-and-pull (PaP) loss to attract in-distribution (ID) instances toward respective class centroids while repelling OOD samples from them. Extensive experiments show that MOOD achieves superior performance compared with other state-of-the-art methods and exhibits robustness across data with different imbalanced ratios and OOD proportions. The source code is available at: https://github.com/xlhuang132/MOODv2.
Yang Lu 0009, Xiaolin Huang, Mengke Li 0001, Yan Yan 0001, Chen Gong 0002, Hanzi Wang
IEEE Trans. Neural Networks Learn. Syst.2
2025 Decentralized Kernel Ridge Regression Based on Data-Dependent Random Feature
abstract
Random feature (RF) has been widely used for node consistency in decentralized kernel ridge regression (KRR). Currently, the consistency is guaranteed by imposing constraints on coefficients of features, necessitating that the RFs on different nodes are identical. However, in many applications, data on different nodes vary significantly on the number or distribution, which calls for adaptive and data-dependent methods that generate different RFs. To tackle the essential difficulty, we propose a new decentralized KRR algorithm that pursues consensus on decision functions, which allows great flexibility and well adapts data on nodes. The convergence is rigorously given, and the effectiveness is numerically verified: by capturing the characteristics of the data on each node, while maintaining the same communication costs as other methods, we achieved an average regression accuracy improvement of 25.5% across six real-world datasets.
Ruikai Yang, Mingzhen He, Jie Yang 0002, Xiaolin Huang
IEEE Trans. Neural Networks Learn. Syst.5
2024 Progressive Stepwise Diffusion Model with Dual Decoders for Semi-Supervised Medical Image Segmentation
abstract
Semi-supervised medical image segmentation tasks aim to harness the potential of vast amounts of unlabeled data using a limited amount of annotated data. Denoising Diffusion Probabilistic Models, which have achieved significant success in image generation, are gradually being explored for their potential in semantic image segmentation. However, their application in semi-supervised medical image segmentation is still in its early stages. Initially, due to the high randomness of diffusion models, the pseudo-labels generated during the early training phase may mislead the processing of unlabeled data. Additionally, the use of fixed-time steps for random sampling during training limits the ability of the model to learn effective denoising functions at an early stage. To address these issues, we propose an innovative framework named Progressive Stepwise Diffusion Network with Dual Decoders (PSDD) for semi-supervised medical image segmentation. This framework incorporates an additional normal decoder into the denoising diffusion encoder-decoder structure to provide more accurate labels and employs a Progressive Incremental Step strategy to gradually train the model for longer generation processes. Evaluated on two 2D colon polyp segmentation datasets and a 3D Left Atrium dataset, the experimental results demonstrate significant performance improvements over current advanced methods, thereby validating the effectiveness and potential of this framework in handling complex semi-supervised learning scenarios.
Xiaolin Huang, Jingchun Lin, Bingzhi Chen, Guangming Lu 0002
BIBM1
2024 OrthCaps: An Orthogonal CapsNet with Sparse Attention Routing and Pruning
abstract
Redundancy is a persistent challenge in Capsule Networks (CapsNet), leading to high computational costs and parameter counts. Although previous studies have introduced pruning after the initial capsule layer, dynamic routing's fully connected nature and non-orthogonal weight matrices reintroduce redundancy in deeper layers. Besides, dynamic routing requires iterating to converge, further increasing computational demands. In this paper, we propose an Orthogonal Capsule Network (OrthCaps) to reduce redundancy, improve routing performance and decrease parameter counts. Firstly, an efficient pruned capsule layer is introduced to discard redundant capsules. Secondly, dynamic routing is replaced with orthogonal sparse attention routing, eliminating the need for iterations and fully connected structures. Lastly, weight matrices during routing are orthogonalized to sustain low capsule similarity, which is the first approach to use Householder orthogonal decomposition to enforce orthogonality in CapsNet. Our experiments on baseline datasets affirm the efficiency and robustness of OrthCaps in classification tasks, in which ablation studies validate the criticality of each component. OrthCaps-Shallow outperforms other Capsule Network benchmarks on four datasets, utilizing only 110k parameters - a mere 1.25% of a standard Capsule Network's total. To the best of our knowledge,$it$achieves the smallest parameter count among existing Capsule Networks. Similarly, OrthCaps-Deep demonstrates competitive performance across four datasets, utilizing only 1.2% of the parameters required by its counterparts.
Xinyu Geng, Jiawei Gong, Yuerong Xue, Jun Xu 0008, Fanglin Chen 0001, Xiaolin Huang
CVPR7
2024 Friendly Sharpness-Aware Minimization
abstract
Sharpness-Aware Minimization (SAM) has been instrumental in improving deep neural network training by minimizing both training loss and loss sharpness. Despite the practical success, the mechanisms behind SAM's generalization enhancements remain elusive, limiting its progress in deep learning optimization. In this work, we investigate SAM's core components for generalization improvement and introduce “Friendly-SAM” (F-SAM) to further enhance SAM's generalization. Our investigation reveals the key role of batch-specific stochastic gradient noise within the adversarial perturbation, i.e., the current minibatcli gradient, which significantly influences SAM's generalization performance. By decomposing the adversarial perturbation in SAM into full gradient and stochastic gradient noise components, we discover that relying solely on the full gradient component degrades generalization while excluding it leads to improved performance. The possible reason lies in the full gradient component's increase in sharpness loss for the entire dataset, creating inconsistencies with the subsequent sharpness minimization step solely on the current minibatcli data. Inspired by these insights, F-SAM aims to mitigate the negative effects of the full gradient component. It removes the full gradient estimated by an exponentially moving average (EMA) of historical stochastic gradients, and then leverages stochastic gradient noise for improved generalization. Moreover, we provide theoretical validation for the EMA approximation and prove the convergence of F-SAM on non-convex problems. Extensive experiments demonstrate the superior generalization performance and robustness of F-SAM over vanilla SAM. Code is available at https://github.com/nblt/F-SAM.
Tao Li 0054, Pan Zhou 0002, Zhengbao He, Xinwen Cheng, Xiaolin Huang
CVPR5
2024 Learning Scalable Model Soup on a Single GPU: An Efficient Subspace Training Strategy
Weisen Jiang, Fanghui Liu 0001, Xiaolin Huang, James T. Kwok
ECCV (65)4
2024 Phase Retrieval by Tensor Total Least Squares
abstract
Phase retrieval seeks to reconstruct a series of image sequences from measurements that only capture their magnitudes. Current approaches either flatten and stack the image sequences, disregarding their multidimensional structural information, or fail to account for errors within the sensing vectors/tensors. To address these two issues simultaneously, we propose a unified framework for the phase retrieval problem, namely tensor total least squares (TTLS). Specifically, we set up a tensor representation for image sequences and the corresponding measurement model, and for the first time employ the advanced tensor ring network to effectively explore the inherent multidimensional structure for more accurate estimation. Moreover, in addition to the additive noise, the multiplicative errors within the sensing tensor can be also well-corrected, leading to a more robust estimation. Experimental results on both simulated data and real videos demonstrate the superiority of the proposed method.
Jiani Liu 0002, Ce Zhu, Xiaolin Huang, Yipeng Liu 0001
ICASSP4
2024 Efficient Black-Box Adversarial Attack on Deep Clustering Models
abstract
Despite the significant progress made by deep clustering models in high-dimensional data processing, they remain vulnerable to adversarial examples. However, research on adversarial attacks against deep clustering algorithms appears to be relatively underexplored. To fill this gap, we propose a query-efficient black-box attack on deep clustering models, which leverages the transferability between different deep clustering models. Initially, we train a generator using a substitute deep clustering model, reducing the number of queries to the target model. Subsequently, when targeting an unknown deep clustering model, we employ the target query information to update both the substitute deep clustering model and the generator. Experimental evaluations on four state-of-the-art deep clustering models across three datasets demonstrate the efficacy of our method in disrupting clustering performance. The results indicate that our approach surpasses the performance of existing methods.
Zhen Long, Xiaolin Huang, Ce Zhu, Yipeng Liu 0001
ICIP4
2024 Ambiguity Consistency and Uncertainty Minimization for Semi-Supervised Medical Image Segmentation
abstract
Co-training and pseudo-supervision are two common strategies in semi-supervised medical image segmentation. However, co-training may lead to a ’resonance’ problem, and the effectiveness of generating pseudo-labels by setting thresholds may greatly depend on manual efforts. To address these issues, we propose an innovative framework for Ambiguity Consistency and Uncertainty Minimization (ACUM) in semi-supervised medical image segmentation. Specifically, ACUM comprises two main components: (1) Ambiguity Consistency Constraint (ACC), which encourages model differentiation and applies dynamic pixel-level consistency constraints through ambiguous areas between sub-networks; (2) Pixel Uncertainty Minimization (PUM), which generates high-confidence pseudo-labels by selecting labels with relatively low uncertainty based on the uncertainty maps of sub-networks. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of our proposed ACUM approach over state-of-the-art techniques.
Xiaolin Huang, Yujiang Yao, Bingzhi Chen
ICME1
2024 Enhancing Semi-supervised Medical Image Segmentation with Asymmetric and Adversarial Cooperative Training
Xiaolin Huang, Binzhi Chen, Jingchun Lin, Ruihua Nie, Jun Liang 0002
ICONIP (8)1
2024 Noise Perturbation Based Graph Contrastive Learning via Flexible Filters for Node Classification
abstract
Graph neural networks (GNNs), as a powerful deep learning framework for modeling graph-structured data, have attracted lots of attention recently. Most of existing GNNs need a lot of labeled data. However, constructing generalizable and robust representation from unlabeled graph data remains a challenge for GNNs. Existing graph contrastive learning (GCL) methods either try to uniformly drop edges, or intend to remove unimportant nodes and edges, which heavily relies on the specific structure of the data. Another thing is that vanilla graph convolutional network only utilize low-pass filter (adjacency matrix), which ignores the middle and high frequency information of the graph structural data. To tackle existing challenges in the GCL methods, instead, we propose a noise perturbation based general GCL framework via flexible filters. Specifically, we first add various types of noise to the nodes and edges. Subsequently, we design flexible filters, which are the combination of low, middle and high-pass filters. Our investigation systematically examines the impact of noise and filters, with an initial theoretical analysis linking these elements to the triplet loss function, shedding light on their roles. Extensive experiments in node classification showcase that our proposed approach surpasses existing state-of-the-art baselines. Surprisingly, we find that moderate levels of noise effectively alleviate the over-smoothing problem encountered in GNNs, while the use of flexible filters notably enhances model performance.
Zhilong Xiong, Ranhui Yan, Xiaolin Huang
IJCNN4
2024 Kernel PCA for Out-of-Distribution Detection
abstract
Out-of-Distribution (OoD) detection is vital for the reliability of Deep Neural Networks (DNNs). Existing works have shown the insufficiency of Principal Component Analysis (PCA) straightforwardly applied on the features of DNNs in detecting OoD data from In-Distribution (InD) data. The failure of PCA suggests that the network features residing in OoD and InD are not well separated by simply proceeding in a linear subspace, which instead can be resolved through proper non-linear mappings. In this work, we leverage the framework of Kernel PCA (KPCA) for OoD detection, and seek suitable non-linear kernels that advocate the separability between InD and OoD data in the subspace spanned by the principal components. Besides, explicit feature mappings induced from the devoted task-specific kernels are adopted so that the KPCA reconstruction error for new test samples can be efficiently obtained with large-scale data. Extensive theoretical and empirical results on multiple OoD data sets and network structures verify the superiority of our KPCA detector in efficiency and efficacy with state-of-the-art detection performance.
Kun Fang 0004, Qinghua Tao, Kexin Lv, Mingzhen He, Xiaolin Huang, Jie Yang 0002
NeurIPS5
2024 Unified Gradient-Based Machine Unlearning with Remain Geometry Enhancement
abstract
Machine unlearning (MU) has emerged to enhance the privacy and trustworthiness of deep neural networks. Approximate MU is a practical method for large-scale models. Our investigation into approximate MU starts with identifying the steepest descent direction, minimizing the output Kullback-Leibler divergence to exact MU inside a parameters' neighborhood. This probed direction decomposes into three components: weighted forgetting gradient ascent, fine-tuning retaining gradient descent, and a weight saliency matrix. Such decomposition derived from Euclidean metric encompasses most existing gradient-based MU methods. Nevertheless, adhering to Euclidean space may result in sub-optimal iterative trajectories due to the overlooked geometric structure of the output probability space. We suggest embedding the unlearning update into a manifold rendered by the remaining geometry, incorporating second-order Hessian from the remaining data. It helps prevent effective unlearning from interfering with the retained performance. However, computing the second-order Hessian for large-scale models is intractable. To efficiently leverage the benefits of Hessian modulation, we propose a fast-slow parameter update strategy to implicitly approximate the up-to-date salient unlearning direction. Free from specific modal constraints, our approach is adaptable across computer vision unlearning tasks, including classification and generation. Extensive experiments validate our efficacy and efficiency. Notably, our method successfully performs class-forgetting on ImageNet using DiT and forgets a class on CIFAR-10 using DDPM in just 50 steps, compared to thousands of steps required by previous methods. Code is available at [Unified-Unlearning-w-Remain-Geometry](https://github.com/K1nght/Unified-Unlearning-w-Remain-Geometry).
Zhehao Huang, Xinwen Cheng, JingHao Zheng, Zhengbao He, Xiaolin Huang
NeurIPS7
2024 Revisiting Deep Ensemble for Out-of-Distribution Detection: A Loss Landscape Perspective
Kun Fang 0004, Qinghua Tao, Xiaolin Huang, Jie Yang 0002
Int. J. Comput. Vis.3
2024 Fast Global Image Smoothing via Quasi Weighted Least Squares
Wei Liu 0044, Hongxing Qin, Xiaolin Huang, Jie Yang 0002, Michael Kwok-Po Ng
Int. J. Comput. Vis.4
2024 Boosting certified robustness via an expectation-based similarity regularization
Kun Fang 0004, Xiaolin Huang, Jie Yang 0002
Image Vis. Comput.3
2024 Random fourier features for asymmetric kernels
Mingzhen He, Fanghui Liu 0001, Xiaolin Huang
Mach. Learn.4
2024 Sparse Generalized Canonical Correlation Analysis: Distributed Alternating Iteration-Based Approach
abstract
Sparse canonical correlation analysis (CCA) is a useful statistical tool to detect latent information with sparse structures. However, sparse CCA, where the sparsity could be considered as a Laplace prior on the canonical variates, works only for two data sets, that is, there are only two views or two distinct objects. To overcome this limitation, we propose a sparse generalized canonical correlation analysis (GCCA), which could detect the latent relations of multiview data with sparse structures. Specifically, we convert the GCCA into a linear system of equations and impose ℓ1 minimization penalty to pursue sparsity. This results in a nonconvex problem on the Stiefel manifold. Based on consensus optimization, a distributed alternating iteration approach is developed, and consistency is investigated elaborately under mild conditions. Experiments on several synthetic and real-world data sets demonstrate the effectiveness of the proposed algorithm.
Kexin Lv, Junyi Huo, Xiaolin Huang, Jie Yang 0002
Neural Comput.5
2024 Low-Dimensional Gradient Helps Out-of-Distribution Detection
abstract
Detecting out-of-distribution (OOD) samples is essential for ensuring the reliability of deep neural networks (DNNs) in real-world scenarios. While previous research has predominantly investigated the disparity between in-distribution (ID) and OOD data through forward information analysis, the discrepancy in parameter gradients during the backward process of DNNs has received insufficient attention. Existing studies on gradient disparities mainly focus on the utilization of gradient norms, neglecting the wealth of information embedded in gradient directions. To bridge this gap, in this paper, we conduct a comprehensive investigation into leveraging the entirety of gradient information for OOD detection. The primary challenge arises from the high dimensionality of gradients due to the large number of network parameters. To solve this problem, we propose performing linear dimension reduction on the gradient using a designated subspace that comprises principal components. This innovative technique enables us to obtain a low-dimensional representation of the gradient with minimal information loss. Subsequently, by integrating the reduced gradient with various existing detection score functions, our approach demonstrates superior performance across a wide range of detection tasks. For instance, on the ImageNet benchmark with ResNet50 model, our method achieves an average reduction of 11.15 % in the false positive rate at 95 % recall (FPR95) compared to the current state-of-the-art approach.
Yingwen Wu, Tao Li 0054, Xinwen Cheng, Jie Yang 0002, Xiaolin Huang
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Towards robust neural networks via orthogonal diversity
Kun Fang 0004, Qinghua Tao, Yingwen Wu, Tao Li 0054, Feipeng Cai, Xiaolin Huang, Jie Yang 0002
Pattern Recognit.7
2024 Consensus-based distributed algorithm for GEP
Kexin Lv, Xiaolin Huang, Jie Yang 0002
Signal Process.3
2024 Improving the Post-Training Neural Network Quantization by Prepositive Feature Quantization
abstract
Post-training neural network quantization (PTQ) is an effective model compression technology that has revolutionized the deployment of deep neural networks on various edge devices. It provides easy-to-use characteristics and allows for generating a quantized model based on a pre-trained counterpart without re-training. Typical PTQ approaches maintain output consistency through layer-wise calibration. However, these approaches still suffer from performance degradation primarily caused by feature quantization in ultra-low bitwidth conditions. To address this issue, we propose a prepositive feature quantization framework that decouples adjacent layers and calibrates the interaction between feature and parameter quantization perturbations. Additionally, we present a feature-loss-aware optimization strategy to solve the corresponding calibration problem. To validate the effectiveness of our method, we conducted extensive experiments on the ImageNet benchmark dataset. Our approach demonstrates a noticeable improvement in PTQ performance under the 2-bit condition.
Zuopeng Yang, Xiaolin Huang
IEEE Trans. Circuits Syst. Video Technol.3
2024 Toward Transferable Adversarial Attacks Against Autoencoder-Based Network Intrusion Detectors
abstract
Deploying machine learning (ML)-based network intrusion detection systems has become a mainstream solution to improve the security of network efficiently. However, recent research has shown that ML models are vulnerable to adversarial examples. It is a formidable challenge for attackers to obtain the structure and gradients of intrusion detectors, thus transferable adversarial attacks that can deceive the unknown models pose a greater threat in practical scenarios. In this work, our goal is to investigate the cross-model transferability of adversarial examples toward autoencoder (AE)-based network intrusion detectors. Unlike adversarial methods in the image domain focusing on the distance between benign input and adversarial example, adversarial algorithms in network field emphasize complying with network protocols and maintaining malicious payload. We first introduce the common adversarial attacks in the image domain into AE-based network intrusion detectors with constraints. The experimental results show that iterative attacks perform better than single-step attacks against different AE-based models. At the same time, we discover that the transferable adversarial attacks in image domain are not very effective in facilitating the transferability of adversarial examples in this scenario because of fewer changeable features. To address this issue, from the perspective of the substitute model, we propose linear autoencoder (LAE) which is simply removed the activation functions of AE model but shares the same main structure with the original model. Extensive experimental evaluation demonstrates that by employing LAE as the source model, the transferability of both gradient-based and optimization-based adversarial attack methods can be improved significantly.
Yihang Zhang 0008, Yingwen Wu, Xiaolin Huang
IEEE Trans. Ind. Informatics3
2024 Self-Similarity Prior Distillation for Unsupervised Remote Physiological Measurement
abstract
Remote photoplethysmography (rPPG) is a non-invasive technique that aims to capture subtle variations in facial pixels caused by changes in blood volume resulting from cardiac activities. Most existing unsupervised methods for rPPG tasks focus on the contrastive learning between samples while neglecting the inherent self-similarity prior in physiological signals. In this paper, we propose a Self-Similarity Prior Distillation (SSPD) framework for unsupervised rPPG estimation, which capitalizes on the intrinsic temporal self-similarity of cardiac activities. Specifically, we first introduce a physical-prior embedded augmentation technique to mitigate the effect of various types of noise. Then, we tailor a self-similarity-aware network to disentangle more reliable self-similar physiological features. Finally, we develop a hierarchical self-distillation paradigm for self-similarity-aware learning and rPPG signal decoupling. Comprehensive experiments demonstrate that the unsupervised SSPD framework achieves comparable or even superior performance compared to the state-of-the-art supervised methods. Meanwhile, SSPD has the lowest inference time and computation cost among end-to-end models.
Weiyu Sun, Hao Lu 0009, Ying Chen 0006, Xiaolin Huang, Ying-Cong Chen
IEEE Trans. Multim.6
2024 Global Search and Analysis for the Nonconvex Two-Level ℓ₁ Penalty
abstract
Imposing suitably designed nonconvex regularization is effective to enhance sparsity, but the corresponding global search algorithm has not been well established. In this article, we propose a global search algorithm for the nonconvex two-level$\ell _{1}$penalty based on its piecewise linear property and apply it to machine learning tasks. With the search capability, the optimization performance of the proposed algorithm could be improved, resulting in better sparsity and accuracy than most state-of-the-art global and local algorithms. Besides, we also provide an approximation analysis to demonstrate the effectiveness of our global search algorithm in sparse quantile regression.
Mingzhen He, Lei Shi 0010, Xiaolin Huang
IEEE Trans. Neural Networks Learn. Syst.4
2023 Better Loss Landscape Visualization for Deep Neural Networks with Trajectory Information
Ruiqi Ding, Xiaolin Huang
ACML3
2023 Measuring the Transferability of ℓ∞ Attacks by the ℓ2 Norm
abstract
Deep neural networks could be fooled by adversarial examples with trivial differences to original samples. To keep the difference imperceptible in human eyes, researchers bound the adversarial perturbations by the ℓ∞norm, which is now commonly served as the standard to align the strength of different attacks for a fair comparison. However, we propose that using the ℓ∞norm alone is not sufficient in measuring the attack strength, because even with a fixed ℓ∞distance, the ℓ2distance also greatly affects the attack transferability between models. Through the discovery, we reach more in-depth understandings towards the attack mechanism, i.e., several existing methods attack black-box models better partly because they craft perturbations with 70% to 130% larger ℓ2distances. Since larger perturbations naturally lead to better transferability, we thereby advocate that the strength of attacks should be simultaneously measured by both the ℓ∞and ℓ2norm. Our proposal is firmly supported by extensive experiments on ImageNet dataset from 7 attacks, 4 white-box models, and 9 black-box models.
Sizhe Chen, Qinghua Tao, Zhixing Ye, Xiaolin Huang
ICASSP4
2023 Self-Ensemble Protection: Training Checkpoints Are Good Data Protectors
Sizhe Chen, Geng Yuan, Xinwen Cheng, Yifan Gong 0004, Minghai Qin, Yanzhi Wang 0001, Xiaolin Huang
ICLR7
2023 Trainable Weight Averaging: Efficient Training by Optimizing Historical Solutions
Tao Li 0054, Zhehao Huang, Qinghua Tao, Yingwen Wu, Xiaolin Huang
ICLR5
2023 One-Pixel Shortcut: On the Learning Preference of Deep Neural Networks
Shutong Wu, Sizhe Chen, Cihang Xie, Xiaolin Huang
ICLR4
2023 Consensus-Based Distributed Kernel One-class Support Vector Machine for Anomaly Detection
abstract
One-class support vector machine (OCSVM) is one of the most widely used methods for learning from imbalanced data and has been successfully applied to numerous tasks such as anomaly detection. However, the study on decentralized OCSVM is currently limited to linear cases. The main challenge is how to communicate the non-parametric and local-data-dependent decision functions between neighboring nodes. To tackle it, this paper proposes a projection consensus constraint to formulate a decentralized OCSVM, where local solutions are assumed to be the projection of the global optimum on local reproducing kernel Hilbert spaces. A fast non-parametric solving algorithm is then designed based on alternating direction method of multipliers. Experiments on real-world anomaly datasets indicate that our method outperforms the existing distributed OCSVM methods while reducing the communication cost.
Tianyao Wang, Ruikai Yang, Zhixing Ye, Xiaolin Huang
IJCNN5
2023 FG-UAP: Feature-Gathering Universal Adversarial Perturbation
abstract
Deep Neural Networks (DNNs) are susceptible to elaborately designed perturbations, whether such perturbations are dependent or independent of images. The latter one, called Universal Adversarial Perturbation (UAP), is very attractive for model robustness analysis, since its independence of input reveals the intrinsic characteristics of the model. Relatively, another interesting observation is Neural Collapse (NC), which means the feature variability may collapse during the terminal phase of training. Motivated by this, we propose to generate UAP by attacking the layer where NC phenomenon happens. Because of NC, the proposed attack could gather all the natural images’ features to its surrounding, which is hence called Feature-Gathering UAP (FG-UAP). We evaluate the effectiveness our proposed algorithm on abundant experiments, including untargeted and targeted universal attacks, attacks under limited dataset, and transfer-based black-box attacks among different architectures including Vision Transformers, which are believed to be more robust. After that, we empirically verify the effectiveness of NC's conclusion on UAP by attacking on only 10% of the dataset while keeping comparable performance. Finally, we investigate FG-UAP in the view of NC by analyzing the labels and extracted features of adversarial examples, finding that collapse phenomenon becomes stronger after the model is corrupted. Codes for the project are available at https://2ithub.com/yzx1213/FG-UAP.
Zhixing Ye, Xinwen Cheng, Xiaolin Huang
IJCNN3
2023 Resolve Domain Conflicts for Generalizable Remote Physiological Measurement
abstract
Remote photoplethysmography (rPPG) technology has become increasingly popular due to its non-invasive monitoring of various physiological indicators, making it widely applicable in multimedia interaction, healthcare, and emotion analysis. Existing rPPG methods utilize multiple datasets for training to enhance the generalizability of models. However, they often overlook the underlying conflict issues in the rPPG field, such as (1) label conflict resulting from different phase delays between physiological signal labels and face videos at the instance level, and (2) attribute conflict stemming from distribution shifts caused by head movements, illumination changes, skin types, etc. To address this, we introduce the DOmain-HArmonious framework (DOHA). Specifically, we first propose a harmonious phase strategy to eliminate uncertain phase delays and preserve the temporal variation of physiological signals. Next, we design a harmonious hyperplane optimization that reduces irrelevant attribute shifts and encourages the model's optimization towards a global solution that fits more valid scenarios. Our experiments demonstrate that DOHA significantly improves the performance of existing methods under multiple protocols.
Weiyu Sun, Hao Lu 0009, Ying Chen 0006, Xiaolin Huang, Ying-Cong Chen
ACM Multimedia6
2023 Multi-Frame Self-Supervised Depth Estimation with Multi-Scale Feature Fusion in Dynamic Scenes
abstract
Monocular depth estimation is a fundamental task in computer vision and multimedia. The self-supervised learning pipeline makes it possible to train the monocular depth network with no need of depth labels. In this paper, a multi-frame depth model with multi-scale feature fusion is proposed for strengthening texture features and spatial-temporal features, which improves the robustness of depth estimation between frames with large camera ego-motion. A novel dynamic object detecting method with geometry explainability is proposed. The detected dynamic objects are excluded during training, which guarantees the static environment assumption and relieves the accuracy degradation problem of the multi-frame depth estimation. Robust knowledge distillation with a consistent teacher network and reliability guarantee is proposed, which improves the multi-frame depth estimation without an increase in computation complexity during the test. The experiments show that our proposed methods achieve great performance improvement on the multi-frame depth estimation.
Jiquan Zhong, Xiaolin Huang, Xiao Yu 0002
ACM Multimedia2
2023 Diffusion Representation for Asymmetric Kernels via Magnetic Transform
abstract
As a nonlinear dimension reduction technique, the diffusion map (DM) has been widely used. In DM, kernels play an important role for capturing the nonlinear relationship of data. However, only symmetric kernels can be used now, which prevents the use of DM in directed graphs, trophic networks, and other real-world scenarios where the intrinsic and extrinsic geometries in data are asymmetric. A promising technique is the magnetic transform which converts an asymmetric matrix to a Hermitian one. However, we are facing essential problems, including how diffusion distance could be preserved and how divergence could be avoided during diffusion process. Via theoretical proof, we successfully establish a diffusion representation framework with the magnetic transform, named MagDM. The effectiveness and robustness for dealing data endowed with asymmetric proximity are demonstrated on three synthetic datasets and two trophic networks.
Mingzhen He, Ruikai Yang, Xiaolin Huang
NeurIPS4
2023 DeCAB: Debiased Semi-supervised Learning for Imbalanced Open-Set Data
Xiaolin Huang, Mengke Li 0001, Yang Lu 0009, Hanzi Wang
PRCV (9)1
2023 Improving adversarial robustness through a curriculum-guided reliable distillation
Kun Fang 0004, Xiaolin Huang, Jie Yang 0002
Comput. Secur.3
2023 An end-to-end lower limb activity recognition framework based on sEMG data augmentation and enhanced CapsNet
Changhe Zhang, Yangan Li, Zidong Yu, Xiaolin Huang
Expert Syst. Appl.4
2023 ADSP: An adaptive sample pooling strategy for diagnostic testing
Xuekui Zhang, Xiaolin Huang
J. Biomed. Informatics2
2023 Weighted neural tangent kernel: a generalized and improved network-induced kernel
Shutong Wu, Wenxing Zhou, Xiaolin Huang
Mach. Learn.4
2023 Learning With Asymmetric Kernels: Least Squares and Feature Interpretation
abstract
Asymmetric kernels naturally exist in real life, e.g., for conditional probability and directed graphs. However, most of the existing kernel-based learning methods require kernels to be symmetric, which prevents the use of asymmetric kernels. This paper addresses the asymmetric kernel-based learning in the framework of the least squares support vector machine named AsK-LS, resulting in the first classification method that can utilize asymmetric kernels directly. We will show that AsK-LS can learn with asymmetric features, namely source and target features, while the kernel trick remains applicable, i.e., the source and target features exist but are not necessarily known. Besides, the computational burden of AsK-LS is as cheap as dealing with symmetric kernels. Experimental results on various tasks, including Corel, PASCAL VOC, Satellite, directed graphs, and UCI database, all show that in the case asymmetric information is crucial, the proposed AsK-LS can learn with asymmetric kernels and performs much better than the existing kernel methods that rely on symmetrization to accommodate asymmetric kernels.
Mingzhen He, Lei Shi 0010, Xiaolin Huang, Johan A. K. Suykens
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Low Dimensional Trajectory Hypothesis is True: DNNs Can Be Trained in Tiny Subspaces
abstract
Deep neural networks (DNNs) usually contain massive parameters, but there is redundancy such that it is guessed that they could be trained in low-dimensional subspaces. In this paper, we propose a Dynamic Linear Dimensionality Reduction (DLDR) based on the low-dimensional properties of the training trajectory. The reduction method is efficient, supported by comprehensive experiments: optimizing DNNs in 40-dimensional spaces can achieve comparable performance as regular training over thousands or even millions of parameters. Since there are only a few variables to optimize, we develop an efficient quasi-Newton-based algorithm, obtain robustness to label noise, and improve the performance of well-trained models, which are three follow-up experiments that can show the advantages of finding such low-dimensional subspaces. The code is released (Pytorch: https://github.com/nblt/DLDR and Mindspore: https://gitee.com/mindspore/docs/tree/r1.6/docs/sample_code/dimension_reduce_training).
Tao Li 0054, Zhehao Huang, Qinghua Tao, Yipeng Liu 0001, Xiaolin Huang
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 End-to-end kernel learning via generative random Fourier features
Kun Fang 0004, Fanghui Liu 0001, Xiaolin Huang, Jie Yang 0002
Pattern Recognit.3
2023 Improving the adversarial robustness of quantized neural networks via exploiting the feature diversity
abstract
Quantized neural networks (QNNs) have become one of the most prevalent approaches in deep learning model compression due to their computational and storage efficiency. However, there is a lack of research specialized in the adversarial robustness of QNNs, which is important for applications in security-critical domains. Existing defenses focus on conventional full-precision networks, which can result in behavioral disparities and degrade the expected performance when directly transferred to QNNs. A novel defensive strategy promotes feature diversity through an orthogonal constraint, which can synergize well with quantization. Inspired by this intuition, we propose an orthogonal regularization with quantization to improve the adversarial robustness of QNNs in this paper. Moreover, we observe that quantization serves as an implicit regularization and is able to alleviate orthogonal degeneration. The proposed orthogonal regularization with quantization is validated on several typical network architectures and benchmark datasets. The results demonstrate that the proposed method can notably enhance adversarial robustness against both white-box and black-box attacks.
Kun Fang 0004, Jie Yang 0002, Xiaolin Huang
Pattern Recognit. Lett.4
2023 Learning non-parametric kernel via matrix decomposition for logistic regression
Mingzhen He, Xiaolin Huang
Pattern Recognit. Lett.4
2023 Unifying Gradients to Improve Real-World Robustness for Deep Networks
abstract
The wide application of deep neural networks (DNNs) demands an increasing amount of attention to their real-world robustness, i.e., whether a DNN resists black-box adversarial attacks, among which score-based query attacks (SQAs) are the most threatening since they can effectively hurt a victim network with only access to model outputs. Defending against SQAs requires a slight but artful variation of outputs due to the service purpose for users, who share the same output information with SQAs. In this article, we propose a real-world defense by Unifying Gradients (UniG) of different data so that SQAs could only probe a much weaker attack direction that is similar for different samples. Since such universal attack perturbations have been validated as less aggressive than the input-specific perturbations, UniG protects real-world DNNs by indicating to attackers a twisted and less informative attack direction. We implement UniG efficiently by a Hadamard product module, which is plug-and-play. According to extensive experiments on 5 SQAs, 2 adaptive attacks and 7 defense baselines, UniG significantly improves real-world robustness without hurting clean accuracy on CIFAR10 and ImageNet. For instance, UniG maintains a model of 77.80% accuracy under a 2500-query Square attack while the state-of-the-art adversarially trained model only has 67.34% on CIFAR10. Simultaneously, UniG outperforms all compared baselines in terms of clean accuracy and achieves the smallest modification of the model output. The code is released at https://github.com/snowien/UniG-pytorch .
Yingwen Wu, Sizhe Chen, Kun Fang 0004, Xiaolin Huang
ACM Trans. Intell. Syst. Technol.4
2022 Handling Sparse Longitudinal Data with Irregular Missing Data - Analysis of Fecal Coliform Bacteria Data
abstract
Fecal coliform bacteria are commonly used as an indicator to reflect the fecal contamination level in the water. To manage a healthy and thriving aquaculture industry, the Canadian Shellfish Sanitation Program (CSSP) was established in 1948. As a part of the CSSP mandate, fecal coliform bacteria levels in shellfish growing habitat have been monitored at nearly 15, 000 shellfish harvesting sites across the six coastal provinces of Canada over 40 years (1980–2019). The irregular sparseness in the measurement data presented a critical challenge for reliable analysis of fecal contamination patterns along Canada's coastline. This paper illustrates a preprocessing approach to handle the irregular sparseness in the measurement data of fecal coliform bacteria levels and demonstrates the effectiveness of data filtering, pooling, binning, partitioning, and the application of functional principal component analysis. We managed to transform the irregularly sparse measurements at a site into surrogate variables without missing data that represent the differences in the site's contamination amplitude and seasonal variation from the average. The surrogate variables will be used for downstream analyses, such as associating contamination with climate change.
Shuai You, Xiaolin Huang, Youlian Pan, Xuekui Zhang
CIBCB2
2022 Subspace Adversarial Training
abstract
Single-step adversarial training (AT) has received wide attention as it proved to be both efficient and robust. However, a serious problem of catastrophic overfitting exists, i.e., the robust accuracy against projected gradient descent (PGD) attack suddenly drops to 0% during the training. In this paper, we approach this problem from a novel perspective of optimization and firstly reveal the close link between the fast-growing gradient of each sample and overfitting, which can also be applied to understand robust overfitting in multi-step AT. To control the growth of the gradient, we propose a new AT method, Subspace Adversarial Training (Sub-AT), which constrains AT in a carefully extracted subspace. It successfully resolves both kinds of overfitting and significantly boosts the robustness. In subspace, we also allow single-step AT with larger steps and larger radius, further improving the robustness performance. As a result, we achieve state-of-the-art single-step AT performance. Without any regularization term, our single-step AT can reach over 51 % robust accuracy against strong PGD-50 attack of radius 8/255 on CIFAR-10, reaching a competitive performance against standard multi-step PGD-10 AT with huge computational advantages. The code is released at https://github.com/nblt/Sub-AT.
Tao Li 0054, Yingwen Wu, Sizhe Chen, Kun Fang 0004, Xiaolin Huang
CVPR5
2022 PCR-CG: Point Cloud Registration via Deep Explicit Color and Geometry
Yu Zhang 0280, Junle Yu, Xiaolin Huang, Wenhui Zhou 0001, Ji Hou
ECCV (10)3
2022 Mutual Diverse-Label Adversarial Training
Sizhe Chen, Xiaolin Huang
ICONIP (1)3
2022 Adversarial Attack on Attackers: Post-Process to Mitigate Black-Box Score-Based Query Attacks
abstract
The score-based query attacks (SQAs) pose practical threats to deep neural networks by crafting adversarial perturbations within dozens of queries, only using the model's output scores. Nonetheless, we note that if the loss trend of the outputs is slightly perturbed, SQAs could be easily misled and thereby become much less effective. Following this idea, we propose a novel defense, namely Adversarial Attack on Attackers (AAA), to confound SQAs towards incorrect attack directions by slightly modifying the output logits. In this way, (1) SQAs are prevented regardless of the model's worst-case robustness; (2) the original model predictions are hardly changed, i.e., no degradation on clean accuracy; (3) the calibration of confidence scores can be improved simultaneously. Extensive experiments are provided to verify the above advantages. For example, by setting $\ell_\infty=8/255$ on CIFAR-10, our proposed AAA helps WideResNet-28 secure 80.59% accuracy under Square attack (2500 queries), while the best prior defense (i.e., adversarial training) only attains 67.44%. Since AAA attacks SQA's general greedy strategy, such advantages of AAA over 8 defenses can be consistently observed on 8 CIFAR-10/ImageNet models under 6 SQAs, using different attack targets, bounds, norms, losses, and strategies. Moreover, AAA calibrates better without hurting the accuracy. Our code is available at https://github.com/Sizhe-Chen/AAA.
Sizhe Chen, Zhehao Huang, Qinghua Tao, Yingwen Wu, Cihang Xie, Xiaolin Huang
NeurIPS6
2022 Universal Adversarial Attack on Attention and the Resulting Dataset DAmageNet
abstract
Adversarial attacks on deep neural networks (DNNs) have been found for several years. However, the existing adversarial attacks have high success rates only when the information of the victim DNN is well-known or could be estimated by the structure similarity or massive queries. In this paper, we propose to Attack on Attention (AoA), a semantic property commonly shared by DNNs. AoA enjoys a significant increase in transferability when the traditional cross entropy loss is replaced with the attention loss. Since AoA alters the loss function only, it could be easily combined with other transferability-enhancement techniques and then achieve SOTA performance. We apply AoA to generate 50000 adversarial samples from ImageNet validation set to defeat many neural networks, and thus name the dataset as DAmageNet. 13 well-trained DNNs are tested on DAmageNet, and all of them have an error rate over 85 percent. Even with defenses or adversarial training, most models still maintain an error rate over 70 percent on DAmageNet. DAmageNet is the first universal adversarial dataset. It could be downloaded freely and serve as a benchmark for robustness testing and adversarial training.
Sizhe Chen, Zhengbao He, Chengjin Sun, Jie Yang 0002, Xiaolin Huang
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Random Features for Kernel Approximation: A Survey on Algorithms, Theory, and Beyond
abstract
The class of random features is one of the most popular techniques to speed up kernel methods in large-scale problems. Related works have been recognized by the NeurIPS Test-of-Time award in 2017 and the ICML Best Paper Finalist in 2019. The body of work on random features has grown rapidly, and hence it is desirable to have a comprehensive overview on this topic explaining the connections among various algorithms and theoretical results. In this survey, we systematically review the work on random features from the past ten years. First, the motivations, characteristics and contributions of representative random features based algorithms are summarized according to their sampling schemes, learning procedures, variance reduction properties and how they exploit training data. Second, we review theoretical results that center around the following key question: how many random features are needed to ensure a high approximation quality or no loss in the empirical/expected risks of the learned estimator. Third, we provide a comprehensive evaluation of popular random features based algorithms on several large-scale benchmark datasets and discuss their approximation quality and prediction performance for classification. Last, we discuss the relationship between random features and modern over-parameterized deep neural networks (DNNs), including the use of high dimensional random features in the analysis of DNNs as well as the gaps between current theoretical and empirical results. This survey may serve as a gentle introduction to this topic, and as a users' guide for practitioners interested in applying the representative algorithms and understanding theoretical results under various technical assumptions. We hope that this survey will facilitate discussion on the open problems in this topic, and more importantly, shed light on future research directions. Due to the page limit, we suggest the readers refer to the full version of this survey https://arxiv.org/abs/2004.11154.
Fanghui Liu 0001, Xiaolin Huang, Yudong Chen 0001, Johan A. K. Suykens
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Towards a Unified Quadrature Framework for Large-Scale Kernel Machines
abstract
In this paper, we develop a quadrature framework for large-scale kernel machines via a numerical integration representation. Considering that the integration domain and measure of typical kernels, e.g., Gaussian kernels, arc-cosine kernels, are fully symmetric, we leverage a numerical integration technique, deterministic fully symmetric interpolatory rules, to efficiently compute quadrature nodes and associated weights for kernel approximation. Thanks to the full symmetric property, the applied interpolatory rules are able to reduce the number of needed nodes while retaining a high approximation accuracy. Further, we randomize the above deterministic rules by the classical Monte-Carlo sampling and control variates techniques with two merits: 1) The proposed stochastic rules make the dimension of the feature mapping flexibly varying, such that we can control the discrepancy between the original and approximate kernels by tuning the dimnension. 2) Our stochastic rules have nice statistical properties of unbiasedness and variance reduction. In addition, we elucidate the relationship between our deterministic/stochastic interpolatory rules and current typical quadrature based rules for kernel approximation, thereby unifying these methods under our framework. Experimental results on several benchmark datasets show that our methods compare favorably with other representative kernel approximation based methods.
Fanghui Liu 0001, Xiaolin Huang, Yudong Chen 0001, Johan A. K. Suykens
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 A Generalized Framework for Edge-Preserving and Structure-Preserving Image Smoothing
abstract
Image smoothing is a fundamental procedure in applications of both computer vision and graphics. The required smoothing properties can be different or even contradictive among different tasks. Nevertheless, the inherent smoothing nature of one smoothing operator is usually fixed and thus cannot meet the various requirements of different applications. In this paper, we first introduce the truncated Huber penalty function which shows strong flexibility under different parameter settings. A generalized framework is then proposed with the introduced truncated Huber penalty function. When combined with its strong flexibility, our framework is able to achieve diverse smoothing natures where contradictive smoothing behaviors can even be achieved. It can also yield the smoothing behavior that can seldom be achieved by previous methods, and superior performance is thus achieved in challenging cases. These together enable our framework capable of a range of applications and able to outperform the state-of-the-art approaches in several tasks. In addition, an efficient numerical solution is provided and its convergence is theoretically guaranteed even the optimization framework is non-convex and non-smooth. A simple yet effective approach is further proposed to reduce the computational cost of our method while maintaining its performance. The effectiveness and superior performance of our approach are validated through comprehensive experiments in a range of applications.
Wei Liu 0044, Yinjie Lei, Xiaolin Huang, Jie Yang 0002, Michael Kwok-Po Ng
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Relevance attack on detectors
Sizhe Chen, Xiaolin Huang, Kun Zhang 0001
Pattern Recognit.3
2022 DFR-ST: Discriminative feature representation with spatio-temporal cues for vehicle re-identification
Jingzheng Tu, Cailian Chen, Xiaolin Huang, Jianping He 0001, Xin-Ping Guan
Pattern Recognit.3
2022 Multivariate uncertain risk aversion with application to accounts receivables pricing
Ke Wang 0004, Xiaolin Huang, Hongwei Wang 0009, Mingxuan Zhao, Jian Zhou 0003
Soft Comput.2
2022 Unsupervised Image Restoration With Quality-Task-Perception Loss
abstract
Image restoration includes various kinds of tasks, such as image denoising, image deraining and low-light image enhancement, etc. Due to the domain shift problem of current supervised methods, researchers tend to adopt unsupervised image restoration methods. However, fake color or blur image, insufficient restoration and missing semantic information are three common problems when utilizing these methods. In this paper, we propose a new hybrid loss named Quality-Task-Perception (QTP) to deal with these three problems simultaneously. Specifically, this hybrid loss includes three components: quality, task and perception. The quality part overcomes the fake color or blur image problem by enforcing image quality scores of the restored images and those of the unpaired clean images to be similar. For the task part, we tackle the insufficient restoration problem by proposing to apply a task probability network to convert the unsupervised image restoration into a supervised classification problem, and this task probability network is learned from our proposed pipeline. The perception part handles the missing semantic information by restricting the multi-scale phase consistency between the degraded image and its restored version. Comprehensive experiments on both supervised and unsupervised datasets in three image restoration tasks demonstrate the superiority of our proposed approach.
Wei Xu 0050, Haoming Guo, Xiaolin Huang, Wei Liu 0044
IEEE Trans. Circuits Syst. Video Technol.4
2022 Hierarchical Superpixel Segmentation by Parallel CRTrees Labeling
abstract
This paper proposes a hierarchical superpixel segmentation by representing an image as a hierarchy of 1-nearest neighbor (1-NN) graphs with pixels/superpixels denoting the graph vertices. The 1-NN graphs are built from the pixel/superpixel adjacent matrices to ensure connectivity. To determine the next-level superpixel hierarchy, inspired by FINCH clustering, the weakly connected components (WCCs) of the 1-NN graph are labeled as superpixels. We reveal that the WCCs of a 1-NN graph consist of a forest of cycle-root-trees (CRTrees). The forest-like structure inspires us to propose a two-stage parallel CRTrees labeling which first links the child vertices to the cycle-roots and then labels all the vertices by the cycle-roots. We also propose an inter-inner superpixel distance penalization and a Lab color lightness penalization base on the property that the distance of a CRTree decreases monotonically from the child to root vertices. Experiments show the parallel CRTrees labeling is several times faster than recent advanced sequential and parallel connected components labeling algorithms. The proposed hierarchical superpixel segmentation has comparable performance to the best performer ETPS (state-of-the-arts) on the BSDS500, NYUV2, and Fash datasets. At the same time, it can achieve 200FPS for 480P video streams.
Tingman Yan, Xiaolin Huang, Qunfei Zhao
IEEE Trans. Image Process.2
2022 Adaptive Temporal-Frequency Network for Time-Series Forecasting
abstract
A novel adaptive temporal-frequency network (ATFN), which is an end-to-end hybrid model incorporating deep learning networks and frequency patterns, is proposed for mid- and long-term time series forecasting. Within the framework of the ATFN, an augmented sequence to sequence model is used to learn the trend feature of complicated nonstationary time series, a frequency-domain block is used to capture dynamic and complicated periodic patterns of time series data, and a fully connected neural network is used to combine the trend and periodic features for producing a final forecast. An adaptive frequency mechanism consisting of phase adaption, frequency adaption, and amplitude adaption is designed for mapping the frequency spectrum of the current sliding window to that of the forecasting interval. The multilayer neural networks conduct a transformation similar to the inverse discrete Fourier transform for generating a periodic feature forecast. Synthetic data and real-world data with different periodic characteristics are used to evaluate the effectiveness of the proposed model. The experimental results indicate that the ATFN has promising performance and strong adaptability for long-term time-series forecasting.
Zhangjing Yang, Weiwu Yan, Xiaolin Huang, Lin Mei 0001
IEEE Trans. Knowl. Data Eng.3
2022 Toward Robust Histology-Prior Embedding for Endomicroscopy Image Classification
abstract
Representation learning is the critical task for medical image analysis in computer-aided diagnosis. However, it is challenging to learn discriminative features due to the limited size of the dataset and the lack of labels. In this paper, we propose a stochastic routing normalization and neighborhood embedding framework with application to breast tissue classification by learning discriminative features of probe-based confocal laser endomicroscopy. In order to align the low-level and mid-level of pCLE and histology domain, we firstly build the domain-specific normalization module with stochastic activation strategy considering both depth-wise and feature-wise criterion. For high-level features, the latent centers are learned from the histology domain as the template for feature matching. The proposed method is evaluated on a clinical database with 700 pCLE mosaics. The accuracy of image classification with limited training samples demonstrates that the proposed method can outperform previous works on domain alignment.
Yun Gu, Yunze Xu, Xiaolin Huang, Jie Yang 0002, Guang-Zhong Yang
IEEE Trans. Medical Imaging3
2021 Fast Learning in Reproducing Kernel Krein Spaces via Signed Measures
abstract
In this paper, we attempt to solve a long-lasting open question for non-positive definite (non-PD) kernels in machine learning community: can a given non-PD kernel be decomposed into the difference of two PD kernels (termed as positive decomposition)? We cast this question as a distribution view by introducing the signed measure, which transforms positive decomposition to measure decomposition: a series of non-PD kernels can be associated with the linear combination of specific finite Borel measures. In this manner, our distribution-based framework provides a sufficient and necessary condition to answer this open question. Specifically, this solution is also computationally implementable in practice to scale non-PD kernels in large sample cases, which allows us to devise the first random features algorithm to obtain an unbiased estimator. Experimental results on several benchmark datasets verify the effectiveness of our algorithm over the existing methods.
Fanghui Liu 0001, Xiaolin Huang, Yingyi Chen, Johan A. K. Suykens
AISTATS2
2021 Improving The Robustness Of Convolutional Neural Networks Via Sketch Attention
abstract
The convolutional neural networks (CNNs) is biased towards texture while human eyes relying heavily on the general structure. The inconformity leads to the vulnerability of CNNs. The convolutional results is determined by the local patterns and delicate adversarial perturbation would be amplified layer-wise. Meanwhile the image context and object structure, which can be represented by sketch, stay almost unchanged. Therefore we propose that the sketch information is weak but more robust. In order to transfer the robustness from sketch to image and improve the capture of global structure, a sketch attention guided CNNs (SAG-CNNs) pipeline is constructed. The experiments on CFAR-10 demonstrate that the defensive capability against black-box attacks of SAG-CNNs outperforms other counterparts evidently and achieve preferable trade-off between generalization and robustness.
Zuopeng Yang, Jie Yang 0002, Xiaolin Huang
ICIP4
2021 ADVMIX: Data Augmentation for Accurate Scene Text Spotting
abstract
Accurate scene text spotting models ask for effective data augmentation algorithms. This paper presents a novel data augmentation algorithm called AdvMix that integrates techniques of adversarial attack and mixup. First, it utilizes the PGD method to synthesize adversarial samples. Second, it is the first work to implement feature-wise mixup between original data and the associated adversarial sample to enhance data augmentation. AdvMix has been evaluated over ICDAR2013 and ICDAR2015 datasets, showing its superior performance in improving accuracy of scene text spotting models.
Yizhang Huang, Kun Fang 0004, Xiaolin Huang, Jie Yang 0002
ICIP3
2021 Residual Enhanced Multi-Hypergraph Neural Network
abstract
Hypergraphs are a generalized data structure of graphs to model higher-order correlations among entities, which have been successfully adopted into various researching fields. Meanwhile, HyperGraph Neural Network (HGNN) is currently the de-facto method for hypergraph representation learning. However, HGNN aims at single hypergraph learning and uses a pre-concatenation approach when confronting multi-modal datasets, which leads to sub-optimal exploitation of the inter-correlations of multi-modal hypergraphs. HGNN also suffers the over-smoothing issue, that is, its performance drops significantly when layers are stacked up. To resolve these issues, we propose the Residual enhanced Multi-Hypergraph Neural Network, which can not only fuse multi-modal information from each hypergraph effectively, but also circumvent the over-smoothing issue associated with HGNN. We conduct experiments on two 3D benchmarks, the NTU and the ModelNet40 datasets, and compare against multiple state-of-the-art methods. Experimental results demonstrate that both the residual hypergraph convolutions and the multi-fusion architecture can improve the performance of the base model and the combined model achieves a new state-of-the-art. Code is available at https://github.com/OneForward/ResMHGNN.
Xiaolin Huang, Jie Yang 0002
ICIP2
2021 Towards Unbiased Random Features with Lower Variance For Stationary Indefinite Kernels
abstract
Random Fourier Features (RFF) demonstrate well-appreciated performance in kernel approximation for large-scale situations but restrict kernels to be stationary and positive definite. And for non-stationary kernels, the corresponding RFF could be converted to that for stationary indefinite kernels when the inputs are restricted to the unit sphere. Numerous methods provide accessible ways to approximate stationary but indefinite kernels. However, they are either biased or possess large variance. In this article, we propose the generalized orthogonal random features, an unbiased estimation with lower variance. Experimental results on various datasets and kernels verify that our algorithm achieves lower variance and approximation error compared with the existing kernel approximation methods. With better approximation to the originally selected kernels, improved classification accuracy and regression ability is obtained with our approximation algorithm in the framework of support vector machine and regression.
Kun Fang 0004, Jie Yang 0002, Xiaolin Huang
IJCNN4
2021 Generalization Properties of hyper-RKHS and its Applications
abstract
This paper generalizes regularized regression problems in a hyper-reproducing kernel Hilbert space (hyper-RKHS), illustrates its utility for kernel learning and out-of-sample extensions, and proves asymptotic convergence results for the introduced regression models in an approximation theory view. Algorithmically, we consider two regularized regression models with bivariate forms in this space, including kernel ridge regression (KRR) and support vector regression (SVR) endowed with hyper-RKHS, and further combine divide-and-conquer with Nyström approximation for scalability in large sample cases. This framework is general: the underlying kernel is learned from a broad class, and can be positive definite or not, which adapts to various requirements in kernel learning. Theoretically, we study the convergence behavior of regularized regression algorithms in hyper-RKHS and derive the learning rates, which goes beyond the classical analysis on RKHS due to the non-trivial independence of pairwise samples and the characterisation of hyper-RKHS. Experimentally, results on several benchmarks suggest that the employed framework is able to learn a general kernel function form an arbitrary similarity matrix, and thus achieves a satisfactory performance on classification tasks.
Fanghui Liu 0001, Lei Shi 0010, Xiaolin Huang, Jie Yang 0002, Johan A. K. Suykens
J. Mach. Learn. Res.3
2021 Analysis of regularized least-squares in reproducing kernel Kreĭn spaces
Fanghui Liu 0001, Lei Shi 0010, Xiaolin Huang, Jie Yang 0002, Johan A. K. Suykens
Mach. Learn.3
2021 Adversarial Attack Type I: Cheat Classifiers by Significant Changes
abstract
Despite the great success of deep neural networks, the adversarial attack can cheat some well-trained classifiers by small permutations. In this paper, we propose another type of adversarial attack that can cheat classifiers by significant changes. For example, we can significantly change a face but well-trained neural networks still recognize the adversarial and the original example as the same person. Statistically, the existing adversarial attack increases Type II error and the proposed one aims at Type I error, which are hence named as Type II and Type I adversarial attack, respectively. The two types of attack are equally important but are essentially different, which are intuitively explained and numerically evaluated. To implement the proposed attack, a supervised variation autoencoder is designed and then the classifier is attacked by updating the latent variables using gradient information. Besides, with pre-trained generative models, Type I attack on latent spaces is investigated as well. Experimental results show that our method is practical and effective to generate Type I adversarial examples on large-scale image datasets. Most of these generated examples can pass detectors designed for defending Type II attack and the strengthening strategy is only efficient with a specific type attack, both implying that the underlying reasons for Type I and Type II attack are different.
Sanli Tang, Xiaolin Huang, Chengjin Sun, Jie Yang 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Mixed-precision quantized neural networks with progressively decreasing bitwidth
Jie Yang 0002, Xiaolin Huang
Pattern Recognit.4
2021 One-Shot Distributed Algorithm for PCA With RBF Kernels
abstract
This letter proposes a one-shot algorithm for feature-distributed kernel PCA. Our algorithm is inspired by the dual relationship between sample-distributed and feature-distributed scenarios. This interesting relationship makes it possible to establish distributed kernel PCA for feature-distributed cases from ideas of distributed PCA in the sample-distributed scenario. In the theoretical part, we analyze the approximation error for both linear and RBF kernels. The result suggests that when eigenvalues decay fast, the proposed algorithm gives high-quality results with low communication cost. This result is also verified by numerical experiments, showing the effectiveness of our algorithm in practice.
Kexin Lv, Jie Yang 0002, Xiaolin Huang
IEEE Signal Process. Lett.4
2021 Learning Tubule-Sensitive CNNs for Pulmonary Airway and Artery-Vein Segmentation in CT
abstract
Training convolutional neural networks (CNNs) for segmentation of pulmonary airway, artery, and vein is challenging due to sparse supervisory signals caused by the severe class imbalance between tubular targets and background. We present a CNNs-based method for accurate airway and artery-vein segmentation in non-contrast computed tomography. It enjoys superior sensitivity to tenuous peripheral bronchioles, arterioles, and venules. The method first uses a feature recalibration module to make the best use of features learned from the neural networks. Spatial information of features is properly integrated to retain relative priority of activated regions, which benefits the subsequent channel-wise recalibration. Then, attention distillation module is introduced to reinforce representation learning of tubular objects. Fine-grained details in high-resolution attention maps are passing down from one layer to its previous layer recursively to enrich context. Anatomy prior of lung context map and distance transform map is designed and incorporated for better artery-vein differentiation capacity. Extensive experiments demonstrated considerable performance gains brought by these components. Compared with state-of-the-art methods, our method extracted much more branches while maintaining competitive overall segmentation performance. Codes and models are available at http://www.pami.sjtu.edu.cn/News/56.
Yulei Qin, Hao Zheng 0008, Yun Gu, Xiaolin Huang, Jie Yang 0002, Lihui Wang 0002, Yue Min Zhu, Guang-Zhong Yang
IEEE Trans. Medical Imaging4
2020 Random Fourier Features via Fast Surrogate Leverage Weighted Sampling
abstract
In this paper, we propose a fast surrogate leverage weighted sampling strategy to generate refined random Fourier features for kernel approximation. Compared to the current state-of-the-art method that uses the leverage weighted scheme (Li et al. 2019), our new strategy is simpler and more effective. It uses kernel alignment to guide the sampling process and it can avoid the matrix inversion operator when we compute the leverage function. Given n observations and s random features, our strategy can reduce the time complexity for sampling from O(ns2+s3) to O(ns2), while achieving comparable (or even slightly better) prediction performance when applied to kernel ridge regression (KRR). In addition, we provide theoretical guarantees on the generalization performance of our approach, and in particular characterize the number of random features required to achieve statistical guarantees in KRR. Experiments on several benchmark datasets demonstrate that our algorithm achieves comparable prediction performance and takes less time cost when compared to (Li et al. 2019).
Fanghui Liu 0001, Xiaolin Huang, Yudong Chen 0001, Jie Yang 0002, Johan A. K. Suykens
AAAI2
2020 A Generalized Framework for Edge-Preserving and Structure-Preserving Image Smoothing
abstract
Image smoothing is a fundamental procedure in applications of both computer vision and graphics. The required smoothing properties can be different or even contradictive among different tasks. Nevertheless, the inherent smoothing nature of one smoothing operator is usually fixed and thus cannot meet the various requirements of different applications. In this paper, a non-convex non-smooth optimization framework is proposed to achieve diverse smoothing natures where even contradictive smoothing behaviors can be achieved. To this end, we first introduce the truncated Huber penalty function which has seldom been used in image smoothing. A robust framework is then proposed. When combined with the strong flexibility of the truncated Huber penalty function, our framework is capable of a range of applications and can outperform the state-of-the-art approaches in several tasks. In addition, an efficient numerical solution is provided and its convergence is theoretically guaranteed even the optimization framework is non-convex and non-smooth. The effectiveness and superior performance of our approach are validated through comprehensive experimental results in a range of applications.
Wei Liu 0044, Yinjie Lei, Xiaolin Huang, Jie Yang 0002, Ian D. Reid 0001
AAAI4
2020 A Real-time PCB Defect Detector Based on Supervised and Semi-supervised Learning
Sanli Tang, Siamak Mehrkanoon, Xiaolin Huang, Jie Yang 0002
ESANN4
2020 Learning from partially labeled data
Siamak Mehrkanoon, Xiaolin Huang, Johan A. K. Suykens
ESANN2
2020 Type I Attack For Generative Models
abstract
Generative models are popular tools with a wide range of applications. Nevertheless, it is as vulnerable to adversarial samples as classifiers. The existing attack methods mainly focus on generating adversarial examples by adding imperceptible perturbations to input, which leads to wrong result. However, we focus on another aspect of attack, i.e., cheating models by significant changes. The former induces Type II error and the latter causes Type I error. In this paper, we propose Type I attack to generative models such as VAE and GAN. One example given in VAE is that we can change an original image significantly to a meaningless one but their reconstruction results are similar. To implement the Type I attack, we destroy the original one by increasing the distance in input space while keeping the output similar because different inputs may correspond to similar features for the property of deep neural network. Experimental results show that our attack method is effective to generate Type I adversarial examples for generative models on large-scale image datasets.
Chengjin Sun, Sizhe Chen, Xiaolin Huang
ICIP4
2020 Edge-Aware Graph Attention Network for Ratio of Edge-User Estimation in Mobile Networks
abstract
Estimating the Ratio of Edge-Users (REU) is an important issue in mobile networks, as it helps the subsequent adjustment of loads in different cells. However, existing approaches usually determine the REU manually, which are experience-dependent and labor-intensive, and thus the estimated REU might be imprecise. Considering the inherited graph structure of mobile networks, in this paper, we utilize a graph-based deep learning method for automatic REU estimation, where the practical cells are deemed as nodes and the load switchings among them constitute edges. Concretely, Graph Attention Network (GAT) is employed as the backbone of our method due to its impressive generalizability in dealing with networked data. Nevertheless, conventional GAT cannot make full use of the information in mobile networks, since it only incorporates node features to infer the pairwise importance and conduct graph convolutions, while the edge features that are actually critical in our problem are disregarded. To accommodate this issue, we propose an Edge-Aware Graph Attention Network (EAGAT), which is able to fuse the node features and edge features for REU estimation. Extensive experimental results on two real-world mobile network datasets demonstrate the superiority of our EAGAT approach to several state-of-the-art methods.
Jiehui Deng, Sheng Wan, Enmei Tu, Xiaolin Huang, Jie Yang 0002, Chen Gong 0002
ICPR5
2020 Learning Bronchiole-Sensitive Airway Segmentation CNNs by Feature Recalibration and Attention Distillation
Yulei Qin, Hao Zheng 0008, Yun Gu, Xiaolin Huang, Jie Yang 0002, Lihui Wang 0002, Yue Min Zhu
MICCAI (1)4
2020 Learning Data-adaptive Non-parametric Kernels
abstract
In this paper, we propose a data-adaptive non-parametric kernel learning framework in margin based kernel methods. In model formulation, given an initial kernel matrix, a data-adaptive matrix with two constraints is imposed in an entry-wise scheme. Learning this data-adaptive matrix in a formulation-free strategy enlarges the margin between classes and thus improves the model flexibility. The introduced two constraints are imposed either exactly (on small data sets) or approximately (on large data sets) in our model, which provides a controllable trade-off between model flexibility and complexity with theoretical demonstration. In algorithm optimization, the objective function of our learning framework is proven to be gradient-Lipschitz continuous. Thereby, kernel and classifier/regressor learning can be efficiently optimized in a unified framework via Nesterov's acceleration. For the scalability issue, we study a decomposition-based approach to our model in the large sample case. The effectiveness of this approximation is illustrated by both empirical studies and theoretical guarantees. Experimental results on various classification and regression benchmark data sets demonstrate that our non-parametric kernel learning framework achieves good performance when compared with other representative kernel learning based algorithms.
Fanghui Liu 0001, Xiaolin Huang, Chen Gong 0002, Jie Yang 0002, Li Li 0013
J. Mach. Learn. Res.2
2020 Embedding Bilateral Filter in Least Squares for Efficient Edge-Preserving Image Smoothing
abstract
Edge-preserving smoothing is a fundamental procedure for many computer vision and graphic applications. This can be achieved with either local methods or global methods. In most cases, global methods can yield superior performance over the local ones. However, local methods usually run much faster than the global ones. In this paper, we propose a new global method that embeds the bilateral filter (BLF) in the least squares (LS) model for efficient edge-preserving smoothing. The proposed method can show comparable performance with the state-of-the-art global method. Meanwhile, since the proposed method can take advantages of the efficiency of the BLF and the LS model, it runs much faster. In addition, we show the flexibility of our method which can be easily extended by replacing the BLF with its variants. They can be further modified to handle more applications. We validate the effectiveness and efficiency of the proposed method through comprehensive experiments in a range of applications.
Wei Liu 0044, Chunhua Shen, Xiaolin Huang, Jie Yang 0002
IEEE Trans. Circuits Syst. Video Technol.5
2020 Toward Making Unsupervised Graph Hashing Discriminative
abstract
Recently, hashing has attracted much attention in visual information retrieval due to its low storage cost and fast query speed. The goal of hashing is to map original high-dimensional data into a low-dimensional binary-code space where the similar data points are assigned similar hash codes and dissimilar points are far away from each other. Existing unsupervised hashing methods mainly focus on recovering the pairwise similarity of the original data in hash space, but do not take specific measures to make the generated binary codes to be discriminative. To address this problem, this paper proposes a novel unsupervised hashing method, named “Discriminative Unsupervised Graph Hashing” (DUGH), which takes both similarity and dissimilarity of original data into consideration to learn discriminative binary codes. In particular, a probabilistic model is utilized to learn the encoding of original data in low-dimensional space, which models the original neighbor structure through both positive and negative edges in the KNN graph and then maximizes the likelihood of observing these edges. To efficiently and accurately measure the neighbor structure for largescale datasets, we propose an effective KNN graph construction algorithm based on the random projection tree and neighbor exploring techniques. The experimental results on one synthetic dataset and four typical real-world image datasets demonstrate that the proposed method significantly outperforms the state-of-the-art unsupervised hashing methods.
Chao Ma 0005, Chen Gong 0002, Xiang Li 0041, Xiaolin Huang, Wei Liu 0005, Jie Yang 0002
IEEE Trans. Multim.4
2020 Online Robust Principal Component Analysis With Change Point Detection
abstract
Robust principal component analysis (PCA) is a key technique for dynamical high-dimensional data analysis, including background subtraction for surveillance video. Typically, robust PCA requires all observations to be stored in memory before processing. The batch manner makes robust PCA inefficient for big data. In this paper, we develop an efficient online robust PCA method, namely, online moving window robust principal component analysis (OMWRPCA). Unlike the existing algorithms, OMWRPCA can successfully track not only slowly changing subspaces but also abruptly changing subspaces. Embedding hypothesis testing into the algorithm enables OMWRPCA to detect change points of the underlying subspaces. Extensive numerical experiments, including real-time background subtraction, demonstrate the superior performance of OMWRPCA compared with other state-of-the-art approaches.
Wei Xiao 0001, Xiaolin Huang, Jorge Silva 0002, Saba Emrani, Arin Chaudhuri
IEEE Trans. Multim.2
2020 A Double-Variational Bayesian Framework in Random Fourier Features for Indefinite Kernels
abstract
Random Fourier features (RFFs) have been successfully employed to kernel approximation in large-scale situations. The rationale behind RFF relies on Bochner's theorem, but the condition is too strict and excludes many widely used kernels, e.g., dot-product kernels (violates the shift-invariant condition) and indefinite kernels [violates the positive definite (PD) condition]. In this article, we present a unified RFF framework for indefinite kernel approximation in the reproducing kernel Kreĭn spaces (RKKSs). Besides, our model is also suited to approximate a dot-product kernel on the unit sphere, as it can be transformed into a shift-invariant but indefinite kernel. By the Kolmogorov decomposition scheme, an indefinite kernel in RKKS can be decomposed into the difference of two unknown PD kernels. The spectral distribution of each underlying PD kernel can be formulated as a nonparametric Bayesian Gaussian mixtures model. Based on this, we propose a double-infinite Gaussian mixture model in RFF by placing the Dirichlet process prior. It takes full advantage of high flexibility on the number of components and has the capability of approximating indefinite kernels on a wide scale. In model inference, we develop a non-conjugate variational algorithm with a sub-sampling scheme for the posterior inference. It allows for the non-conjugate case in our model and is quite efficient due to the sub-sampling strategy. Experimental results on several large classification data sets demonstrate the effectiveness of our nonparametric Bayesian model for indefinite kernel approximation when compared to other representative random feature-based methods.
Fanghui Liu 0001, Xiaolin Huang, Lei Shi 0010, Jie Yang 0002, Johan A. K. Suykens
IEEE Trans. Neural Networks Learn. Syst.2
2020 Real-time Image Smoothing via Iterative Least Squares
abstract
Edge-preserving image smoothing is a fundamental procedure for many computer vision and graphic applications. There is a tradeoff between the smoothing quality and the processing speed: the high smoothing quality usually requires a high computational cost, which leads to the low processing speed. In this article, we propose a new global optimization based method, named iterative least squares (ILS), for efficient edge-preserving image smoothing. Our approach can produce high-quality results but at a much lower computational cost. Comprehensive experiments demonstrate that the proposed method can produce results with little visible artifacts. Moreover, the computation of ILS can be highly parallel, which can be easily accelerated through either multi-thread computing or the GPU hardware. With the acceleration of a GTX 1080 GPU, it is able to process images of 1080p resolution (1920 × 1080) at the rate of 20fps for color images and 47fps for gray images. In addition, the ILS is flexible and can be modified to handle more applications that require different smoothing properties. Experimental results of several applications show the effectiveness and efficiency of the proposed method. The code is available at https://github.com/wliusjtu/Real-time-Image-Smoothing-via-Iterative-Least-Squares.
Wei Liu 0044, Xiaolin Huang, Jie Yang 0002, Chunhua Shen, Ian D. Reid 0001
ACM Trans. Graph.3
2019 AirwayNet: A Voxel-Connectivity Aware Approach for Accurate Airway Segmentation Using Convolutional Neural Networks
Yulei Qin, Hao Zheng 0008, Yun Gu, Mali Shen, Jie Yang 0002, Xiaolin Huang, Yue Min Zhu, Guang-Zhong Yang
MICCAI (6)7
2019 A liver fibrosis staging method using cross-contrast network
Yudan Huang, Ying Chen 0006, Haochuan Zhu, Weifeng Li 0004, Xiaolin Huang
Expert Syst. Appl.6
2019 Sparse Kernel Regression with Coefficient-based $\ell_q-$regularization
abstract
In this paper, we consider the $\ell_q-$regularized kernel regression with $0 < q \leq 1$. In form, the algorithm minimizes a least-square loss functional adding a coefficient-based $\ell_q-$penalty term over a linear span of features generated by a kernel function. We study the asymptotic behavior of the algorithm under the framework of learning theory. The contribution of this paper is two-fold. First, we derive a tight bound on the $\ell_2-$empirical covering numbers of the related function space involved in the error analysis. Based on this result, we obtain the convergence rates for the $\ell_1-$regularized kernel regression which is the best so far. Second, for the case $0 < q < 1$, we show that the regularization parameter plays a role as a trade-off between sparsity and convergence rates. Under some mild conditions, the fraction of non-zero coefficients in a local minimizer of the algorithm will tend to $0$ at a polynomial decay rate when the sample size $m$ becomes large. As the concerned algorithm is non-convex, we also discuss how to generate a minimizing sequence iteratively, which can help us to search a local minimizer around any initial point.
Lei Shi 0010, Xiaolin Huang, Yunlong Feng, Johan A. K. Suykens
J. Mach. Learn. Res.2
2019 Robust mixed one-bit compressive sensing
Xiaolin Huang, Yixing Huang, Lei Shi 0010, Andreas K. Maier, Ming Yan 0006
Signal Process.1
2019 Varifocal-Net: A Chromosome Classification Approach Using Deep Convolutional Networks
abstract
Chromosome classification is critical for karyotyping in abnormality diagnosis. To expedite the diagnosis, we present a novel method named Varifocal-Net for simultaneous classification of chromosome's type and polarity using deep convolutional networks. The approach consists of one global-scale network (G-Net) and one local-scale network (L-Net). It follows three stages. The first stage is to learn both global and local features. We extract global features and detect finer local regions via the G-Net. By proposing a varifocal mechanism, we zoom into local parts and extract local features via the L-Net. Residual learning and multi-task learning strategies are utilized to promote high-level feature extraction. The detection of discriminative local parts is fulfilled by a localization subnet of the G-Net, whose training process involves both supervised and weakly supervised learning. The second stage is to build two multi-layer perceptron classifiers that exploit features of both two scales to boost classification performance. The third stage is to introduce a dispatch strategy of assigning each chromosome to a type within each patient case, by utilizing the domain knowledge of karyotyping. The evaluation results from 1909 karyotyping cases showed that the proposed Varifocal-Net achieved the highest accuracy per patient case (%) of 99.2 for both type and polarity tasks. It outperformed state-of-the-art methods, demonstrating the effectiveness of our varifocal mechanism, multi-scale feature ensemble, and dispatch strategy. The proposed method has been applied to assist practical karyotype diagnosis.
Yulei Qin, Hao Zheng 0008, Xiaolin Huang, Jie Yang 0002, Yue Min Zhu, Lingqian Wu, Guang-Zhong Yang
IEEE Trans. Medical Imaging4
2019 Indefinite Kernel Logistic Regression With Concave-Inexact-Convex Procedure
abstract
In kernel methods, the kernels are often required to be positive definitethat restricts the use of many indefinite kernels. To consider those nonpositive definite kernels, in this paper, we aim to build an indefinite kernel learning framework for kernel logistic regression (KLR). The proposed indefinite KLR (IKLR) model is analyzed in the reproducing kernel Kreĭn spaces and then becomes nonconvex. Using the positive decomposition of a nonpositive definite kernel, the derived IKLR model can be decomposed into the difference of two convex functions. Accordingly, a concave-convex procedure (CCCP) is introduced to solve the nonconvex optimization problem. Since the CCCP has to solve a subproblem in each iteration, we propose a concave-inexact-convex procedure (CCICP) algorithm with an inexact solving scheme to accelerate the solving process. Besides, we propose a stochastic variant of CCICP to efficiently obtain a proximal solution, which achieves the similar purpose with the inexact solving scheme in CCICP. The convergence analyses of the above-mentioned two variants of CCCP are conducted. By doing so, our method works effectively not only in a deterministic setting but also in a stochastic setting. Experimental results on several benchmarks suggest that the proposed IKLR model performs favorably against the standard (positive definite) KLR and other competitive indefinite learning-based algorithms.
Fanghui Liu 0001, Xiaolin Huang, Chen Gong 0002, Jie Yang 0002, Johan A. K. Suykens
IEEE Trans. Neural Networks Learn. Syst.2
2018 Nonlinear Pairwise Layer and Its Training for Kernel Learning
abstract
Kernel learning is a fundamental technique that has been intensively studied in the past decades. For the complicated practical tasks, the traditional "shallow" kernels (e.g., Gaussian kernel and sigmoid kernel) are not flexible enough to produce satisfactory performance. To address this shortcoming, this paper introduces a nonlinear layer in kernel learning to enhance the model flexibility. This layer is pairwise, which fully considers the coupling information among examples. So our model contains a fixed single mapping layer (i.e. a Gaussian kernel) as well as a nonlinear pairwise layer, thereby achieving better flexibility than the existing kernel structures. Moreover, the proposed structure can be seamlessly embedded to Support Vector Machines (SVM), of which the training process can be formulated as a joint optimization problem including nonlinear function learning and standard SVM optimization. We theoretically prove that the objective function is gradient-Lipschitz continuous, which further guides us how to accelerate the optimization process in a deep kernel architecture. Experimentally, we find that the proposed structure outperforms other state-ofthe-art kernel-based algorithms on various benchmark datasets, and thus the effectiveness of the incorporated pairwise layer with its training approach is demonstrated.
Fanghui Liu 0001, Xiaolin Huang, Chen Gong 0002, Jie Yang 0002, Li Li 0013
AAAI2
2018 Non-destructive Digitization of Soiled Historical Chinese Bamboo Scrolls
abstract
For about 2000 years, no paper was used as a media in China but writings and drawings were captured on bamboo and wooden slips. Several slips were bound together with strips and rolled up to a scroll. The writings and drawings were either brushed or even carved into the wood. Those documents are very precious for culture inheritance and research, but due to aging processes, the discovered pieces are sometimes in a poor condition and also soiled. Because cleaning the slips is not only challenging but also writings could be erased, we developed a method to digitize such historical documents without the need of cleaning. We perform a 3-D X-ray micro-CT scan resulting in a 3-D volume of the complete document. With our approach, we were able to investigate the scroll without any manual labor (e.g. unwrapping or cleaning). We showed that the method also works for heavily soiled scrolls where nothing is readable with the naked eye. This can help conservators to store all writings before they may be erased by the cleaning process. Finally, we present a manual technique to virtually unwrap and post-process the documents resulting in a 2-D image of all bamboo slips.
Daniel Stromer, Vincent Christlein, Andreas K. Maier, Patrick Zippert, Eric Helmecke, Tino Hausotte, Xiaolin Huang
DAS7
2018 Privileged Semi-Supervised Learning
abstract
Semi-Supervised Learning (SSL) aims to leverage unlabeled data to improve performance. Due to the lack of supervised information, previous works mainly focus on how to utilize the available unlabeled data to improve the training quality. However, the estimation of the data distribution revealed by the unlabeled examples might not be accurate as their classes are unknown. Inspired by the framework of Learning Using Privileged Information (LUPI), we propose to introduce an intelligent “teacher” which can utilize the privileged information (i.e. more precise explanations of training set) to improve the performance of SSL. This developed method is named as Privileged Semi-Supervised Learning (PSSL). Moreover, our method can be efficiently solved with l2-loss. Experimentally, we compare our PSSL with several typical algorithms on three real-world datasets, and the results suggest that our approach is able to achieve the state-of-the-art performance.
Chen Gong 0002, Chao Ma 0005, Xiaolin Huang, Jie Yang 0002
ICIP4
2018 Discrete Locally-Linear Preserving Hashing
abstract
Recently, hashing has attracted considerable attention for nearest neighbor search due to its fast query speed and low storage cost. However, existing unsupervised hashing algorithms have two problems in common. Firstly, the widely utilized anchor graph construction algorithm has inherent limitations in local weight estimation. Secondly, the locally linear structure in the original feature space is seldom taken into account for binary encoding. Therefore, in this paper, we propose a novel unsupervised hashing method, dubbed “dis-crete locally-linear preserving hashing”, which effectively calculates the adjacent matrix while preserving the locally linear structure in the obtained hash space. Specifically, a novel local anchor embedding algorithm is adopted to construct the approximate adjacent matrix. After that, we directly minimize the reconstruction error with the discrete constrain to learn the binary codes. Experimental results on two typical image datasets indicate that the proposed method significantly outperforms the state-of-the-art unsupervised methods.
Xiang Li 0041, Chao Ma 0005, Jie Yang 0002, Xiaolin Huang
ICIP4
2018 Small Lesion Classification in Dynamic Contrast Enhancement MRI for Breast Cancer Early Detection
Hao Zheng 0008, Yun Gu, Yulei Qin, Xiaolin Huang, Jie Yang 0002, Guang-Zhong Yang
MICCAI (2)4
2018 New Motion Estimation with Angular-Distance Median Filter for Frame Interpolation
HuangKai Cai, Xiaolin Huang, Jie Yang 0002
PRCV (1)3
2018 Violence Detection Based on Spatio-Temporal Feature and Fisher Vector
HuangKai Cai, Xiaolin Huang, Jie Yang 0002, Xiangjian He
PRCV (1)3
2018 Deep Supervised Auto-encoder Hashing for Image Retrieval
Sanli Tang, Haoyuan Chi, Jie Yang 0002, Xiaolin Huang, Masoumeh Zareapoor
PRCV (2)4
2018 Pinball loss minimization for one-bit compressive sensing: Convex models and algorithms
Xiaolin Huang, Lei Shi 0010, Ming Yan 0006, Johan A. K. Suykens
Neurocomputing1
2018 Multi-modal self-paced learning for image classification
Wei Xu 0050, Wei Liu 0044, Xiaolin Huang, Jie Yang 0002, Song Qiu
Neurocomputing3
2018 Indefinite kernel spectral learning
Siamak Mehrkanoon, Xiaolin Huang, Johan A. K. Suykens
Pattern Recognit.2
2018 Nonconvex penalties with analytical solutions for one-bit compressive sensing
Xiaolin Huang, Ming Yan 0006
Signal Process.1
2018 Multi-task classification with sequential instances and tasks
Wei Xu 0050, Wei Liu 0044, Haoyuan Chi, Xiaolin Huang, Jie Yang 0002
Signal Process. Image Commun.4
2018 Fast Signal Recovery From Saturated Measurements by Linear Loss and Nonconvex Penalties
abstract
Sign information is the key for overcoming the inevitable saturation error in compressive sensing system, which causes loss of information and may result in great bias. For sparse signal recovery from saturation, we propose to use linear loss to improve the effectiveness from the existing methods that utilize hard constraints/hinge loss for sign consistency. Due to the use of linear loss, analytical solution in the update progress is obtained and some nonconvex penalties are applicable, e.g., minimax concave penalty, ℓ0norm, and sorted ℓ1norm. Theoretical analysis reveals that the estimation error can still be bounded. Generally, with linear loss and nonconvex penalties, the recovery performance can be significantly improved and the computational time is also largely saved, which is verified by the numerical experiments.
Xiaolin Huang, Yipeng Liu 0001, Ming Yan 0006
IEEE Signal Process. Lett.2
2018 Robust Visual Tracking Revisited: From Correlation Filter to Template Matching
abstract
In this paper, we propose a novel matching based tracker by investigating the relationship between template matching and the recent popular correlation filter based trackers (CFTs). Compared to the correlation operation in CFTs, a sophisticated similarity metric termed mutual buddies similarity is proposed to exploit the relationship of multiple reciprocal nearest neighbors for target matching. By doing so, our tracker obtains powerful discriminative ability on distinguishing target and background as demonstrated by both empirical and theoretical analyses. Besides, instead of utilizing single template with the improper updating scheme in CFTs, we design a novel online template updating strategy named memory, which aims to select a certain amount of representative and reliable tracking results in history to construct the current stable and expressive template set. This scheme is beneficial for the proposed tracker to comprehensively understand the target appearance variations, recall some stable results. Both qualitative and quantitative evaluations on two benchmarks suggest that the proposed tracking method performs favorably against some recently developed CFTs and other competitive trackers.
Fanghui Liu 0001, Chen Gong 0002, Xiaolin Huang, Tao Zhou 0002, Jie Yang 0002, Dacheng Tao
IEEE Trans. Image Process.3
2018 Modified Sparse Linear-Discriminant Analysis via Nonconvex Penalties
abstract
This paper considers the linear-discriminant analysis (LDA) problem in the undersampled situation, in which the number of features is very large and the number of observations is limited. Sparsity is often incorporated in the solution of LDA to make a well interpretation of the results. However, most of the existing sparse LDA algorithms pursue sparsity by means of the $\ell _{1}$ -norm. In this paper, we give elaborate analysis for nonconvex penalties, including the $\ell _{0}$ -based and the sorted $\ell _{1}$ -based LDA methods. The latter one can be regarded as a bridge between the $\ell _{0}$ and $\ell _{1}$ penalties. These nonconvex penalty-based LDA algorithms are evaluated on the gene expression array and face database, showing high classification accuracy on real-world problems.
Xiaolin Huang
IEEE Trans. Neural Networks Learn. Syst.2
2018 Classification With Truncated $\ell _{1}$ Distance Kernel
abstract
This brief proposes a truncated distance (TL1) kernel, which results in a classifier that is nonlinear in the global region but is linear in each subregion. With this kernel, the subregion structure can be trained using all the training data and local linear classifiers can be established simultaneously. The TL1 kernel has good adaptiveness to nonlinearity and is suitable for problems which require different nonlinearities in different areas. Though the TL1 kernel is not positive semidefinite, some classical kernel learning methods are still applicable which means that the TL1 kernel can be directly used in standard toolboxes by replacing the kernel evaluation. In numerical experiments, the TL1 kernel with a pregiven parameter achieves similar or better performance than the radial basis function kernel with the parameter tuned by cross validation, implying the TL1 kernel a promising nonlinear kernel for classification tasks.
Xiaolin Huang, Johan A. K. Suykens, Shuning Wang, Joachim Hornegger, Andreas K. Maier
IEEE Trans. Neural Networks Learn. Syst.1
2017 Self-paced least square semi-coupled dictionary learning for person re-identification
abstract
Person re-identification aims to match people across disjoint camera views. It has been reported that Least Square Semi-Coupled Dictionary Learning (LSSCDL) based sample-specific SVM learning framework has obtained the state of the art performance. However, the objective function of the LSSCDL, the algorithm of learning the pairs (feature, weight) dictionaries and the mapping function between feature space and weight space, is non-convex, which usually result in suboptimal solutions with the bad local minima of the objective function. To tackle with this constraint, we present Self-Paced Least Square Semi-Coupled Dictionary Learning (SLSSCDL) algorithm, which is inspired by previous works on self-paced learning, a framework able to improve the accuracy of conventional learning models by presenting the training data in a meaningful order to get a better local minima, i.e. easy samples are provided first. In addition, a graph based regularization term is also introduced to preserve the local similarities in each space. Experimental results show that SLSSCDL gains competitive performance on two challenging datasets.
Wei Xu 0050, Haoyuan Chi, Lei Zhou 0003, Xiaolin Huang, Jie Yang 0002
ICIP4
2017 Robust Kernel Approximation for Classification
Fanghui Liu 0001, Xiaolin Huang, Cheng Peng 0004, Jie Yang 0002, Nikola K. Kasabov
ICONIP (1)2
2017 A New Bayesian Method for Jointly Sparse Signal Recovery
Xiaolin Huang, Cheng Peng 0004, Jie Yang 0002, Li Li 0013
ICONIP (4)2
2017 Indefinite Kernel Logistic Regression
abstract
Traditionally, kernel learning methods require positive definitiveness on the kernel, which is too strict and excludes many sophisticated similarities, that are indefinite. To utilize those indefinite kernels, indefinite learning methods are of great interests. This paper aims at the extension of the logistic regression from positive definite kernels to indefinite ones. The proposed model, named indefinite kernel logistic regression (IKLR), keeps consistency to the regular KLR in formulation but it essentially becomes non-convex. Thanks to the positive decomposition of an indefinite kernel, IKLR can be transformed into a difference of two convex models, which follows the use of concave-convex procedure. Moreover, aiming at large-scale problems in practice, a concave-inexact-convex procedure (CCICP) algorithm with an inexact solving scheme is proposed with convergence guarantees. Experimental results on multi-modal datasets demonstrate the superiority of the proposed IKLR model over kernel logistic regression with positive definite kernels and other state-of-the-art indefinite learning based methods.
Fanghui Liu 0001, Xiaolin Huang, Jie Yang 0002
ACM Multimedia2
2017 Robust kernel canonical correlation analysis with applications to information retrieval
Xiaolin Huang
Eng. Appl. Artif. Intell.2
2017 Adaptive block coordinate DIRECT algorithm
Qinghua Tao, Xiaolin Huang, Shuning Wang, Li Li 0013
J. Glob. Optim.2
2017 Efficient and Robust Corner Detectors Based on Second-Order Difference of Contour
abstract
As one of the most significant local features of image, corner is widely used in many computer vision tasks. Corner detection aims to achieve the highest possible detection accuracy while minimizing the computational complexity. In this letter, we first introduce a new measurement termed as second-order difference of contour (SODC), and then examine its regular distribution, which is found to provide useful information to distinguish corners from noncorners. Based on the SODC distribution characteristics, we propose two novel corner detectors to measure the response of contour points using Manhattan distance and Euclidean distance, respectively. Numerical experiments demonstrate that the Manhattan detector greatly decreases the computational complexity, while the Euclidean detector outperforms the state-of-the-art corner detectors in terms of repeatability and localization error.
Ce Zhu, Qian Zhang 0047, Xiaolin Huang, Yipeng Liu 0001
IEEE Signal Process. Lett.4
2017 Hybrid CS-DMRI: Periodic Time-Variant Subsampling and Omnidirectional Total Variation Based Reconstruction
abstract
Compressive sensing (CS) has been used to accelerate dynamic magnetic resonance imaging (DMRI). Currently, the online CS-DMRI is faster, whereas the offline CS-DMRI provides higher accuracy for image reconstruction. To achieve good image reconstruction performance in terms of both speed and accuracy, we propose a hybrid CS-DMRI method using periodic time-variant subsampling for different frames. In each period, there is one reference frame that is sampled at a higher subsampling ratio. The two nearby reference frames with good reconstruction quality can be used to provide rough predictions of the other frames between them. To finely recover the current frame, one structural regularization in the optimization model for reconstruction is a 2-D omnidirectional total variation (OTV) for exploiting the sparsity of the difference between the predicted and estimated frames, and the other is a 3-D OTV as a regularization term for exploiting the bilateral spatio-temporal coherence between the forward reference frame, current frame, and backward reference frame. Compared with classical total variation, the proposed OTV fully utilizes the correlations of all the possible directions of the data. The formulated optimization model can be solved using iterative reweighted least squares with the pre-conditioned conjugate gradient method. Numerical experiments demonstrate that the proposed method has better reconstruction accuracy than all the existing methods and low computational complexity that is comparable to the existing online methods.
Yipeng Liu 0001, Xiaolin Huang, Ce Zhu
IEEE Trans. Medical Imaging3
2017 Dynamic 2-D/3-D Rigid Registration Framework Using Point-To-Plane Correspondence Model
abstract
In image-guided interventional procedures, live 2-D X-ray images can be augmented with preoperative 3-D computed tomography or MRI images to provide planning landmarks and enhanced spatial perception. An accurate alignment between the 3-D and 2-D images is a prerequisite for fusion applications. This paper presents a dynamic rigid 2-D/3-D registration framework, which measures the local 3-D-to-2-D misalignment and efficiently constrains the update of both planar and non-planar 3-D rigid transformations using a novel point-to-plane correspondence model. In the simulation evaluation, the proposed method achieved a mean 3-D accuracy of 0.07 mm for the head phantom and 0.05 mm for the thorax phantom using single-view X-ray images. In the evaluation on dynamic motion compensation, our method significantly increases the accuracy comparing with the baseline method. The proposed method is also evaluated on a publicly-available clinical angiogram data set with "gold-standard" registrations. The proposed method achieved a mean 3-D accuracy below 0.8 mm and a mean 2-D accuracy below 0.3 mm using single-view X-ray images. It outperformed the state-of-the-art methods in both accuracy and robustness in single-view registration. The proposed method is intuitive, generic, and suitable for both initial and dynamic registration scenarios.
Jian Wang 0009, Roman Schaffert, Anja Borsdorf, Benno Heigl, Xiaolin Huang, Joachim Hornegger, Andreas K. Maier
IEEE Trans. Medical Imaging5
2017 Solution Path for Pin-SVM Classifiers With Positive and Negative τ Values
abstract
Applying the pinball loss in a support vector machine (SVM) classifier results in pin-SVM. The pinball loss is characterized by a parameter τ . Its value is related to the quantile level and different τ values are suitable for different problems. In this paper, we establish an algorithm to find the entire solution path for pin-SVM with different τ values. This algorithm is based on the fact that the optimal solution to pin-SVM is continuous and piecewise linear with respect to τ . We also show that the nonnegativity constraint on τ is not necessary, i.e., τ can be extended to negative values. First, in some applications, a negative τ leads to better accuracy. Second, τ = -1 corresponds to a simple solution that links SVM and the classical kernel rule. The solution for τ = -1 can be obtained directly and then be used as a starting point of the solution path. The proposed method efficiently traverses τ values through the solution path, and then achieves good performance by a suitable τ . In particular, τ = 0 corresponds to C-SVM, meaning that the traversal algorithm can output a result at least as good as C-SVM with respect to validation error.
Xiaolin Huang, Lei Shi 0010, Johan A. K. Suykens
IEEE Trans. Neural Networks Learn. Syst.1
2016 Cost-sensitive sparse linear regression for crowd counting with imbalanced training data
abstract
Video-based crowd counting (VCC) is a high demanded technique in many video applications. Existing supervised VCC methods essentially learn an intrinsic mapping function between image features and corresponding crowd counts. However, imbalanced training dataset degrades the performance of VCC significantly. Encouraged by recent success in cost-sensitive learning for image classification with imbalance dataset, we propose a novel cost-sensitive sparse linear regression VCC method (CS-SLR-VCC). Specifically, a sparse linear regression (SLR) model is firstly learned and the modelling errors associated with each training data are calculated accordingly. Then, aiming to eliminate the adverse effect of the high modelling errors of SLR model due to imbalanced data, all modelling errors are taken as prior knowledge to design sample-related weighting factors. Thus, a cost-sensitive SLR model is reformulated and its optimal solution is derived. Extensive experiments conducted on public UCSD and Mall benchmarks demonstrate the superior performance of our proposed CS-SLR-VCC method.
Xiaolin Huang, Yuexian Zou, Yi Wang 0033
ICME1
2016 Example-based visual object counting with a sparsity constraint
abstract
For existing mainstream visual object counting (VOC) methods, training data insufficiency will lead to significant performance degradation. To address this challenge, we propose a novel sparsity-constrained example-based VOC method. Given a test image, its counts are estimated by integrating over its density map, and our method will predict such density map based on patch using training examples. Specifically, image patches and their counterpart density maps generated from annotated training images share similar local geometry on manifolds. Such local geometry can be captured by locally linear embedding (LLE) only when data are well-sampled. However, training data are poorly sampled due to their insufficiency. To handle this problem, we impose sparsity on the local optimization based on LLE, where the chosen examples favor the similar structure of input patches. Extensive experiments on public datasets demonstrate the effectiveness and competitiveness of our method by using simple features and a few training images.
Yi Wang 0033, Yuexian Zou, Xiaolin Huang, Cheng Cai
ICME4
2016 Robust Support Vector Machines for Classification with Nonconvex and Smooth Losses
abstract
This letter addresses the robustness problem when learning a large margin classifier in the presence of label noise. In our study, we achieve this purpose by proposing robustified large margin support vector machines. The robustness of the proposed robust support vector classifiers (RSVC), which is interpreted from a weighted viewpoint in this work, is due to the use of nonconvex classification losses. Besides the robustness, we also show that the proposed RSCV is simultaneously smooth, which again benefits from using smooth classification losses. The idea of proposing RSVC comes from M-estimation in statistics since the proposed robust and smooth classification losses can be taken as one-sided cost functions in robust statistics. Its Fisher consistency property and generalization ability are also investigated. Besides the robustness and smoothness, another nice property of RSVC lies in the fact that its solution can be obtained by solving weighted squared hinge loss-based support vector machine problems iteratively. We further show that in each iteration, it is a quadratic programming problem in its dual space and can be solved by using state-of-the-art methods. We thus propose an iteratively reweighted type algorithm and provide a constructive proof of its convergence to a stationary point. Effectiveness of the proposed classifiers is verified on both artificial and real data sets.
Yunlong Feng, Xiaolin Huang, Siamak Mehrkanoon, Johan A. K. Suykens
Neural Comput.3
2016 Coordinate Descent Algorithm for Ramp Loss Linear Programming Support Vector Machines
Xiangming Xi, Xiaolin Huang, Johan A. K. Suykens, Shuning Wang
Neural Process. Lett.2
2016 Multiple Gaussian graphical estimation with jointly sparse penalty
Qinghua Tao, Xiaolin Huang, Shuning Wang, Xiangming Xi, Li Li 0013
Signal Process.2
2015 Sequential minimal optimization for SVM with pinball loss
Xiaolin Huang, Lei Shi 0010, Johan A. K. Suykens
Neurocomputing1
2015 Learning with the maximum correntropy criterion induced losses for regression
Yunlong Feng, Xiaolin Huang, Lei Shi 0010, Johan A. K. Suykens
J. Mach. Learn. Res.2
2015 Two-level ℓ1 minimization for compressed sensing
Xiaolin Huang, Yipeng Liu 0001, Lei Shi 0010, Sabine Van Huffel, Johan A. K. Suykens
Signal Process.1
2015 Signal recovery for jointly sparse vectors with different sensing matrices
Li Li 0013, Xiaolin Huang, Johan A. K. Suykens
Signal Process.2
2014 Non-parallel support vector classifiers with different loss functions
Siamak Mehrkanoon, Xiaolin Huang, Johan A. K. Suykens
Neurocomputing2
2014 Ramp loss linear programming support vector machine
Xiaolin Huang, Lei Shi 0010, Johan A. K. Suykens
J. Mach. Learn. Res.1
2014 Support Vector Machine Classifier With Pinball Loss
abstract
Traditionally, the hinge loss is used to construct support vector machine (SVM) classifiers. The hinge loss is related to the shortest distance between sets and the corresponding classifier is hence sensitive to noise and unstable for re-sampling. In contrast, the pinball loss is related to the quantile distance and the result is less sensitive. The pinball loss has been deeply studied and widely applied in regression but it has not been used for classification. In this paper, we propose a SVM classifier with the pinball loss, called pin-SVM, and investigate its properties, including noise insensitivity, robustness, and misclassification error. Besides, insensitive zone is applied to the pin-SVM for a sparse model. Compared to the SVM with the hinge loss, the proposed pin-SVM has the same computational complexity and enjoys noise insensitivity and re-sampling stability.
Xiaolin Huang, Lei Shi 0010, Johan A. K. Suykens
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 Fixed-size Pegasos for hinge and pinball loss SVM
abstract
Pegasos has become a widely acknowledged algorithm for learning linear Support Vector Machines. It utilizes properties of hinge loss and theory of strongly convex optimization problems for fast convergence rates and lower computational and memory costs. In this paper we adopt the recently proposed pinball loss for the Pegasos algorithm and show some advantages of using it in a variety of classification problems. First we present the newly derived Pegasos optimization objective with respect to pinball loss and analyze its properties and convergence rates. Additionally we present extensions of the Pegasos algorithm applied to the kernel-induced and Nyström approximated feature maps which introduce non-linearity in the input space. This is done using a Fixed-Size kernel method approach. Second we give experimental results for publicly available UCI datasets to justify the advantages and the importance of pinball loss for achieving a better classification accuracy and greater numerical stability in the partially or fully stochastic setting. Finally we conclude our paper with a brief discussion of the applicability of pinball loss to real-life problems.
Vilen Jumutc, Xiaolin Huang, Johan A. K. Suykens
IJCNN2
2013 Support vector machines with piecewise linear feature mapping
Xiaolin Huang, Siamak Mehrkanoon, Johan A. K. Suykens
Neurocomputing1
2013 Hierarchical Minutiae Matching for Fingerprint and Palmprint Identification
abstract
Fingerprints and palmprints are the most common authentic biometrics for personal identification, especially for forensic security. Previous research have been proposed to speed up the searching process in fingerprint and palmprint identification systems, such as those based on classification or indexing, in which the deterioration of identification accuracy is hard to avert. In this paper, a novel hierarchical minutiae matching algorithm for fingerprint and palmprint identification systems is proposed. This method decomposes the matching step into several stages and rejects many false fingerprints or palmprints on different stages, thus it can save much time while preserving a high identification rate. Experimental results show that the proposed algorithm can save almost 50% searching time compared with traditional methods and illustrate its effectiveness.
Fanglin Chen 0001, Xiaolin Huang, Jie Zhou 0001
IEEE Trans. Image Process.2
2013 Hinging Hyperplanes for Time-Series Segmentation
abstract
Division of a time series into segments is a common technique for time-series processing, and is known as segmentation. Segmentation is traditionally done by linear interpolation in order to guarantee the continuity of the reconstructed time series. The interpolation-based segmentation methods may perform poorly for data with a level of noise because interpolation is noise sensitive. To handle the problem, this paper establishes an explicit expression for segmentation from a compact representation for piecewise linear functions using hinging hyperplanes. This expression enables the use of regression to obtain a continuous reconstructed signal and, as a consequence, application of advanced techniques in segmentation. In this paper, a least squares support vector machine with lasso using a hinging feature map is given and analyzed, based on which a segmentation algorithm and its online version are established. Numerical experiments conducted on synthetic and real-world datasets demonstrate the advantages of our methods compared to existing segmentation algorithms.
Xiaolin Huang, Marin Matijas, Johan A. K. Suykens
IEEE Trans. Neural Networks Learn. Syst.1
2012 Nonlinear system identification with continuous piecewise linear neural network
Xiaolin Huang, Jun Xu 0008, Shuning Wang
Neurocomputing1
2010 Operation optimization for centrifugal chiller plants using continuous piecewise linear programming
abstract
Centrifugal chiller plants (CCP) are widely used in air conditioning systems, its operation optimization can save lots of energy and has great significance in environmental protection. The optimization is a large-scale nonlinear problem and there is no practical algorithm until now. This paper proposes a new method to do this operation optimization using continuous piecewise linear programming (CPWLP). The main idea is transforming the nonlinear problem into a series of linear programmings by approximating the original system using piecewise linear representation. For CPWLP, some properties are discussed and an algorithm is given. Using CPWLP, CCP system is optimized and its energy performance is improved significantly.
Xiaolin Huang, Jun Xu 0008, Shuning Wang
SMC1
2010 A neural network of smooth hinge functions
abstract
Smooth hinging hyperplane (SHH) has been proposed as an improvement over the well-known hinging hyperplane (HH) by the fact that it retains the useful features of HH while overcoming HH's drawback of nondifferentiability. This paper introduces a formal characterization of smooth hinge function (SHF), which can be used to generate SHH as a neural network. A method for the general construction of SHF is also given. Furthermore, the work proves that SHH is better than HH in functional approximation, i.e., the optimal error of SHH approximating a general function is always smaller or equal to that of HH. Particularly, in the case that the SHF is generated via the integration of a class of sigmoidal functions, it is further proven that the corresponding SHH of the 2m SHFs would outperform a neural network with m of the sigmoidal function from which the SHF is derived. Any upper bound established on the approximation error of a neural network of m sigmoidal activation functions can hence be translated to the SHH of m SHFs by replacing m with [m/2]. The work also includes an algorithm for the identification of SHH making use of its differentiability property. Simulation experiments are presented to validate the theoretical conclusions to possible extent.
Shuning Wang, Xiaolin Huang, Yeung Yam
IEEE Trans. Neural Networks2
2008 Configuration of Continuous Piecewise-Linear Neural Networks
abstract
The problem of constructing a general continuous piecewise-linear neural network is considered in this paper. It is shown that every projection domain of an arbitrary continuous piecewise-linear function can be partitioned into convex polyhedra by using difference functions of its local linear functions. Based on these convex polyhedra, a group of continuous piecewise-linear basis functions are formulated. It is proven that a linear combination of these basis functions plus a constant, which we call a standard continuous piecewise-linear neural network, can represent all continuous piecewise-linear functions. In addition, the proposed standard continuous piecewise-linear neural network is applied to solve some function approximation problems. A number of numerical experiments are presented to illustrate that the standard continuous piecewise-linear neural network can be a promising tool for function approximation.
Shuning Wang, Xiaolin Huang, Junaid M. Khan
IEEE Trans. Neural Networks2