Hui Xue 0002

dblp:27/3541-2 · DBLP profile ↗
← Back
81ranked-venue papers
18as first author
44since 2021 · last 2026
0000-0002-5856-4445ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 61 · 16 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 4 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 4 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Intra-Image Mining and Symmetric Maximum Concept Matching for Few Shot Out-of-Distribution Detection
abstract
Recent vision-language model (VLM)-based methods have achieved promising results in zero-shot out-of-distribution (OOD) detection by effectively leveraging the local patch features. However, the zero-shot nature inherently comes with two limitations: 1) imperfect local feature prototypes; 2) lack of OOD prototypes. In this paper, we propose Intra-Image Mining (IIM), a lightweight framework designed to overcome these limitations in a few-shot manner. IIM is motivated by the fact that local patches within an image often exhibit diverse semantics, with some patches deviating from the main class concept. Therefore, for each image, we first select the top-k class prototype-related patches as positive samples and leverage them to refine and optimize the local feature prototype. Then, the next top-k among the remaining patches are selected as negatives—serving as OOD signals to construct OOD prototypes. This process yields coherent local positives and challenging negatives, effectively enhancing the model’s local feature discrimination. Besides, we propose a novel inference strategy named Symmetric Maximum Concept Matching (S-MCM). While existing approaches typically adopt an image-to-text scheme—comparing the image features to textual class prototypes—S-MCM further incorporate a text-to-image perspective, leading to more reliable OOD detection. We also propose two benchmarks to analyze the impact of semantic diversity within ID dataset. Built on a frozen VLM, IIM, in conjunction with S-MCM, achieves consistent gains in OOD detection on ImageNet-1k and other benchmarks, outperforming prior methods in FPR95 and AUROC across various few-shot settings.
Kaixiang Chen, Pengfei Fang, Hui Xue 0002
AAAI3
2026 HAP: Harmonized Amplitude Perturbation for Cross-Domain Few-Shot Learning
abstract
Cross-Domain Few-Shot Learning (CD-FSL) remains a significant challenge due to substantial distribution shifts between source and target domains. While prior approaches primarily focus on spatial alignment, they often overlook discrepancies in the frequency domain. In this paper, we reveal frequency band discretization as a key phenomenon, characterized by intra-domain low-frequency dominance, inter-domain amplitude divergence, and limited high-frequency variation. This spectral disharmony biases models toward low-frequency components, leading to spectral collapse. We quantify spectral collapse via the effective rank, a principled measure of spectral diversity. To mitigate spectral collapse, we propose Harmonized Amplitude Perturbation (HAP), a frequency-domain augmentation strategy that perturbs the amplitude spectrum via frequency-aware gains sampled from Harmonized Distributions, while fixing the phase spectrum to maintain semantic integrity. Extensive experiments on both Cross-Domain Few-Shot Image Classification and Object Detection benchmarks demonstrate that HAP effectively increases spectral diversity and consistently improves generalization, outperforming state-of-the-art methods without introducing extra model complexity.
Wenqian Li 0005, Pengfei Fang, Hui Xue 0002
AAAI3
2026 Adaptive Hyperbolic Kernels: Modulated Embedding in de Branges-Rovnyak Spaces
abstract
Hierarchical data pervades diverse machine learning applications, including natural language processing, computer vision, and social network analysis. Hyperbolic space, characterized by its negative curvature, has demonstrated strong potential in such tasks due to its capacity to embed hierarchical structures with minimal distortion. Previous evidence indicates that the hyperbolic representation capacity can be further enhanced through kernel methods. However, existing hyperbolic kernels still suffer from mild geometric distortion or lack adaptability. This paper addresses these issues by introducing a curvature-aware de Branges–Rovnyak space, a reproducing kernel Hilbert space (RKHS) that is isometric to a Poincaré ball. We design an adjustable multiplier to select the appropriate RKHS corresponding to the hyperbolic space with any curvature adaptively. Building on this foundation, we further construct a family of adaptive hyperbolic kernels, including the novel adaptive hyperbolic radial kernel, whose learnable parameters modulate hyperbolic features in a task-aware manner. Extensive experiments on visual and language benchmarks demonstrate that our proposed kernels outperform existing hyperbolic kernels in modeling hierarchical dependencies.
Leping Si, Meimei Yang, Hui Xue 0002, Shipeng Zhu, Pengfei Fang
AAAI3
2026 InsNet: Deep indefinite spectral kernel network
Yanfang Xue, Hui Xue 0002, Shipeng Zhu
Pattern Recognit.2
2026 Conditional Self-Supervised Learning for Few-Shot Learning
abstract
Few-shot learning seeks to emulate humans by grasping a new concept with a few examples. However, it is often a tricky problem to completely learn a new concept and avoid falling into overfitting with few examples. Recently, self-supervision has been introduced as an auxiliary pretext task to enhance the generalization of few-shot learning models. However, in few-shot scenarios, self-supervised learning may easily learn undesirable shortcuts and biased representations due to insufficient information, negatively impacting the decision boundary of few-shot models. In this paper, we propose a new learning paradigm called conditional self-supervised learning (CSS), where the blind self-supervised tasks are constrained. Specifically, we first utilize the embedding knowledge of supervised learning as a condition to guide the representation learning process of self-supervised tasks. In order to optimize the overall objectives, we formulate the CSS as a multiple objective optimization problem to find Pareto optimal solutions. Moreover, we combine the meaningful information extracted from both supervised and self-supervised learning into a unified distribution, further enriching the original representation. Extensive experiments show that our method, without any fine-tuning, significantly improves accuracy in few-shot tasks compared to state-of-the-art methods.
Yuexuan An, Hui Xue 0002, Xingyu Zhao 0002
IEEE Trans. Circuits Syst. Video Technol.2
2025 SVasP: Self-Versatility Adversarial Style Perturbation for Cross-Domain Few-Shot Learning
abstract
Cross-Domain Few-Shot Learning (CD-FSL) aims to transfer knowledge from seen source domains to unseen target domains, which is crucial for evaluating the generalization and robustness of models. Recent studies focus on utilizing visual styles to bridge the domain gap between different domains. However, the serious dilemma of gradient instability and local optimization problem occurs in those style-based CD-FSL methods. This paper addresses these issues and proposes a novel crop-global style perturbation method, called Self-Versatility Adversarial Style Perturbation (SVasP), which enhances the gradient stability and escapes from poor sharp minima jointly. Specifically, SVasP simulates more diverse potential target domain adversarial styles via diversifying input patterns and aggregating localized crop style gradients, to serve as global style perturbation stabilizers within one image, a concept we refer to as self-versatility. Then a novel objective function is proposed to maximize visual discrepancy while maintaining semantic consistency between global, crop, and adversarial features. Having the stabilized global style perturbation in the training phase, one can obtain a flattened minima in the loss landscape, boosting the transferability of the model to the target domains. Extensive experiments on multiple benchmark datasets demonstrate that our method significantly outperforms existing state-of-the-art methods.
Wenqian Li 0005, Pengfei Fang, Hui Xue 0002
AAAI3
2025 PEARL: Input-Agnostic Prompt Enhancement with Negative Feedback Regulation for Class-Incremental Learning
abstract
Class-incremental learning (CIL) aims to continuously introduce novel categories into a classification system without forgetting previously learned ones, thus adapting to evolving data distributions. Researchers are currently focusing on leveraging the rich semantic information of pre-trained models (PTMs) in CIL tasks. Prompt learning has been adopted in CIL for its ability to adjust data distribution to better align with pre-trained knowledge. This paper critically examines the limitations of existing methods from the perspective of prompt learning, which heavily rely on input information. To address this issue, we propose a novel PTM-based CIL method called Input-Agnostic Prompt Enhancement with NegAtive Feedback ReguLation (PEARL). In PEARL, we implement an input-agnostic global prompt coupled with an adaptive momentum update strategy to reduce the model's dependency on data distribution, thereby effectively mitigating catastrophic forgetting. Guided by negative feedback regulation, this adaptive momentum update addresses the parameter sensitivity inherent in fixed-weight momentum updates. Furthermore, it fosters the continuous enhancement of the prompt for new tasks by harnessing correlations between different tasks in CIL. Experiments on six benchmarks demonstrate that our method achieves state-of-the-art performance.
Yongchun Qin, Pengfei Fang, Hui Xue 0002
AAAI3
2025 KAnoCLIP: Zero-Shot Anomaly Detection through Knowledge-Driven Prompt Learning and Enhanced Cross-Modal Integration
abstract
Zero-shot anomaly detection (ZSAD) identifies anomalies without needing training samples from the target dataset, essential for scenarios with privacy concerns or limited data. Vision-language models like CLIP show potential in ZSAD but have limitations: relying on manually crafted fixed textual descriptions or anomaly prompts is time-consuming and prone to semantic ambiguity, and CLIP struggles with pixel-level anomaly segmentation, focusing more on global semantics than local details. To address these limitations, We introduce KAnoCLIP, a novel ZSAD framework that leverages vision-language models. KAnoCLIP combines general knowledge from a Large Language Model (GPT-3.5) and fine-grained, image-specific knowledge from a Visual Question Answering system (Llama3) via Knowledge-Driven Prompt Learning (KnPL). KnPL uses a knowledge-driven (KD) loss function to create learnable anomaly prompts, removing the need for fixed text prompts and enhancing generalization. KAnoCLIP includes the CLIP visual encoder with V-V attention (CLIP-VV), Bi-Directional Cross-Attention for Multi-Level Cross-Modal Interaction (Bi-CMCI), and Conv-Adapter. These components preserve local visual semantics, improve local cross-modal fusion, and align global visual features with textual information, enhancing pixel-level anomaly detection. KAnoCLIP achieves state-of-the-art performance in ZSAD across 12 industrial and medical datasets, demonstrating superior generalization compared to existing methods.
Suyang Zhou, Jieping Kong, Lei Qi 0001, Hui Xue 0002
ICASSP5
2025 Precision in Pathology: PMA-DETR Elevates Tumor Lesion Detection
abstract
The DETR series, known for its end-to-end object detection models, has gained significant attention for its performance. RT-DETR excels with higher accuracy and faster real-time inference. However, applying these models to medical imaging poses challenges, such as low-contrast and complex lesion structures, which can reduce effectiveness. When detecting tumors, models may overfit due to the distinct differences and variability between different cases, affecting generalization and accuracy. To address these challenges, we propose a multi-view parallel feature extraction module, specifically for tumor detection. This module includes adaptive preprocessing, joint axial and channel attention, multi-pooling angular attention to enhance relevant features and reduce redundancy. Additionally, axial dynamic deformable convolution is used to improve adaptability and robustness. The resulting PMA-DETR architecture achieves state-of-the-art tumor detection while maintaining real-time processing.
Yecheng Zhao, Lei Qi 0001, Hui Xue 0002
ICASSP4
2025 DePro: Domain Ensemble using Decoupled Prompts for Universal Cross-Domain Retrieval
abstract
This paper investigates the potential of vision-language models (VLMs) in addressing the challenges of universal cross-domain retrieval (UCDR), where queries originate from unseen domains or classes. A common approach to adapting VLMs for downstream tasks involves prompt tuning, which alleviates the computational burden of full fine-tuning. However, this approach often struggles with the domain and semantic shifts inherent in UCDR. To overcome these limitations, we propose a novel prompt decoupling strategy that separates prompts into universal domain prompts (UDPs) and class prompts (CPs). Specifically, UDPs are designed to unify features from both seen and unseen domains into a cohesive universal domain, while CPs are tailored to capture class-specific visual characteristics, enabling robust retrieval across both known and unknown classes. To ensure effective decoupling, we introduce a dedicated decoupling loss that enforces the domain-agnostic nature of CPs. Additionally, we employ a regulation loss to align features from the frozen CLIP domain with those of the universal domain by selectively integrating or excluding UDPs. This mechanism fosters a synergistic domain ensemble effect, enhancing retrieval generalization across diverse domains. Finally, we propose the domain-aware triplet-hard (DaTri) loss to mitigate overfitting by reducing the risk of class collapse. The proposed framework, referred to as Domain Ensemble using Decoupled Prompts (DePro), demonstrates state-of-the-art performance and effectively enhances the model's generalization capacity across unseen domains and classes, as validated through extensive experiments. Code is here.
Kaixiang Chen, Pengfei Fang, Hui Xue 0002
SIGIR3
2025 OAFN: An efficient open-world audio few-shot learning network for event classification
Fei Chen 0015, Hui Xue 0002, Pengfei Fang
Knowl. Based Syst.2
2025 Toward Few-Shot Learning in the Open World: A Review and Beyond
abstract
Human intelligence is characterized by our ability to absorb and apply knowledge from the world around us, especially in rapidly acquiring new concepts from minimal examples, underpinned by prior knowledge. Few-shot learning (FSL) aims to mimic this capacity by enabling significant generalizations and transferability. However, traditional FSL frameworks often rely on assumptions of clean, complete, and static data, conditions that are seldom met in real-world environments. Such assumptions falter in the inherently uncertain, incomplete, and dynamic contexts of the open world. This paper presents a comprehensive review of recent advancements designed to adapt FSL to open-world environments. We categorize existing methods into three distinct types of FSL in the open world: those involving varying instances, varying classes, and varying distributions. Each category is discussed in terms of its specific challenges and methods, as well as its strengths and weaknesses. We standardize experimental settings and metric benchmarks across scenarios and provide a comparative analysis of the performance of various methods. In conclusion, we outline potential future research directions for this evolving field. It is our hope that this review will catalyze further development of effective solutions to these complex challenges, thereby advancing the field of artificial intelligence.
Hui Xue 0002, Yuexuan An, Yongchun Qin, Wenqian Li 0005, Yixin Wu 0004, Yongjuan Che, Pengfei Fang, Min-Ling Zhang
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 On inferring prototypes for multi-label few-shot learning via partial aggregation
Pengfei Fang, Hui Xue 0002
Pattern Recognit.3
2025 Learning Noisy Few-Shot Classification Without Relying on Pseudo-Noise Data
abstract
Recently, noisy few-shot learning (NFSL) has been exploring the model robustness to label noise, breaking the limitation of completely accurate labeling in small sample scenarios. Existing NFSL methods directly employ delicately designed pseudo-noise to simulate and adapt to noisy environments. However, determining the optimal combination of pseudo-noise is challenging and improperly configuring pseudo-noise may lead to adverse effects on the training models. To deal with the problems, this letter proposes a novelAdaptiveMultI-viewDenoisingEvaluation (AMIDE) framework, which establishes an adaptive and robust embedding and classifier without relying on pseudo-noise. In the training phase, we design an adaptive label smoothing scheme, where soft labels with learnable smooth coefficients are inferred from data distribution to mitigate overconfident labeling. In the testing stage, we propose a multi-view fused evaluation scheme, where different network layers are treated as distinct views to generate potential clean features and modify prototypes, thereby enhancing the accuracy of evaluation. In this way, the impact of noise is effectively alleviated from two perspectives. Extensive experiments on several few-shot classification benchmarks show the superiority and robustness of our method.
Yixin Wu 0004, Hui Xue 0002, Yuexuan An, Pengfei Fang
IEEE Signal Process. Lett.2
2025 On Modulating Motion-Aware Visual-Language Representation for Few-Shot Action Recognition
abstract
This paper focuses on few-shot action recognition (FSAR), where the machine is required to understand human actions, with each only seeing a few video samples. Even with only a few explorations, the most cutting-edge methods employ the action textual features, pre-trained by a visual-language model (VLM), as a cue to optimize video prototypes. However, the action textual features used in these methods are generated from a static prompt, causing the network to overlook rich motion cues within videos. To tackle this issue, we propose a novel framework, namely, motion-aware visual-language representation modulation network (MoveNet). The proposed MoveNet utilizes dynamic motion cues within videos to integrate motion-aware textual and visual feature representations, as a way to modulate the video prototypes. In doing so, a long short motion aggregation module (LSMAM) is first proposed to capture diverse motion cues. Having the motion cues at hand, a motion-conditional prompting module (MCPM) utilizes the motion cues as conditions to boost the semantic associations between textual features and action classes. One further develops a motion-guided visual refinement module (MVRM) that adopts motion cues as guidance in enhancing local frame features. The proposed components compensate for each other and contribute to significant performance gains over the FASR task. Thorough experiments on five standard benchmarks demonstrate the effectiveness of the proposed method, considerably outperforming current state-of-the-art methods.
Pengfei Fang, Hui Xue 0002
IEEE Trans. Circuits Syst. Video Technol.4
2025 Leveraging Bilateral Correlations for Multi-Label Few-Shot Learning
abstract
Multi-label few-shot learning (ML-FSL) refers to the task of tagging previously unseen images with a set of relevant labels, giving a small number of training examples. Modeling the correlations between instances and labels, formulated in the existing methods, allows us to extract more available knowledge from limited examples. However, they simply explore the instance and label correlations with a uniform importance assumption without considering the discrepancy of importance in different instances or labels, making the utilization of instance and label correlations a bottleneck for ML-FSL. To tackle the issue, we propose a unified framework named bilateral correlation reconstruction (BCR) to enable the network to effectively mine underlying instance and label correlations with varying importance information from both instance-to-label and label-to-instance perspectives. Specifically, from the instance-to-label perspective, we refine prototypes per category by reweighting each image with its specific instance-importance degree extracted from the similarity between the instance and the corresponding category. From the label-to-instance perspective, we smooth labels for each image by recovering latent label-importance with considering the integrated topology of all samples in a task. Experimental results on multiple benchmarks validate that BCR could outperform existing ML-FSL methods by large margins.
Yuexuan An, Hui Xue 0002, Xingyu Zhao 0002, Ning Xu 0009, Pengfei Fang, Xin Geng 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Text Image Inpainting via Global Structure-Guided Diffusion Models
abstract
Real-world text can be damaged by corrosion issues caused by environmental or human factors, which hinder the preservation of the complete styles of texts, e.g., texture and structure. These corrosion issues, such as graffiti signs and incomplete signatures, bring difficulties in understanding the texts, thereby posing significant challenges to downstream applications, e.g., scene text recognition and signature identification. Notably, current inpainting techniques often fail to adequately address this problem and have difficulties restoring accurate text images along with reasonable and consistent styles. Formulating this as an open problem of text image inpainting, this paper aims to build a benchmark to facilitate its study. In doing so, we establish two specific text inpainting datasets which contain scene text images and handwritten text images, respectively. Each of them includes images revamped by real-life and synthetic datasets, featuring pairs of original images, corrupted images, and other assistant information. On top of the datasets, we further develop a novel neural framework, Global Structure-guided Diffusion Model (GSDM), as a potential solution. Leveraging the global structure of the text as a prior, the proposed GSDM develops an efficient diffusion model to recover clean texts. The efficacy of our approach is demonstrated by thorough empirical study, including a substantial boost in both recognition accuracy and image quality. These findings not only highlight the effectiveness of our method but also underscore its potential to enhance the broader field of text image understanding and processing. Code and datasets are available at: https://github.com/blackprotoss/GSDM.
Shipeng Zhu, Pengfei Fang, Chenjie Zhu, Zuoyan Zhao, Hui Xue 0002
AAAI6
2024 TSDNet: An Efficient Light Footprint Keyword Spotting Deep Network Base on Tempera Segment Normalization
abstract
Keyword spotting plays a pivotal role in smart devices equipped with speech services, serving as the hands-free interface for the activation of pre-specified tasks. It brings the challenge of maintaining high precision with minimal parameters becomes particularly pronounced in devices constrained by limited hardware resources. In this paper, we delve into the research of Temporal Segment Normalization (TSN) and introduce a highly scalable TSN-based ResNet, termed TSDNet, tailored for deployment on target smart devices. Our approach involves meticulous experimental design, wherein we systematically configure experiments to train distinct scaled models. This not only allows us to validate the effectiveness of TSN but also assesses the performance of the proposed network under varying scales. Experimental results, conducted on the Google Speech Commands dataset v1 and v2, demonstrate that the small, medium, and large-scale models consistently attain stateof-the-art results compared to similar methods. Notably, the smallest model achieves a remarkable 97% accuracy on dataset v1 and 97.1% on dataset v2, while operating with a minimum of 7.4K Bytes of parameters. Furthermore, we undertake retraining of ResNet using our proposed training method, achieving new accuracy margins.
Fei Chen 0015, Hui Xue 0002, Pengfei Fang
CSCWD2
2024 Copula-Nested Spectral Kernel Network
abstract
Spectral Kernel Networks (SKNs) emerge as a promising approach in machine learning, melding solid theoretical foundations of spectral kernels with the representation power of hierarchical architectures. At its core, the spectral density function plays a pivotal role by revealing essential patterns in data distributions, thereby offering deep insights into the underlying framework in real-world tasks. Nevertheless, prevailing designs of spectral density often overlook the intricate interactions within data structures. This phenomenon consequently neglects expanses of the hypothesis space, thus curtailing the performance of SKNs. This paper addresses the issues through a novel approach, the **Co**pula-Nested Spectral **Ke**rnel **Net**work (**CokeNet**). Concretely, we first redefine the spectral density with the form of copulas to enhance the diversity of spectral densities. Next, the specific expression of the copula module is designed to allow the excavation of complex dependence structures. Finally, the unified kernel network is proposed by integrating the corresponding spectral kernel and the copula module. Through rigorous theoretical analysis and experimental verification, CokeNet demonstrates superior performance and significant advancements over SOTA algorithms in the field.
Jinyue Tian, Hui Xue 0002, Yanfang Xue, Pengfei Fang
ICML2
2024 PEAN: A Diffusion-Based Prior-Enhanced Attention Network for Scene Text Image Super-Resolution
abstract
Scene text image super-resolution (STISR) aims at simultaneously increasing the resolution and readability of low-resolution scene text images, thus boosting the performance of the downstream recognition task. Two factors in scene text images, visual structure and semantic information, affect the recognition performance significantly. To mitigate the effects from these factors, this paper proposes a Prior-Enhanced Attention Network (PEAN). Specifically, an attention-based modulation module is leveraged to understand scene text images by neatly perceiving the local and global dependence of images, despite the shape of the text. Meanwhile, a diffusion-based module is developed to enhance the text prior, hence offering better guidance for the SR network to generate SR images with higher semantic accuracy. Additionally, a multi-task learning paradigm is employed to optimize the network, enabling the model to generate legible SR images. As a result, PEAN establishes new SOTA results on the TextZoom benchmark. Experiments are also conducted to analyze the importance of the enhanced text prior as a means of improving the performance of the SR network. Code is available at https://github.com/jdfxzzy/PEAN.
Zuoyan Zhao, Hui Xue 0002, Pengfei Fang, Shipeng Zhu
ACM Multimedia2
2024 Reproducing the Past: A Dataset for Benchmarking Inscription Restoration
abstract
Inscriptions on ancient steles, as carriers of culture, encapsulate the humanistic thoughts and aesthetic values of our ancestors. However, these relics often deteriorate due to environmental and human factors, resulting in significant information loss. Since the advent of inscription rubbing technology over a millennium ago, archaeologists and epigraphers have devoted immense effort to manually restoring these cultural imprints, endeavoring to unlock the storied past within each rubbing. This paper approaches this challenge as a multi-modal task, aiming to establish a novel benchmark for the inscription restoration from rubbings. In doing so, we construct the Chinese Inscription Rubbing Image (CIRI) dataset, which includes a wide variety of real inscription rubbing images characterized by diverse calligraphy styles, intricate character structures, and complex degradation forms. Furthermore, we develop a synthesis approach to generate "intact-degraded'' paired data, mirroring real-world degradation faithfully. On top of the datasets, we propose a baseline framework that achieves visual consistency and textual integrity through global and local diffusion-based restoration processes and explicit incorporation of domain knowledge. Comprehensive evaluations confirm the effectiveness of our pipeline, demonstrating significant improvements in visual presentation and textual integrity. The project is available at: https://github.com/blackprotoss/CIRI.
Shipeng Zhu, Hui Xue 0002, Na Nie, Chenjie Zhu, Haiyue Liu, Pengfei Fang
ACM Multimedia2
2024 Dynamic functional connections analysis with spectral learning for brain disorder detection
Yanfang Xue, Hui Xue 0002, Pengfei Fang, Shipeng Zhu, Lishan Qiao, Yuexuan An
Artif. Intell. Medicine2
2024 Towards kernelizing the classifier for hyperbolic data
Meimei Yang, Xinkai Sun, Na Shi, Hui Xue 0002
Frontiers Comput. Sci.5
2024 Reinforced Self-Supervised Training for Few-Shot Learning
abstract
Few-shot learning is an open problem to learning a new concept with little supervision from limited labeled data. As an alternative knowledge for few-shot learning, self-supervised learning can extract supervisory signals directly from unlabeled data. However, existing self-supervised few-shot methods which directly take the summation of two tasks, have two fundamental bottlenecks: 1) representation bias: how to extract efficacious supervisory signals in self-supervision and eliminate the disturbance of undesirable shortcuts with limited examples, 2) objective conflict: how to adaptively trade-off the self-supervision and supervision to achieve the optimal model performance. To address the above problems, in this paper, we propose a novel approach named ReInforced SElf-supervised training (RISE) for few-shot learning. RISE leverages agent-relative supervision to eliminate the undesirable shortcut learning of self-supervised training. Meanwhile, it dynamically explores the balance between supervisory signals from self-supervised tasks and inherent supervision from few-shot tasks to avoid the trade-off dilemma. Therefore, the new pattern for training self-supervision can be resilient to few-shot learning and enhance the performance for few-shot identification. Extensive experiments on several public benchmark datasets verify the effectiveness of our approach.
Zhichao Yan 0003, Yuexuan An, Hui Xue 0002
IEEE Signal Process. Lett.3
2024 DoubleAUG: Single-domain Generalized Object Detector in Urban via Color Perturbation and Dual-style Memory
abstract
Object detection in urban scenarios is crucial for autonomous driving in intelligent traffic systems. However, unlike conventional object detection tasks, urban-scene images vary greatly in style. For example, images taken on sunny days differ significantly from those taken on rainy days. Therefore, models trained on sunny-day images may not generalize well to rainy-day images. In this article, we aim to solve the single-domain generalizable object detection task in urban scenarios, meaning that a model trained on images from one weather condition should be able to perform well on images from any other weather conditions. To address this challenge, we propose a novel Double AUGmentation (DoubleAUG) method that includes image- and feature-level augmentation schemes. In the image-level augmentation, we consider the variation in color information across different weather conditions and propose a Color Perturbation (CP) method that randomly exchanges the RGB channels to generate various images. In the feature-level augmentation, we propose to utilize a Dual-Style Memory (DSM) to explore the diverse style information on the entire dataset, further enhancing the model’s generalization capability. Extensive experiments demonstrate that our proposed method outperforms state-of-the-art methods. Furthermore, ablation studies confirm the effectiveness of each module in our proposed method. Moreover, our method is plug-and-play and can be integrated into existing methods to further improve model performance.
Lei Qi 0001, Tan Xiong, Hui Xue 0002, Xin Geng 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Improving Scene Text Retrieval via Stylized Middle Modality
abstract
Scene text retrieval addresses the challenge of localizing and searching for all text instances within scene images based on a query text. This cross-modal task has significant applications in various domains, such as intelligent transportation systems and social media analysis. In practice, ensuring consistency of the same content between two modalities is crucial in improving retrieval accuracy. This article addresses the issue by introducing a stylized middle modality, which fuses the graphical query text with the style of the extracted text proposal. To this end, we propose a stylized middle modality learning (SM 2 L) framework. The proposed stylized middle modality enables the network to jointly enforce constraints on visual feature coherence and text semantic feature consistency in the optimization phase, thereby minimizing the modality gap in the retrieval space. This brings in two major advantages: (1) SM 2 L will pave the way to seamlessly benefit the scene text retrieval and (2) the proposed learning paradigm enables the machine to avoid adding redundant computing resources in the inference phase. Substantial experiments demonstrate that the proposed method outperforms the state-of-the-art retrieval performance considerably.
Shipeng Zhu, Pengfei Fang, Hui Xue 0002
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Improving Scene Text Image Super-resolution via Dual Prior Modulation Network
abstract
Scene text image super-resolution (STISR) aims to simultaneously increase the resolution and legibility of the text images, and the resulting images will significantly affect the performance of downstream tasks. Although numerous progress has been made, existing approaches raise two crucial issues: (1) They neglect the global structure of the text, which bounds the semantic determinism of the scene text. (2) The priors, e.g., text prior or stroke prior, employed in existing works, are extracted from pre-trained text recognizers. That said, such priors suffer from the domain gap including low resolution and blurriness caused by poor imaging conditions, leading to incorrect guidance. Our work addresses these gaps and proposes a plug-and-play module dubbed Dual Prior Modulation Network (DPMN), which leverages dual image-level priors to bring performance gain over existing approaches. Specifically, two types of prior-guided refinement modules, each using the text mask or graphic recognition result of the low-quality SR image from the preceding layer, are designed to improve the structural clarity and semantic accuracy of the text, respectively. The following attention mechanism hence modulates two quality-enhanced images to attain a superior SR result. Extensive experiments validate that our method improves the image quality and boosts the performance of downstream tasks over five typical approaches on the benchmark. Substantial visualizations and ablation studies demonstrate the advantages of the proposed DPMN. Code is available at: https://github.com/jdfxzzy/DPMN.
Shipeng Zhu, Zuoyan Zhao, Pengfei Fang, Hui Xue 0002
AAAI4
2023 Learning to Learn from Corrupted Data for Few-Shot Learning
abstract
Few-shot learning which aims to generalize knowledge learned from annotated base training data to recognize unseen novel classes has attracted considerable attention. Existing few-shot methods rely on completely clean training data. However, in the real world, the training data are always corrupted and accompanied by noise due to the disturbance in data transmission and low-quality annotation, which severely degrades the performance and generalization capability of few-shot models. To address the problem, we propose a unified peer-collaboration learning (PCL) framework to extract valid knowledge from corrupted data for few-shot learning. PCL leverages two modules to mimic the peer collaboration process which cooperatively evaluates the importance of each sample. Specifically, each module first estimates the importance weights of different samples by encoding the information provided by the other module from both global and local perspectives. Then, both modules leverage the obtained importance weights to guide the reevaluation of the loss value of each sample. In this way, the peers can mutually absorb knowledge to improve the robustness of few-shot models. Experiments verify that our framework combined with different few-shot methods can significantly improve the performance and robustness of original models.
Yuexuan An, Xingyu Zhao 0002, Hui Xue 0002
IJCAI3
2023 Boosting Few-Shot Open-Set Recognition with Multi-Relation Margin Loss
abstract
Few-shot open-set recognition (FSOSR) has become a great challenge, which requires classifying known classes and rejecting the unknown ones with only limited samples. Existing FSOSR methods mainly construct an ambiguous distribution of known classes from scarce known samples without considering the latent distribution information of unknowns, which degrades the performance of open-set recognition. To address this issue, we propose a novel loss function called multi-relation margin (MRM) loss that can plug in few-shot methods to boost the performance of FSOSR. MRM enlarges the margin between different classes by extracting the multi-relationship of paired samples to dynamically refine the decision boundary for known classes and implicitly delineate the distribution of unknowns. Specifically, MRM separates the classes by enforcing a margin while concentrating samples of the same class on a hypersphere with a learnable radius. In order to better capture the distribution information of each class, MRM extracts the similarity and correlations among paired samples, ameliorating the optimization of the margin and radius. Experiments on public benchmarks reveal that methods with MRM loss can improve the unknown detection of AUROC by a significant margin while correctly classifying the known classes.
Yongjuan Che, Yuexuan An, Hui Xue 0002
IJCAI3
2023 Expanding the Hyperbolic Kernels: A Curvature-aware Isometric Embedding View
abstract
Modeling data relation as a hierarchical structure has proven beneficial for many learning scenarios, and the hyperbolic space, with negative curvature, can encode such data hierarchy without distortion. Several recent studies also show that the representation power of the hyperbolic space can be further improved by endowing the kernel methods. Unfortunately, the known kernel methods, developed in hyperbolic space, are limited by the adaptation capacity or distortion issues. This paper addresses the issues through a novel embedding function. To this end, we propose a curvature-aware isometric embedding, which establishes an isometry from the Poincar\'e model to a special reproducing kernel Hilbert space (RKHS). Then we can further define a series of kernels on this RKHS, including several positive definite kernels and an indefinite kernel. Thorough experiments are conducted to demonstrate the superiority of our proposals over existing-known hyperbolic and Euclidean kernels in various learning tasks, e.g., graph learning and zero-shot learning.
Meimei Yang, Pengfei Fang, Hui Xue 0002
IJCAI3
2023 CosNet: A Generalized Spectral Kernel Network
abstract
Complex-valued representation exists inherently in the time-sequential data that can be derived from the integration of harmonic waves. The non-stationary spectral kernel, realizing a complex-valued feature mapping, has shown its potential to analyze the time-varying statistical characteristics of the time-sequential data, as a result of the modeling frequency parameters. However, most existing spectral kernel-based methods eliminate the imaginary part, thereby limiting the representation power of the spectral kernel. To tackle this issue, we propose a generalized spectral kernel network, namely, \underline{Co}mplex-valued \underline{s}pectral kernel \underline{Net}work (CosNet), which includes spectral kernel mapping generalization (SKMG) module and complex-valued spectral kernel embedding (CSKE) module. Concretely, the SKMG module is devised to generalize the spectral kernel mapping in the real number domain to the complex number domain, recovering the inherent complex-valued representation for the real-valued data. Then a following CSKE module is further developed to combine the complex-valued spectral kernels and neural networks to effectively capture long-range or periodic relations of the data. Along with the CosNet, we study the effect of the complex-valued spectral kernel mapping via theoretically analyzing the bound of covering number and generalization error. Extensive experiments demonstrate that CosNet performs better than the mainstream kernel methods and complex-valued neural networks.
Yanfang Xue, Pengfei Fang, Jinyue Tian, Shipeng Zhu, Hui Xue 0002
NeurIPS5
2023 On learning distribution alignment for video-based visible-infrared person re-identification
Pengfei Fang, Yaojun Hu, Shipeng Zhu, Hui Xue 0002
Comput. Vis. Image Underst.4
2023 FATE: a three-stage method for arithmetical exercise correction
Qipeng Zhu, Zhuoyan Luo, Shipeng Zhu, Zihang Xu, Hui Xue 0002
Neural Comput. Appl.6
2023 From Instance to Metric Calibration: A Unified Framework for Open-World Few-Shot Learning
abstract
Robust few-shot learning (RFSL), which aims to address noisy labels in few-shot learning, has recently gained considerable attention. Existing RFSL methods are based on the assumption that the noise comes from known classes (in-domain), which is inconsistent with many real-world scenarios where the noise does not belong to any known classes (out-of-domain). We refer to this more complex scenario as open-world few-shot learning (OFSL), where in-domain and out-of-domain noise simultaneously exists in few-shot datasets. To address the challenging problem, we propose a unified framework to implement comprehensive calibration from instance to metric. Specifically, we design a dual-networks structure composed of a contrastive network and a meta network to respectively extract feature-related intra-class information and enlarged inter-class variations. For instance-wise calibration, we present a novel prototype modification strategy to aggregate prototypes with intra-class and inter-class instance reweighting. For metric-wise calibration, we present a novel metric to implicitly scale the per-class prediction by fusing two spatial metrics respectively constructed by the two networks. In this way, the impact of noise in OFSL can be effectively mitigated from both feature space and label space. Extensive experiments on various OFSL settings demonstrate the robustness and superiority of our method. Our source codes is available at https://github.com/anyuexuan/IDEAL.
Yuexuan An, Hui Xue 0002, Xingyu Zhao 0002, Jing Wang 0113
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Automatically Gating Multi-Frequency Patterns through Rectified Continuous Bernoulli Units with Theoretical Principles
abstract
Different nonlinearities are only suitable for responding to different frequency signals. The locally-responding ReLU is incapable of modeling high-frequency features due to the spectral bias, whereas the globally-responding sinusoidal function is intractable to represent low-frequency concepts cheaply owing to the optimization dilemma. Moreover, nearly all the practical tasks are composed of complex multi-frequency patterns, whereas there is little prospect of designing or searching a heterogeneous network containing various types of neurons matching the frequencies, because of their exponentially-increasing combinatorial states. In this paper, our contributions are three-fold: 1) we propose a general Rectified Continuous Bernoulli (ReCB) unit paired with an efficient variational Bayesian learning paradigm, to automatically detect/gate/represent different frequency responses; 2) our numerically-tight theoretical framework proves that ReCB-based networks can achieve the optimal representation ability, which is O(m^{η/(d^2)}) times better than that of popular neural networks, for a hidden dimension of m, an input dimension of d, and a Lipschitz constant of η; 3) we provide comprehensive empirical evidence showing that ReCB-based networks can keenly learn multi-frequency patterns and push the state-of-the-art performance.
Zheng-Fan Wu, Yi-Nan Feng, Hui Xue 0002
IJCAI3
2022 Self-corrected unsupervised domain adaptation
Hui Xue 0002, Songcan Chen
Frontiers Comput. Sci.3
2022 Re-Weighting Large Margin Label Distribution Learning for Classification
abstract
Label ambiguity has attracted quite some attention among the machine learning community. The latterly proposed Label Distribution Learning (LDL) can handle label ambiguity and has found wide applications in real classification problems. In the training phase, an LDL model is learned first. In the test phase, the top label(s) in the label distribution predicted by the learned LDL model is (are) then regarded as the predicted label(s). That is, LDL considers the whole label distribution in the training phase, but only the top label(s) in the test phase, which likely leads to objective inconsistency. To avoid such inconsistency, we propose a new LDL method Re-Weighting Large Margin Label Distribution Learning (RWLM-LDL). First, we prove that the expected$ L_1 $-norm loss of LDL bounds the classification error probability, and thus apply$ L_1 $-norm loss as the learning metric. Second, re-weighting schemes are put forward to alleviate the inconsistency. Third, large margin is introduced to further solve the inconsistency. The theoretical results are presented to showcase the generalization and discrimination of RWLM-LDL. Finally, experimental results show the statistically superior performance of RWLM-LDL against other comparing methods.
Jing Wang 0113, Xin Geng 0001, Hui Xue 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Indefinite twin support vector machine with DC functions programming
Yuexuan An, Hui Xue 0002
Pattern Recognit.2
2021 Conditional Self-Supervised Learning for Few-Shot Classification
abstract
How to learn a transferable feature representation from limited examples is a key challenge for few-shot classification. Self-supervision as an auxiliary task to the main supervised few-shot task is considered to be a conceivable way to solve the problem since self-supervision can provide additional structural information easily ignored by the main task. However, learning a good representation by traditional self-supervised methods is usually dependent on large training samples. In few-shot scenarios, due to the lack of sufficient samples, these self-supervised methods might learn a biased representation, which more likely leads to the wrong guidance for the main tasks and finally causes the performance degradation. In this paper, we propose conditional self-supervised learning (CSS) to use auxiliary information to guide the representation learning of self-supervised tasks. Specifically, CSS leverages supervised information as prior knowledge to shape and improve the learning feature manifold of self-supervision without auxiliary unlabeled data, so as to reduce representation bias and mine more effective semantic information. Moreover, CSS exploits more meaningful information through supervised and the improved self-supervised learning respectively and integrates the information into a unified distribution, which can further enrich and broaden the original representation. Extensive experiments demonstrate that our proposed method without any fine-tuning can achieve a significant accuracy improvement on the few-shot classification scenarios compared to the state-of-the-art few-shot learning methods.
Yuexuan An, Hui Xue 0002, Xingyu Zhao 0002
IJCAI2
2021 Adversarial Spectral Kernel Matching for Unsupervised Time Series Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) has been received increasing attention since it does not require labels in target domain. Most existing UDA methods learn domain-invariant features by minimizing discrepancy distance computed by a certain metric between domains. However, these discrepancy-based methods cannot be robustly applied to unsupervised time series domain adaptation (UTSDA). That is because discrepancy metrics in these methods contain only low-order and local statistics, which have limited expression for time series distributions and therefore result in failure of domain matching. Actually, the real-world time series are always non-local distributions, i.e., with non-stationary and non-monotonic statistics. In this paper, we propose an Adversarial Spectral Kernel Matching (AdvSKM) method, where a hybrid spectral kernel network is specifically designed as inner kernel to reform the Maximum Mean Discrepancy (MMD) metric for UTSDA. The hybrid spectral kernel network can precisely characterize non-stationary and non-monotonic statistics in time series distributions. Embedding hybrid spectral kernel network to MMD not only guarantees precise discrepancy metric but also benefits domain matching. Besides, the differentiable architecture of the spectral kernel network enables adversarial kernel learning, which brings more discriminatory expression for discrepancy matching. The results of extensive experiments on several real-world UTSDA tasks verify the effectiveness of our proposed method.
Hui Xue 0002
IJCAI2
2021 Learning Deeper Non-Monotonic Networks by Softly Transferring Solution Space
abstract
Different from popular neural networks using quasiconvex activations, non-monotonic networks activated by periodic nonlinearities have emerged as a more competitive paradigm, offering revolutionary benefits: 1) compactly characterizing high-frequency patterns; 2) precisely representing high-order derivatives. Nevertheless, they are also well-known for being hard to train, due to easily over-fitting dissonant noise and only allowing for tiny architectures (shallower than 5 layers). The fundamental bottleneck is that the periodicity leads to many poor and dense local minima in solution space. The direction and norm of gradient oscillate continually during error backpropagation. Thus non-monotonic networks are prematurely stuck in these local minima, and leave out effective error feedback. To alleviate the optimization dilemma, in this paper, we propose a non-trivial soft transfer approach. It smooths their solution space close to that of monotonic ones in the beginning, and then improve their representational properties by transferring the solutions from the neural space of monotonic neurons to the Fourier space of non-monotonic neurons as the training continues. The soft transfer consists of two core components: 1) a rectified concrete gate is constructed to characterize the state of each neuron; 2) a variational Bayesian learning framework is proposed to dynamically balance the empirical risk and the intensity of transfer. We provide comprehensive empirical evidence showing that the soft transfer not only reduces the risk of non-monotonic networks on over-fitting noise, but also helps them scale to much deeper architectures (more than 100 layers) achieving the new state-of-the-art performance.
Zheng-Fan Wu, Hui Xue 0002, Weimin Bai
IJCAI2
2021 Pointwise manifold regularization for semi-supervised learning
Yating Shen, Hui Xue 0002
Frontiers Comput. Sci.4
2021 Source-Free Unsupervised Domain Adaptation with Sample Transport Learning
Qing Tian 0002, Feng-Yuan Zhang, Shun Peng, Hui Xue 0002
J. Comput. Sci. Technol.5
2021 Learning graph-level representation from local-structural distribution with Graph Neural Networks
Wei-Xiang Sun, Hui Xue 0002
Knowl. Based Syst.2
2020 BaKer-Nets: Bayesian Random Kernel Mapping Networks
abstract
Recently, deep spectral kernel networks (DSKNs) have attracted wide attention. They consist of periodic computational elements that can be activated across the whole feature spaces. In theory, DSKNs have the potential to reveal input-dependent and long-range characteristics, and thus are expected to perform more competitive than prevailing networks. But in practice, they are still unable to achieve the desired effects. The structural superiority of DSKNs comes at the cost of the difficult optimization. The periodicity of computational elements leads to many poor and dense local minima in loss landscapes. DSKNs are more likely stuck in these local minima, and perform worse than expected. Hence, in this paper, we propose the novel Bayesian random Kernel mapping Networks (BaKer-Nets) with preferable learning processes by escaping randomly from most local minima. Specifically, BaKer-Nets consist of two core components: 1) a prior-posterior bridge is derived to enable the uncertainty of computational elements reasonably; 2) a Bayesian learning paradigm is presented to optimize the prior-posterior bridge efficiently. With the well-tuned uncertainty, BaKer-Nets can not only explore more potential solutions to avoid local minima, but also exploit these ensemble solutions to strengthen their robustness. Systematical experiments demonstrate the significance of BaKer-Nets in improving learning processes on the premise of preserving the structural superiority.
Hui Xue 0002, Zheng-Fan Wu
IJCAI1
2020 Sketch discriminatively regularized online gradient descent classification
Hui Xue 0002
Appl. Intell.1
2020 Non-convex approximation based l0-norm multiple indefinite kernel feature selection
Hui Xue 0002
Appl. Intell.1
2020 A primal perspective for indefinite kernel SVM problem
Hui Xue 0002, Xiaohong Chen 0001
Frontiers Comput. Sci.1
2020 Discrimination-Aware Domain Adversarial Neural Network
Jian-Min Gu, Songcan Chen, Hui Xue 0002
J. Comput. Sci. Technol.5
2020 Multiple indefinite kernel learning for feature selection
Hui Xue 0002
Knowl. Based Syst.1
2019 Adaptive Teacher-and-Student Model for Heterogeneous Domain Adaptation
abstract
In heterogeneous domain adaptation (HDA), since the feature spaces of the source and target domains are different, knowledge transfer from the source to the target domain is really challenging. How to align the different feature spaces and then adaptively transfer the related knowledge is critical for HDA. In this paper, we develop an adaptive teacher-and-student model for heterogeneous domain adaptation (AtsHDA). In AtsHDA, the source domain as a teacher and the target domain as a student are aligned or co-adapted to each other first, so that their correlation can be maximized. Then the target domain adaptively learns from the source domain. Specifically, there is a balance between the learning by the target domain itself and the instruction from the source domain. That is, when the guidance from the source domain is helpful for learning, the learning of target classifier emphasizes the instruction of source knowledge, and considers its own knowledge more, otherwise. Further, an ensemble method is designed to decide such a balance. Finally, empirical results show that AtsHDA can achieve competitive results compared with the state-of-arts.
Xuzhang Chen, Songcan Chen, Hui Xue 0002
ICDM4
2019 Deep Spectral Kernel Learning
abstract
Recently, spectral kernels have attracted wide attention in complex dynamic environments. These advanced kernels mainly focus on breaking through the crucial limitation on locality, that is, the stationarity and the monotonicity. But actually, owing to the inefficiency of shallow models in computational elements, they are more likely unable to accurately reveal dynamic and potential variations. In this paper, we propose a novel deep spectral kernel network (DSKN) to naturally integrate non-stationary and non-monotonic spectral kernels into elegant deep architectures in an interpretable way, which can be further generalized to cover most kernels. Concretely, we firstly deal with the general form of spectral kernels by the inverse Fourier transform. Secondly, DSKN is constructed by embedding the preeminent spectral kernels into each layer to boost the efficiency in computational elements, which can effectively reveal the dynamic input-dependent characteristics and potential long-range correlations by compactly representing complex advanced concepts. Thirdly, detailed analyses of DSKN are presented. Owing to its universality, we propose a unified spectral transform technique to flexibly extend and reasonably initialize domain-related DSKN. Furthermore, the representer theorem of DSKN is given. Systematical experiments demonstrate the superiority of DSKN compared to state-of-the-art relevant algorithms on varieties of standard real-world tasks.
Hui Xue 0002, Zheng-Fan Wu, Wei-Xiang Sun
IJCAI1
2019 A maximum margin clustering algorithm based on indefinite kernels
Hui Xue 0002, Xiaohong Chen 0001
Frontiers Comput. Sci.1
2019 A Primal Framework for Indefinite Kernel Learning
Hui Xue 0002, Songcan Chen
Neural Process. Lett.1
2018 Subclass Maximum Margin Tree Error Correcting Output Codes
Fa Zheng, Hui Xue 0002
PRICAI (1)2
2018 Transfer learning with partial related "instance-feature" knowledge
Jie Zhai, Yun Li 0009, Ke-Jia Chen 0001, Hui Xue 0002
Neurocomputing5
2017 Solving Indefinite Kernel Support Vector Machine with Difference of Convex Functions Programming
abstract
Indefinite kernel support vector machine (IKSVM) has recently attracted increasing attentions in machine learning. Different from traditional SVMs, IKSVM essentially is a non-convex optimization problem. Some algorithms directly change the spectrum of the indefinite kernel matrix at the cost of losing some valuable information involved in the kernels so as to transform the non-convex problem into a convex one. Other algorithms aim to solve the dual form of IKSVM, but suffer from the dual gap between the primal and dual problems in the case of indefinite kernels. In this paper, we directly focus on the non-convex primal form of IKSVM and propose a novel algorithm termed as IKSVM-DC. According to the characteristics of the spectrum for the indefinite kernel matrix, IKSVM-DC decomposes the objective function into the subtraction of two convex functions and thus reformulates the primal problem as a difference of convex functions (DC) programming which can be optimized by the DC algorithm (DCA). In order to accelerate convergence rate, IKSVM-DC further combines the classical DCA with a line search step along the descent direction at each iteration. A theoretical analysis is then presented to validate that IKSVM-DC can converge to a local minimum. Systematical experiments on real-world datasets demonstrate the superiority of IKSVM-DC compared to state-of-the-art IKSVM related algorithms.
Hui Xue 0002, Xiaohong Chen 0001
AAAI2
2017 Multiple Indefinite Kernel Learning for Feature Selection
abstract
Multiple kernel learning for feature selection (MKL-FS) utilizes kernels to explore complex properties of features and performs better in embedded methods. However, the kernels in MKL-FS are generally limited to be positive definite. In fact, indefinite kernels often emerge in actual applications and can achieve better empirical performance. But due to the non-convexity of indefinite kernels, existing MKL-FS methods are usually inapplicable and the corresponding research is also relatively little. In this paper, we propose a novel multiple indefinite kernel feature selection method (MIK-FS) based on the primal framework of indefinite kernel support vector machine (IKSVM), which applies an indefinite base kernel for each feature and then exerts an l1-norm constraint on kernel combination coefficients to select features automatically. A two-stage algorithm is further presented to optimize the coefficients of IKSVM and kernel combination alternately. In the algorithm, we reformulate the non-convex optimization problem of primal IKSVM as a difference of convex functions (DC) programming and transform the non-convex problem into a convex one with the affine minorization approximation. Experiments on real-world datasets demonstrate that MIK-FS is superior to some related state-of-the-art methods in both feature selection and classification performance.
Hui Xue 0002
IJCAI1
2017 Towards Safe Semi-supervised Classification: Adjusted Cluster Assumption via Clustering
Zhenyong Fu, Hui Xue 0002
Neural Process. Lett.4
2017 Semi-supervised manifold regularization with adaptive graph construction
Yun Li 0009, Songcan Chen, Zhenyong Fu, Hui Xue 0002
Pattern Recognit. Lett.6
2016 Logistic Boosting Regression for Label Distribution Learning
abstract
Label Distribution Learning (LDL) is a general learning framework which includes both single label and multi-label learning as its special cases. One of the main assumptions made in traditional LDL algorithms is the derivation of the parametric model as the maximum entropy model. While it is a reasonable assumption without additional information, there is no particular evidence supporting it in the problem of LDL. Alternatively, using a general LDL model family to approximate this parametric model can avoid the potential influence of the specific model. In order to learn this general model family, this paper uses a method called Logistic Boosting Regression (LogitBoost) which can be seen as an additive weighted function regression from the statistical viewpoint. For each step, we can fit individual weighted regression function (base learner) to realize the optimization gradually. The base learners are chosen as weighted regression tree and vector tree, which constitute two algorithms named LDLogitBoost and AOSO-LDLogitBoost in this paper. Experiments on facial expression recognition, crowd opinion prediction on movies and apparent age estimation show that LDLogitBoost and AOSO-LDLogitBoost can achieve better performance than traditional LDL algorithms as well as other LogitBoost algorithms.
Xin Geng 0001, Hui Xue 0002
CVPR3
2016 Maximum Margin Tree Error Correcting Output Codes
Fa Zheng, Hui Xue 0002, Xiaohong Chen 0001
PRICAI2
2016 Human Age Estimation by Considering both the Ordinality and Similarity of Ages
Qing Tian 0002, Hui Xue 0002, Lishan Qiao
Neural Process. Lett.2
2016 Joint Binary Classifier Learning for ECOC-Based Multi-Class Classification
abstract
Error-correcting output coding (ECOC) is one of the most widely used strategies for dealing with multi-class problems by decomposing the original multi-class problem into a series of binary sub-problems. In traditional ECOC-based methods, binary classifiers corresponding to those sub-problems are usually trained separately without considering the relationships among these classifiers. However, as these classifiers are established on the same training data, there may be some inherent relationships among them. Exploiting such relationships can potentially improve the generalization performances of individual classifiers, and, thus, boost ECOC learning algorithms. In this paper, we explore to mine and utilize such relationship through a joint classifier learning method, by integrating the training of binary classifiers and the learning of the relationship among them into a unified objective function. We also develop an efficient alternating optimization algorithm to solve the objective function. To evaluate the proposed method, we perform a series of experiments on eleven datasets from the UCI machine learning repository as well as two datasets from real-world image recognition tasks. The experimental results demonstrate the efficacy of the proposed method, compared with state-of-the-art methods for ECOC-based multi-class classification.
Mingxia Liu 0001, Daoqiang Zhang, Songcan Chen, Hui Xue 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2015 Emotion Distribution Recognition from Facial Expressions
abstract
Most existing facial expression recognition methods assume the availability of a single emotion for each expression in the training set. However, in practical applications, an expression rarely expresses pure emotion, but often a mixture of different emotions. To address this problem, this paper deals with a more common case where multiple emotions are associated to each expression. The key idea is to learn the specific description degrees of all basic emotions for each expression and the mapping from the expression images to the emotion distributions by the proposed emotion distribution learning (EDL) method.The databases used in the experiments are the s-JAFFE database and the s-BU\_3DFE database as they are the databases with explicit scores for each emotion on each expression image. Experimental results show that EDL can effectively deal with the emotion distribution recognition problem and perform remarkably better than the state-of-the-art multi-label learning methods.
Hui Xue 0002, Xin Geng 0001
ACM Multimedia2
2015 Semi-supervised classification learning by discrimination-aware manifold regularization
Songcan Chen, Hui Xue 0002, Zhenyong Fu
Neurocomputing3
2014 l p -norm Multiple Kernel Learning with Diversity of Classes
Dayin Zhang, Hui Xue 0002
PKAW2
2014 Discriminality-driven regularization framework for indefinite kernel machine
Hui Xue 0002, Songcan Chen
Neurocomputing1
2012 Discriminative indefinite kernel classifier from pairwise constraints and unlabeled data
Hui Xue 0002, Songcan Chen, Jijian Huang
ICPR1
2012 Semi-Supervised Discriminatively Regularized Classifier with Pairwise Constraints
Jijian Huang, Hui Xue 0002, Yuqing Zhai
PRICAI2
2012 Can under-exploited structure of original-classes help ECOC-based multi-class classification?
Songcan Chen, Hui Xue 0002
Neurocomputing3
2012 A unified dimensionality reduction framework for semi-paired and semi-supervised multi-view data
Xiaohong Chen 0001, Songcan Chen, Hui Xue 0002
Pattern Recognit.3
2011 Support Vector Machine incorporated with feature discrimination
Songcan Chen, Hui Xue 0002
Expert Syst. Appl.3
2011 Glocalization pursuit support vector machine
Hui Xue 0002, Songcan Chen
Neural Comput. Appl.1
2011 Structural Regularized Support Vector Machine: A Framework for Structural Large Margin Classifier
abstract
Support vector machine (SVM), as one of the most popular classifiers, aims to find a hyperplane that can separate two classes of data with maximal margin. SVM classifiers are focused on achieving more separation between classes than exploiting the structures in the training data within classes. However, the structural information, as an implicit prior knowledge, has recently been found to be vital for designing a good classifier in different real-world problems. Accordingly, using as much prior structural information in data as possible to help improve the generalization ability of a classifier has yielded a class of effective structural large margin classifiers, such as the structured large margin machine (SLMM) and the Laplacian support vector machine (LapSVM). In this paper, we unify these classifiers into a common framework from the concept of "structural granularity" and the formulation for optimization problems. We exploit the quadratic programming (QP) and second-order cone programming (SOCP) methods, and derive a novel large margin classifier, we call the new classifier the structural regularized support vector machine (SRSVM). Unlike both SLMM at the cross of the cluster granularity and SOCP and LapSVM at the cross of the point granularity and QP, SRSVM is located at the cross of the cluster granularity and QP and thus follows the same optimization formulation as LapSVM to overcome large computational complexity and non-sparse solution in SLMM. In addition, it integrates the compactness within classes with the separability between classes simultaneously. Furthermore, it is possible to derive generalization bounds for these algorithms by using eigenvalue analysis of the kernel matrices. Experimental results demonstrate that SRSVM is often superior in classification and generalization performances to the state-of-the-art algorithms in the framework, both with the same and different structural granularities.
Hui Xue 0002, Songcan Chen, Qiang Yang 0001
IEEE Trans. Neural Networks1
2010 Structure-Embedded AUC-SVM
abstract
AUC-SVM directly maximizes the area under the ROC curve (AUC) through minimizing its hinge loss relaxation, and the decision function is determined by those support vector sample pairs playing the same roles as the support vector samples in SVM. Such a learning paradigm generally emphasizes more on the local discriminative information just associated with these support vectors whereas hardly takes the overall view of data into account, thereby it may incur loss of the global distribution information in data favorable for classification. Moreover, due to the high computational complexity of AUC-SVM induced by the large number of training sample pairs quadratic in the number of samples, sampling is usually adopted, incurring a further loss of the distribution information in data. In order to compensate the distribution information loss and simultaneously boost the AUC-SVM performance, in this paper, we develop a novel structure-embedded AUC-SVM (SAUC-SVM for short) through embedding the global structure information in the whole data into AUC-SVM. With such an embedding, the proposed SAUC-SVM incorporates the local discriminative information and global structure information in data into a uniform formulation and consequently guarantees better generalization performance. Comparative experiments on both synthetic and real datasets confirm its effectiveness.
Songcan Chen, Hui Xue 0002
Int. J. Pattern Recognit. Artif. Intell.3
2010 A Novel Regularization Learning for Single-View Patterns: Multi-View Discriminative Regularization
Zhe Wang 0002, Songcan Chen, Hui Xue 0002
Neural Process. Lett.3
2009 Local ridge regression for face recognition
Hui Xue 0002, Yulian Zhu, Songcan Chen
Neurocomputing1
2009 Discriminatively regularized least-squares classification
Hui Xue 0002, Songcan Chen, Qiang Yang 0001
Pattern Recognit.1
2008 Structural Support Vector Machine
Hui Xue 0002, Songcan Chen, Qiang Yang 0001
ISNN (1)1
2008 Classifier learning with a new locality regularization method
Hui Xue 0002, Songcan Chen, Xiaoqin Zeng
Pattern Recognit.1