Jing-Hao Xue

dblp:72/1980 · DBLP profile ↗
← Back
196ranked-venue papers
15as first author
117since 2021 · last 2026
0000-0003-1174-610XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 107 · 12 first-author · 63 since 2021Graphics, computer vision, multimedia, augmented reality and games · 83 · 3 first-author · 55 since 2021Databases, data management, data science and information retrieval · 7 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Security and privacy · 5 · 3 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Dawid-Skene-model-based label-noise mitigation for federated learning
abstract
Federated learning (FL) enables collaborative model training without centralising raw data, but its performance is susceptible to label noise from clients. A common mitigation strategy involves using a clean, labelled public dataset at the server to assess client reliability. However, this approach is impractical due to the unrealistic assumption of availability of a clean, labelled public dataset. To address this issue, we propose FedDS, a novel approach that brings the Dawid-Skene model from statistical analysis to FL, which enables the estimation of the reliability of each client in FL without requiring any labelled data at the server. This approach effectively mitigates the adverse impact of heterogeneous label noise under a weaker and more practical assumption, offering a robust aggregation strategy for real-world FL scenarios with label noise. The code is available at https://github.com/Gia99999/FedDS .
Jia Dong, Rui Zhu 0006, Xinyi Shang, Jing-Hao Xue
Inf. Sci.4
2026 Federated learning with noisy labels: A comprehensive and concise review of current methodologies and future directions
abstract
Federated learning, a vital paradigm in modern machine learning, enables private and decentralised training of models that is crucial for learning from sensitive data. Noisy label learning, another vital paradigm in modern machine learning, addresses the training of models from the data with potentially incorrect labels. Their integration, namely federated learning with noisy labels (FLNL), is an emerging but challenging topic arising from the practice of machine learning, which, however, still lacks a review of its research progress. The aim of this paper is to fill in this gap. We first summarise four core challenges to FLNL: localised label noise, across-client heterogeneity of label noise, localised overfitting to label noise, and inadequate benchmarking. We then propose a taxonomy to categorise current FLNL studies into four types that address the four challenges correspondingly: sample-wise methods, client-wise methods, model-wise methods, and benchmark-wise studies. This work offers the first comprehensive and concise review dedicated to FLNL; moreover, we also provide future research directions for this rapidly evolving and practically significant field.
Jia Dong, Rui Zhu 0006, Xinyi Shang, Jing-Hao Xue
Neural Networks4
2026 UC-PUAL: A universally consistent classifier of positive-unlabelled data
abstract
Positive-unlabelled (PU) learning is a challenging task in pattern recognition, as there are only labelled-positive instances and unlabelled instances available for the training of a classifier. The task becomes even harder when the PU data show an underlying trifurcate pattern that positive instances roughly distribute on both sides of ground-truth negative instances. To address this issue, we propose a universally consistent PU classifier with asymmetric loss (UC-PUAL) on positive instances. We also propose two three-block algorithms for non-convex optimisation to enable UC-PUAL to obtain linear and kernel-induced non-linear decision boundaries, respectively. Theoretical and experimental results verify the superiority of UC-PUAL. The code for UC-PUAL is available at https://github.com/tkks22123/UC-PUAL.
Rui Zhu 0006, Jing-Hao Xue
Pattern Recognit.3
2026 Global-Local Disturbance Decoupling for Federated Facial Expression Recognition
abstract
Most existing facial expression recognition (FER) methods are designed for centralized model training on largescale data. Unfortunately, accessing massive facial expression data can be difficult due to privacy concerns in practice. In this paper, we study an important but little-explored task, federated FER, which allows us to train an FER model with decentralized expression data. To this end, we propose a novel global-local disturbance decoupling (GLDD) method for federated FER. Specifically, for local disturbance decoupling on each client, we develop a dual-branch feature decoupling network consisting of a backbone network, an expression branch, and a disturbance branch, to perform local FER. In the disturbance branch, we design an entropy-guided feature encoding module to extract priorbased disturbance features. This greatly facilitates the extraction of client-specific disturbance features. For global disturbance decoupling on the server, we introduce orthogonal decoupling, which is a global-level feature disentanglement technique that enforces mutual orthogonality between the global expression feature class centers and the global disturbance feature centers, thereby eliminating cross-client disturbances across clients and yields decoupled global expression feature class centers. These centers are then used to retrain the global model, substantially enhancing disturbance invariance and classification performance. By jointly performing disturbance decoupling at local and global levels, our method effectively addresses the unique challenges of heterogeneous expression data and heterogeneous disturbances in federated FER. Experimental results on two real-world facial expression databases show that, on the federated FER task, our method significantly outperforms several state-of-the-art federated learning methods and FER methods. The code will be released soon.
Hu Ding 0005, Yan Yan 0001, Yang Lu 0009, Jing-Hao Xue, Hanzi Wang
IEEE Trans. Affect. Comput.4
2026 Fine-Tuning via Linked Domains: A Closed-Form Dual Alignment Mechanism for Transferring Vision-Language Models
abstract
Adapters and prompt learning have become two de facto strategies to fine-tune pre-trained vision-language models, mitigating the high computational cost of fine-tuning an entire model for downstream tasks. They can align the prediction from the fine-tuned model with that from the pre-trained model. However, the existing methods of these strategies primarily focus on aligning within a single modality, and the exploration of bidirectional interactions between modalities remains limited. To address this issue, we propose a closed-form dual alignment mechanism (DAM) thatnot only ensures the consistency in predictions within a single modality but also achieves the alignment of features across different modalities. In DAM, all alignments are achieved by closed-form solutions to ridge regression, without inducing a massive number of learnable parameters. Experimental results demonstrate that DAM outperforms the state-of-the-art methods on 11 benchmarks over various evaluation metrics. Our codes are available at https://github.com/Peiy-Lu/DAM.
Peiyu Lu, Rui Zhu 0006, Zhanyu Ma, Jie Cao 0014, Jing-Hao Xue
IEEE Trans. Circuits Syst. Video Technol.6
2026 HOH-Net: High-Order Hierarchical Middle-Feature Learning Network for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) is a cross-modality retrieval task that aims to match images of the same person across visible (VIS) and infrared (IR) modalities. Existing VI-ReID methods ignore high-order structure information of features and struggle to learn a reliable common feature space due to the modality discrepancy between VIS and IR images. To alleviate the above issues, we propose a novel high-order hierarchical middle-feature learning network (HOH-Net) for VI-ReID. We introduce a high-order structure learning (HSL) module to explore the high-order relationships of short- and long-range feature nodes, for significantly mitigating model collapse and effectively obtaining discriminative features. We further develop a fine-coarse graph attention alignment (FCGA) module, which efficiently aligns multi-modality feature nodes from node-level and region-level perspectives, ensuring reliable middle-feature representations. Moreover, we exploit a hierarchical middle-feature agent learning (HMAL) loss to hierarchically reduce the modality discrepancy at each stage of the network by using the agents of middle features. The proposed HMAL loss also exchanges detailed and semantic information between low- and high-stage networks. Finally, we introduce a modality-range identity-center contrastive (MRIC) loss to minimize the distances between VIS, IR, and middle features. Extensive experiments demonstrate that the proposed HOH-Net yields state-of-the-art performance on the image-based and video-based VI-ReID datasets. The code is available at: https://github.com/Jaulaucoeng/HOS-Net.
Liuxiang Qiu, Si Chen 0002, Jing-Hao Xue, Dahan Wang, Shunzhi Zhu, Yan Yan 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 A Survey on Text-Driven 360° Panorama Generation
abstract
The advent of text-driven 360-degree panorama generation, enabling the synthesis of 360-degree panoramic images directly from textual descriptions, marks a transformative advancement in immersive visual content creation. This innovation significantly simplifies the traditionally complex process of producing such content. Recent progress in text-to-image diffusion models has accelerated the rapid development in this emerging field. This survey presents a comprehensive review of text-driven 360-degree panorama generation, offering an in-depth analysis of state-of-the-art algorithms. We extend our analysis to two closely related domains: text-driven 360-degree 3D scene generation and text-driven 360-degree panoramic video generation. Furthermore, we critically examine current limitations and propose promising directions for future research. A curated project page with relevant resources and research papers is available at https://littlewhitesea.github.io/Text-Driven-Pano-Gen/.
Hai Wang 0020, Xiaoyu Xiang, Weihao Xia 0001, Jing-Hao Xue
IEEE Trans. Circuits Syst. Video Technol.4
2026 BDC-Occ: Binarized Deep Convolution Unit for Binarized Occupancy Network
abstract
Existing 3D occupancy networks demand significant hardware resources, hindering the deployment of resource-limited devices. Binarized Neural Networks (BNNs) offer a potential solution by substantially reducing computational and memory requirements. However, their performance decrease notably compared to full-precision networks. In addition, it is challenging to enhance the performance of the binarized model by increasing the number of binarized convolutional layers, which limits its practicability for 3D occupancy prediction. In this paper, we reconsider the components in binarized convolutional layers, and structures, for 3D occupancy prediction task. Two original insights into binarized convolution are presented, substantiated with theoretical proofs: (a) 1×1 binarized convolution introduces minimal binarization errors as the network deepens, and (b) binarized convolution is inferior to full-precision convolution in capturing cross-channel feature importance. Building on the above insights, we propose a novel binarized deep convolution (BDC) unit that significantly enhances performance, even when the number of binarized convolutional layers increases to meet the requirements of 3D occupancy networks. Specifically, in the BDC unit, additional binarized convolutional kernels are constrained to 1×1 to minimize the effects of binarization errors. Further, we propose a per-channel refinement branch to reweight the output via first-order approximation. Then, we partition the 3D occupancy networks into four distinct convolutional modules, employing BDC units to explore the effects of binarizing each of these modules. The proposed BDC unit minimizes binarization errors and improves perceptual capability, meeting the stringent requirements for accuracy and computational efficiency in 3D occupancy prediction. Extensive quantitative and qualitative experiments demonstrate that the proposed BDC unit achieves state-of-the-art performance in 3D occupancy prediction and 3D object detection tasks, while significantly reducing parameters and computational costs. This highlights the potential of the BDC unit as an efficient fundamental component in binarized 3D occupancy networks. Code for our paper will be released on “https://github.com/zzk785089755/BDC”.
Zongkai Zhang, Peng Ling, Zidong Xu, Wenming Yang, Qingmin Liao, Jing-Hao Xue
IEEE Trans. Circuits Syst. Video Technol.6
2025 Mind the Gap: Confidence Discrepancy Can Guide Federated Semi-Supervised Learning Across Pseudo-Mismatch
abstract
Federated Semi-Supervised Learning (FSSL) aims to leverage unlabeled data across clients with limited labeled data to train a global model with strong generalization ability. Most FSSL methods rely on consistency regularization with pseudo-labels, converting predictions from local or global models into hard pseudo-labels as supervisory signals. However, we discover that the quality of pseudo-label is largely deteriorated by data heterogeneity, an intrinsic facet of federated learning. In this paper, we study the problem of FSSL in-depth and show that (1) heterogeneity exacerbates pseudo-label mismatches, further degrading model performance and convergence, and (2) local and global models’ predictive tendencies diverge as heterogeneity increases. Motivated by these findings, we propose a simple and effective method called Semi-supervised Aggregation for Globally-Enhanced Ensemble (SAGE), that can flexibly correct pseudo-labels based on confidence discrepancies. This strategy effectively mitigates performance degradation caused by incorrect pseudo-labels and enhances consensus between local and global models. Experimental results demonstrate that SAGE outperforms existing FSSL methods in both performance and convergence. Our code is available at https://github.com/Jay-Codeman/SAGE.
Xinyi Shang, Yiqun Zhang 0006, Yang Lu 0009, Chen Gong 0002, Jing-Hao Xue, Hanzi Wang
CVPR6
2025 Identity-Preserving Diffusion for Face Restoration
abstract
Face restoration is a critical task in computer vision, aiming to restore high-quality facial images from degraded inputs. In existing diffusion models, identity information is not well preserved when confronted with severely degradation. To address this challenge, we propose a Local Patch-Based Identity-Preserving Diffusion (LPIP-Diff) framework. Our local patch-based strategy leverages the interrelationships between neighboring patches to model highly structured facial context, which facilitates the restoration of fine-grained details and the preservation of identity-related features. We also introduce a fusion degradation estimation method that makes each overlapping area restored multiple times by adjacent patches, effectively restoring local details. The experimental results of LPIP-Diff on three publicly available datasets, including one severely degraded dataset, consistently demonstrate its superiority over the state-of-the-art methods in terms of both quantitative and qualitative evaluations, strikes a good balance between realism and fidelity, and enhances robustness against degradation.
Xiaying Bai, Wenming Yang, Rui Zhu 0006, Jing-Hao Xue
ICASSP5
2025 DARL: Mitigating Gradient Conflicts in Long-Tailed Out-of-Distribution Learning
abstract
Long-tailed out-of-distribution learning aims to reduce performance bias in long-tailed in-distribution (ID) data while rejecting out-of-distribution (OOD) samples, which are often mistaken for under-represented tail classes. To achieve OOD detection, existing methods incorporate an outlier exposure (OE) term into the long-tailed recognition (LTR) loss. However, as we prove in this paper, the OE term induces a gradient conflict with the ID objectives, especially for tail classes, thereby contradicting the core motivation of LTR. To avoid the ID-OOD dilemma, we propose Dynamic Ambiguity-aware Recalibration for Logits (DARL), an ambiguity-guided long-tailed OOD learning approach, grounded on two theoretical insights. First, we show that the mixed ID data can mitigate the conflict in OE training and exhibits higher intrinsic ambiguity than the original ID data, thus able to serve as a surrogate for real OOD data. Second, we introduce an ambiguity-aware logit adjustment that can dynamically calibrate the class margins using energy-based ambiguity metrics, effectively reducing early-stage bias while avoiding late-stage overfitting. Extensive experiments show that DARL achieves the overall state-of-the-art performance of long-tailed OOD learning. Moreover, compared with the OE methods, DARL trains solely on the ID data, which can reduce the data requirements by 80%. The code is available in https://github.com/XuanZhang-A/DARL.
Sin Chee Chin, Jing-Hao Xue, Wenming Yang
ACM Multimedia3
2025 WarpGAN: Warping-Guided 3D GAN Inversion with Style-Based Novel View Inpainting
abstract
3D GAN inversion projects a single image into the latent space of a pre-trained 3D GAN to achieve single-shot novel view synthesis, which requires visible regions with high fidelity and occluded regions with realism and multi-view consistency. However, existing methods focus on the reconstruction of visible regions, while the generation of occluded regions relies only on the generative prior of 3D GAN. As a result, the generated occluded regions often exhibit poor quality due to the information loss caused by the low bit-rate latent code. To address this, we introduce the warping-and-inpainting strategy to incorporate image inpainting into 3D GAN inversion and propose a novel 3D GAN inversion method, WarpGAN. Specifically, we first employ a 3D GAN inversion encoder to project the single-view image into a latent code that serves as the input to 3D GAN. Then, we perform warping to a novel view using the depth map generated by 3D GAN. Finally, we develop a novel SVINet, which leverages the symmetry prior and multi-view image correspondence w.r.t. the same latent code to perform inpainting of occluded regions in the warped image. Quantitative and qualitative experiments demonstrate that our method consistently outperforms several state-of-the-art methods.
Kaitao Huang, Yan Yan 0001, Jing-Hao Xue, Hanzi Wang
NeurIPS3
2025 Towards Reliable and Holistic Visual In-Context Learning Prompt Selection
abstract
Visual In-Context Learning (VICL) has emerged as a prominent approach for adapting visual foundation models to novel tasks, by effectively exploiting contextual information embedded in in-context examples, which can be formulated as a global ranking problem of potential candidates. Current VICL methods, such as Partial2Global and VPR, are grounded in the similarity-priority assumption that images more visually similar to a query image serve as better in-context examples. This foundational assumption, while intuitive, lacks sufficient justification for its efficacy in selecting optimal in-context examples. Furthermore, Partial2Global constructs its global ranking from a series of randomly sampled pairwise preference predictions. Such a reliance on random sampling can lead to incomplete coverage and redundant samplings of comparisons, thus further adversely impacting the final global ranking. To address these issues, this paper introduces an enhanced variant of Partial2Global designed for reliable and holistic selection of in-context examples in VICL. Our proposed method, dubbed RH-Partial2Global, leverages a jackknife conformal prediction-guided strategy to construct reliable alternative sets and a covering design-based sampling approach to ensure comprehensive and uniform coverage of pairwise preferences. Extensive experiments demonstrate that RH-Partial2Global achieves excellent performance and outperforms Partial2Global across diverse visual tasks.
Wenxiao Wu, Jing-Hao Xue, Chengming Xu 0001, Chen Liu 0030, Xinwei Sun 0001, Changxin Gao, Nong Sang, Yanwei Fu 0001
NeurIPS2
2025 360PanT: Training-Free Text-Driven 360-Degree Panorama-to-Panorama Translation
abstract
Preserving boundary continuity in the translation of 360-degree panoramas remains a significant challenge for existing text-driven image-to-image translation methods. These methods often produce visually jarring discontinuities at the translated panorama's boundaries, disrupting the immersive experience. To address this issue, we propose 360PanT, a training-free approach to text-based 360-degree panorama-to-panorama translation with boundary continuity. Our 360PanT achieves seamless translations through two key components: boundary continuity encoding and seamless tiling translation with spatial control. Firstly, the boundary continuity encoding embeds critical boundary continuity information of the input 360-degree panorama into the noisy latent representation by constructing an extended input image. Secondly, leveraging this embedded noisy latent representation and guided by a target prompt, the seamless tiling translation with spatial control enables the generation of a translated image with identical left and right halves while adhering to the extended input's structure and semantic layout. This process ensures a final translated 360-degree panorama with seamless boundary continuity. Experimental results on both real-world and synthesized datasets demonstrate the effectiveness of our 360PanT in translating 360-degree panoramas. Code is available at https://github.com/littlewhitesea/360PanT.
Hai Wang 0020, Jing-Hao Xue
WACV2
2025 Interactive triplet attention for few-shot fine-grained image classification
abstract
Few-shot fine-grained classification aims to identify novel fine-grained classes from extremely few examples with ultra-high semantic similarity between classes, hence a notoriously hard task. To extract discriminative features from few samples for recognizing subtle differences between fine-grained classes , it is pivotal to exploit comprehensive interactions across all dimensions in space and channel, which, however, is unexplored yet by state-of-the-art methods in this challenging area. To address this issue, in this paper we show that a simple adjustment to the existing triplet attention module (TAM) can be highly effective for few-shot fine-grained image classification. More specifically, building on TAM which comprises three parallel branches for pairwise interactions between height, width, and channel dimensions, we introduce an additional interaction between the output of these three branches, capable of modeling the dependency across all three dimensions; the revised method is dubbed interactive triplet attention module (ITAM). ITAM is a plug-and-play module, which can be inserted into any metric-based few-shot fine-grained image classifiers for performance enhancement. Extensive experiments, on CUB-200–2011, Flowers, Stanford-Cars, and Stanford-Dogs, showcase the superiority of ITAM against state-of-the-art few-shot fine-grained image classifiers.
Shaoying Xue, Jiyang Xie 0001, Zhanyu Ma, Jing-Hao Xue
Neurocomputing6
2025 PUAL: A classifier on trifurcate positive-unlabelled data
Rui Zhu 0006, Jing-Hao Xue
Neurocomputing4
2025 GKF-PUAL: A group kernel-free approach to positive-unlabeled learning with variable selection
abstract
Variable selection is important for classification of data with many irrelevant predicting variables, but it has not yet been well studied in positive-unlabeled (PU) learning, where classifiers have to be trained without labelled-negative instances. In this paper, we propose a group kernel-free PU classifier with asymmetric loss (GKF-PUAL) to achieve quadratic PU classification with group-lasso regularisation embedded for variable selection. We also propose a five-block algorithm to solve the optimization problem of GKF-PUAL. Our experimental results reveal the superiority of GKF-PUAL in both PU classification and variable selection, improving the baseline PUAL by more than 10% in F1-score across four benchmark datasets and removing over 70% of irrelevant variables on six benchmark datasets. The code for GKF-PUAL is at https://github.com/tkks22123/GKF-PUAL . • We propose a group kernel-free PU classifier (GKF-PUAL) with variable selection. • We propose a five-block algorithm for optimization of GKF-PUAL. • Experimental results verify the superiority of GKF-PUAL.
Rui Zhu 0006, Jing-Hao Xue
Inf. Sci.3
2025 Consistency-driven feature scoring and regularization network for visible-infrared person re-identification
Xueting Chen, Yan Yan 0001, Jing-Hao Xue, Nannan Wang 0001, Hanzi Wang
Pattern Recognit.3
2025 CDN4: A cross-view Deep Nearest Neighbor Neural Network for fine-grained few-shot classification
abstract
The fine-grained few-shot classification is a challenging task in computer vision, aiming to classify images with subtle and detailed differences given scarce labeled samples. A promising avenue to tackle this challenge is to use spatially local features to densely measure the similarity between query and support samples. Compared with image-level global features, local features contain more low-level information that is rich and transferable across categories. However, methods based on spatially localized features have difficulty distinguishing subtle category differences due to the lack of sample diversity. To address this issue, we propose a novel method called Cross-view Deep Nearest Neighbor Neural Network (CDN4). CDN4 applies a random geometric transformation to augment a different view of support and query samples and subsequently exploits four similarities between the original and transformed views of query local features and those views of support local features. The geometric augmentation increases the diversity between samples of the same class, and the cross-view measurement encourages the model to focus more on discriminative local features for classification through the cross-measurements between the two branches. Extensive experiments validate the superiority of CDN4, which achieves new state-of-the-art results in few-shot classification across various fine-grained benchmarks. Code is available at . • A novel fine-grained FSL method for improving local feature discriminativeness. • Construct the Episodic Dual-Branch Structure to enhance sample diversity. • Exploit four cross-view metric pairs to enforce learning of discriminative features. • CDN4 achieves state-of-the-art performance on three fine-grained benchmark datasets.
Jiyang Xie 0001, Zhanyu Ma, Jing-Hao Xue
Pattern Recognit.6
2025 Clarity in chaos: Boosting few-shot classification through information suppression and sparsification
abstract
The advance of deep learning has invigorated the research of few-shot classification. However, the interference of non-target information in feature representations hampers classification generalization. To tackle this issue, we propose an irrelevant information suppression (IIS) module, which is focused on suppressing the weight of unimportant information and elevating the sparsity of feature representations . An IIS network with three consecutive IIS modules is developed, to illustrate the progressive suppression of unimportant information and highlighting of key discriminative features of the target. Extensive experiments showcase the superior performance of our IIS network on five widely-used benchmark datasets. Furthermore, we show that the IIS module can be readily used as a plug-in module by state-of-the-art few-shot classifiers, and can clearly further improve their performance. Our code is available on GitHub at https://github.com/LC4188/IISNet . • We propose an IIS module to progressively suppress non-target information. • The IIS module can be readily used as a plug-in module. • The IIS module can clearly improve the performance of few-shot classifiers.
Luchen Ji, Rui Zhu 0006, Zhanyu Ma, Jing-Hao Xue
Pattern Recognit.5
2025 Self-randomized focuses effectively boost metric-based few-shot classifiers
Zhen Li 0026, Zhongyuan Liu, Dongliang Chang, Aneeshan Sain, Zhanyu Ma, Jing-Hao Xue, Yi-Zhe Song
Pattern Recognit.7
2025 SRML: Structure-relation mutual learning network for few-shot image classification
abstract
Few-shot image classification aims at tackling a challenging but practical classification setting, where only few labelled images are available for training. Metric-based methods are main-stream solutions for few-shot image classification, but many of them extract features that are either irrelevant to target objects in the query images or insufficient to describe the local shape or structural patterns within images, which can lead to mis-identification of the target objects, especially when the images are of multiple objects. To resolve this issue, we propose the structure-relation mutual learning (SRML) network, which first learns both the intra-image structural features and the inter-image relational features in a parallel fashion via two parallel branches, the structural feature extractor (SFE) and the relational feature extractor (RFE), and then harnesses mutual learning to enable knowledge exchange between them. In such a manner, the structural features learnt from the SFE branch not only contain the structural patterns within the images, but also focus more on the target objects, guided by the relational knowledge from the RFE branch. In return, the RFE branch can exploit the more-focused structural knowledge to better match the target objects in the support and query images. We conduct extensive experiments on four few-shot classification benchmark datasets to showcase the superior classification of the proposed SRML network, achieving a 3.17% improvement in classification accuracy over the leading competitor, RENet Kang et al. (2021). The code of this work can be found in https://github.com/Rilliant7/SRML .
Rui Zhu 0006, Zhanyu Ma, Jie Cao 0014, Jing-Hao Xue
Pattern Recognit.6
2025 Uncertainty-Aware Label Refinement on Hypergraphs for Personalized Federated Facial Expression Recognition
abstract
Most facial expression recognition (FER) models are trained on large-scale expression data with centralized learning. Unfortunately, collecting a large amount of centralized expression data is difficult in practice due to privacy concerns of facial images. In this paper, we investigate FER under the framework of personalized federated learning, which is a valuable and practical decentralized setting for real-world applications. To this end, we develop a novel uncertainty-Aware label refineMent on hYpergraphs (AMY) method. For local training, each local model consists of a backbone, an uncertainty estimation (UE) block, and an expression classification (EC) block. In the UE block, we leverage a hypergraph to model complex high-order relationships between expression samples and incorporate these relationships into uncertainty features. A personalized uncertainty estimator is then introduced to estimate reliable uncertainty weights of samples in the local client. In the EC block, we perform label propagation on the hypergraph, obtaining high-quality refined labels for retraining an expression classifier. Based on the above, we effectively alleviate heterogeneous sample uncertainty across clients and learn a robust personalized FER model in each client. Experimental results on two challenging real-world facial expression databases show that our proposed method consistently outperforms several state-of-the-art methods. This indicates the superiority of hypergraph modeling for uncertainty estimation and label refinement on the personalized federated FER task. The source code will be released athttps://github.com/mobei1006/AMY.
Hu Ding 0005, Yan Yan 0001, Yang Lu 0009, Jing-Hao Xue, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.4
2025 Rise by Lifting Others: Interacting Features to Uplift Few-Shot Fine-Grained Classification
abstract
Few-shot fine-grained classification entails notorious subtle inter-class variation. Recent works address this challenge by developing attention mechanisms, such as the task discrepancy maximization (TDM) that can highlight discriminative channels. This paper, however, aims to reveal that, besides designing sophisticated attention modules, a well-designed input scheme, which simply blends two types of features and their interactions capturing different properties of the target object, can also greatly promote the quality of the learnt weights. To illustrate, we design a bi-feature interactive TDM (BiFI-TDM) module to serve as a strong foundation for TDM to discover the most discriminative channels with ease. Specifically, we design a novel mixing strategy to produce four sets of channel weights with different focuses, reflecting the properties of the corresponding input features and their interactions, as well as a proper feature re-weighting scheme. Extensive experiments on four benchmark fine-grained image datasets showcase superior performance of BiFI-TDM in metric-based few-shot methods. Our codes are available athttps://github.com/Peiy-Lu/BiFI-TDM.
Peiyu Lu, Rui Zhu 0006, Zhanyu Ma, Jie Cao 0014, Jing-Hao Xue
IEEE Trans. Circuits Syst. Video Technol.6
2025 Selectively Augmented Attention Network for Few-Shot Image Classification
abstract
Few-shot image classification is a challenging task that aims to learn from a limited number of labelled training images a classification model that can be generalised to unseen classes. Two strategies are usually taken to improve the classification performances of few-shot image classifiers: either applying data augmentation to enlarge the sample size of the training set and reduce overfitting, or involving attention mechanisms to highlight discriminative spatial regions or channels. However, naively applying them to few-shot classifiers directly and separately may lead to undesirable results; for example, some augmented images may focus majorly on the background rather than the object, which brings additional noises to the training process. In this paper, we propose a unified framework, the selectively augmented attention (SAA) network, that carefully integrates the best of the two approaches in an end-to-end fashion via a selective best match module to select the most representative images from the augmented training set. The selected images tend to concentrate on the objects with less irrelevant background, which can assist the subsequent calculation of attentions by alleviating the interference from background. Moreover, we design a joint attention module to jointly learn both the spatial and channel-wise attentions. Experimental results on four benchmark datasets showcase the superior classification performance of the proposed SAA network compared with the state-of-the-arts.
Rui Zhu 0006, Zhanyu Ma, Jie Cao 0014, Jing-Hao Xue
IEEE Trans. Circuits Syst. Video Technol.6
2025 Query-Aware Cross-Mixup and Cross-Reconstruction for Few-Shot Fine-Grained Image Classification
abstract
Few-shot fine-grained image classification is prominent but challenging in computer vision, aiming to distinguish sub-classes under the same parent class but with only a few labeled support samples. Data augmentation techniques were explored to address the few-shot issue, but they often fail to mitigate the bias between support and query samples. Therefore, in this paper we propose a query-aware cross-mixup and cross-reconstruction method to address both few-shot and fine-grained issues. Specifically, in the training phase, we randomly select query samples and mix them with the support samples from the same class to augment the support set. This first strategy ensures the augmented support set query-aware within each sub-class. Then, we reconstruct both query samples and support samples from both original and cross-mixed support samples, thus leveraging both cross-reconstruction and self-reconstruction to enhance classification. This second strategy, enabling the reconstruction also query-aware, further mitigates the bias between support and query samples, leading to more reliable generalization. We evaluate our proposed method on four widely used few-shot fine-grained image classification datasets, and experimental results demonstrate its effectiveness in achieving the state-of-the-art classification performance.
Dongliang Chang, Rui Zhu 0006, Zhanyu Ma, Jing-Hao Xue
IEEE Trans. Circuits Syst. Video Technol.6
2025 Augmentation Matters: A Mix-Paste Method for X-Ray Prohibited Item Detection Under Noisy Annotations
abstract
Automatic X-ray prohibited item detection is vital for public safety. Existing deep learning-based methods all assume that the annotations of training X-ray images are correct. However, obtaining correct annotations is extremely hard if not impossible for large-scale X-ray images, where item overlapping is ubiquitous. As a result, X-ray images are easily contaminated with noisy annotations, leading to performance deterioration of existing methods. In this paper, we address the challenging problem of training a robust prohibited item detector under noisy annotations (including both category noise and bounding box noise) from a novel perspective of data augmentation, and propose an effective label-aware mixed patch paste augmentation method (Mix-Paste). Specifically, for each item patch, we mix several item patches with the same category label from different images and replace the original patch in the image with the mixed patch. In this way, the probability of containing the correct prohibited item within the generated image is increased. Meanwhile, the mixing process mimics item overlapping, enabling the model to learn the characteristics of X-ray images. Moreover, we design an item-based large-loss suppression (LLS) strategy to suppress the large losses corresponding to potentially positive predictions of additional items due to the mixing operation. We show the superiority of our method on X-ray datasets under noisy annotations. In addition, we evaluate our method on the noisy MS-COCO dataset to showcase its generalization ability. These results clearly indicate the great potential of data augmentation to handle noise annotations. The source code is released athttps://github.com/wscds/Mix-Paste.
Ruikang Chen, Yan Yan 0001, Jing-Hao Xue, Yang Lu 0009, Hanzi Wang
IEEE Trans. Inf. Forensics Secur.3
2025 I2OL-Net: Intra-Inter Objectness Learning Network for Point-Supervised X-Ray Prohibited Item Detection
abstract
Automatic detection of prohibited items in X-ray images plays a crucial role in public security. However, existing methods rely heavily on labor-intensive box annotations. To address this, we investigate X-ray prohibited item detection under labor-efficient point supervision and develop an intra-inter objectness learning network (I2OL-Net). I2OL-Net consists of two key modules: an intra-modality objectness learning (intra-OL) module and an inter-modality objectness learning (inter-OL) module. The intra-OL module designs a local focus Gaussian masking block and a global random Gaussian masking block to collaboratively learn the objectness in X-ray images. Meanwhile, the inter-OL module introduces the wavelet decomposition-based adversarial learning block and the objectness block, effectively reducing the modality discrepancy between natural images and X-ray images and transferring the objectness knowledge learned from natural images with box annotations to X-ray images. Based on the above, I2OL-Net greatly alleviates the severe problem of part domination caused by large intra-class variations in X-ray images. Experimental results on four X-ray datasets show that I2OL-Net can achieve superior performance with a significant reduction of annotation cost, thus enhancing its accessibility and practicality. The source code is released athttps://github.com/houjoeng/I2OL-Net.
Yan Yan 0001, Jing-Hao Xue, Hanzi Wang
IEEE Trans. Inf. Forensics Secur.3
2025 Hyperbolic Self-Paced Multi-Expert Network for Cross-Domain Few-Shot Facial Expression Recognition
abstract
Recently, cross-domain few-shot facial expression recognition (CF-FER), which identifies novel compound expressions with a few images in the target domain by using the model trained only on basic expressions in the source domain, has attracted increasing attention. Generally, existing CF-FER methods leverage the multi-dataset to increase the diversity of the source domain and alleviate the discrepancy between the source and target domains. However, these methods learn feature embeddings in the Euclidean space without considering imbalanced expression categories and imbalanced sample difficulty in the multi-dataset. Such a way makes the model difficult to capture hierarchical relationships of facial expressions, resulting in inferior transferable representations. To address these issues, we propose a hyperbolic self-paced multi-expert network (HSM-Net), which contains multiple mixture-of-experts (MoE) layers located in the hyperbolic space, for CF-FER. Specifically, HSM-Net collaboratively trains multiple experts in a self-distillation manner, where each expert focuses on learning a subset of expression categories from the multi-dataset. Based on this, we introduce a hyperbolic self-paced learning (HSL) strategy that exploits sample difficulty to adaptively train the model from easy-to-hard samples, greatly reducing the influence of imbalanced expression categories and imbalanced sample difficulty. Our HSM-Net can effectively model rich hierarchical relationships of facial expressions and obtain a highly transferable feature space. Extensive experiments on both in-the-lab and in-the-wild compound expression datasets demonstrate the superiority of our proposed method over several state-of-the-art methods. Code will be released at https://github.com/cxtjl/HSM-Net.
Xueting Chen, Yan Yan 0001, Jing-Hao Xue, Hanzi Wang
IEEE Trans. Image Process.3
2025 Knowledge Distillation Meets Label Noise Learning: Ambiguity-Guided Mutual Label Refinery
abstract
Knowledge distillation (KD), which aims at transferring the knowledge from a complex network (a teacher) to a simpler and smaller network (a student), has received considerable attention in recent years. Typically, most existing KD methods work on well-labeled data. Unfortunately, real-world data often inevitably involve noisy labels, thus leading to performance deterioration of these methods. In this article, we study a little-explored but important issue, i.e., KD with noisy labels. To this end, we propose a novel KD method, called ambiguity-guided mutual label refinery KD (AML-KD), to train the student model in the presence of noisy labels. Specifically, based on the pretrained teacher model, a two-stage label refinery framework is innovatively introduced to refine labels gradually. In the first stage, we perform label propagation (LP) with small-loss selection guided by the teacher model, improving the learning capability of the student model. In the second stage, we perform mutual LP between the teacher and student models in a mutual-benefit way. During the label refinery, an ambiguity-aware weight estimation (AWE) module is developed to address the problem of ambiguous samples, avoiding overfitting these samples. One distinct advantage of AML-KD is that it is capable of learning a high-accuracy and low-cost student model with label noise. The experimental results on synthetic and real-world noisy datasets show the effectiveness of our AML-KD against state-of-the-art KD methods and label noise learning (LNL) methods. Code is available at https://github.com/Runqing-forMost/ AML-KD.
Runqing Jiang, Yan Yan 0001, Jing-Hao Xue, Si Chen 0002, Nannan Wang 0001, Hanzi Wang
IEEE Trans. Neural Networks Learn. Syst.3
2024 High-Order Structure Based Middle-Feature Learning for Visible-Infrared Person Re-identification
abstract
Visible-infrared person re-identification (VI-ReID) aims to retrieve images of the same persons captured by visible (VIS) and infrared (IR) cameras. Existing VI-ReID methods ignore high-order structure information of features while being relatively difficult to learn a reasonable common feature space due to the large modality discrepancy between VIS and IR images. To address the above problems, we propose a novel high-order structure based middle-feature learning network (HOS-Net) for effective VI-ReID. Specifically, we first leverage a short- and long-range feature extraction (SLE) module to effectively exploit both short-range and long-range features. Then, we propose a high-order structure learning (HSL) module to successfully model the high-order relationship across different local features of each person image based on a whitened hypergraph network. This greatly alleviates model collapse and enhances feature representations. Finally, we develop a common feature space learning (CFL) module to learn a discriminative and reasonable common feature space based on middle features generated by aligning features from different modalities and ranges. In particular, a modality-range identity-center contrastive (MRIC) loss is proposed to reduce the distances between the VIS, IR, and middle features, smoothing the training process. Extensive experiments on the SYSU-MM01, RegDB, and LLCM datasets show that our HOS-Net achieves superior state-of-the-art performance. Our code is available at https://github.com/Jaulaucoeng/HOS-Net.
Liuxiang Qiu, Si Chen 0002, Yan Yan 0001, Jing-Hao Xue, Dahan Wang, Shunzhi Zhu
AAAI4
2024 Spatial-Contextual Discrepancy Information Compensation for GAN Inversion
abstract
Most existing GAN inversion methods either achieve accurate reconstruction but lack editability or offer strong editability at the cost of fidelity. Hence, how to balance the distortion-editability trade-off is a significant challenge for GAN inversion. To address this challenge, we introduce a novel spatial-contextual discrepancy information compensation-based GAN-inversion method (SDIC), which consists of a discrepancy information prediction network (DIPN) and a discrepancy information compensation network (DICN). SDIC follows a ``compensate-and-edit'' paradigm and successfully bridges the gap in image details between the original image and the reconstructed/edited image. On the one hand, DIPN encodes the multi-level spatial-contextual information of the original and initial reconstructed images and then predicts a spatial-contextual guided discrepancy map with two hourglass modules. In this way, a reliable discrepancy map that models the contextual relationship and captures fine-grained image details is learned. On the other hand, DICN incorporates the predicted discrepancy information into both the latent code and the GAN generator with different transformations, generating high-quality reconstructed/edited images. This effectively compensates for the loss of image details during GAN inversion. Both quantitative and qualitative experiments demonstrate that our proposed method achieves the excellent distortion-editability trade-off at a fast inference speed for both image inversion and editing tasks. Our code is available at https://github.com/ZzqLKED/SDIC.
Yan Yan 0009, Jing-Hao Xue, Hanzi Wang
AAAI3
2024 UMBRAE: Unified Multimodal Brain Decoding
Weihao Xia 0001, Raoul de Charette, A. Cengiz Öztireli, Jing-Hao Xue
ECCV (7)4
2024 Reschedule Diffusion-based Bokeh Rendering
Shiyue Yan, Xiaoshi Qiu, Qingmin Liao, Jing-Hao Xue
IJCAI4
2024 Channel-Spatial Support-Query Cross-Attention for Fine-Grained Few-Shot Image Classification
abstract
Few-shot fine-grained image classification aims to use only few labelled samples to successfully recognize subtle sub-classes within the same parent class. This task is extremely challenging, due to the co-occurrence of large inter-class similarity, low intra-class similarity, and only few labelled samples. In this paper, to address these challenges, we propose a new Channel-Spatial Cross-Attention Module (CSCAM), which can effectively drive a model to extract discriminative fine-grained feature representations with only few shots. CSCAM collaboratively integrates a channel cross-attention module and a spatial cross-attention module, for the attentions across support and query samples. In addition, to fit for the characteristics of fine-grained images, a support averaging method is proposed in CSCAM to reduce the intra-class distance and increase the inter-class distance. Extensive experiments on four few-shot fine-grained classification datasets validate the effectiveness of CSCAM. Furthermore, CSCAM is a plug-and-play module, conveniently enabling effective improvement of state-of-the-art methods for few-shot fine-grained image classification.
Shicheng Yang, Dongliang Chang, Zhanyu Ma, Jing-Hao Xue
ACM Multimedia5
2024 Customizing 360-Degree Panoramas through Text-to-Image Diffusion Models
abstract
Personalized text-to-image (T2I) synthesis based on diffusion models has attracted significant attention in recent research. However, existing methods primarily concentrate on customizing subjects or styles, neglecting the exploration of global geometry. In this study, we propose an approach that focuses on the customization of 360-degree panoramas, which inherently possess global geometric properties, using a T2I diffusion model. To achieve this, we curate a paired image-text dataset specifically designed for the task and subsequently employ it to fine-tune a pre-trained T2I diffusion model with LoRA. Nevertheless, the fine-tuned model alone does not ensure the continuity between the leftmost and rightmost sides of the synthesized images, a crucial characteristic of 360-degree panoramas. To address this issue, we propose a method called StitchDiffusion. Specifically, we perform pre-denoising operations twice at each time step of the denoising process on the stitch block consisting of the leftmost and rightmost image regions. Furthermore, a global cropping is adopted to synthesize seamless 360-degree panoramas. Experimental results demonstrate the effectiveness of our customized model combined with the proposed StitchDiffusion in generating high-quality 360-degree panoramic images. Moreover, our customized model exhibits exceptional generalization ability in producing scenes unseen in the fine-tuning dataset. Code is available at https://github.com/littlewhitesea/StitchDiffusion.
Hai Wang 0020, Xiaoyu Xiang, Yuchen Fan 0001, Jing-Hao Xue
WACV4
2024 DREAM: Visual Decoding from REversing HumAn Visual SysteM
abstract
In this work we present DREAM, an fMRI-to-image method for reconstructing viewed images from brain activities, grounded on fundamental knowledge of the human visual system. We craft reverse pathways that emulate the hierarchical and parallel nature of how humans perceive the visual world. These tailored pathways are specialized to decipher semantics, color, and depth cues from fMRI data, mirroring the forward pathways from visual stimuli to fMRI recordings. To do so, two components mimic the inverse processes within the human visual system: the Reverse Visual Association Cortex (R-VAC) which reverses pathways of this brain region, extracting semantics from fMRI data; the Reverse Parallel PKM (R-PKM) component simultaneously predicting color and depth from fMRI signals. The experiments indicate that our method outperforms the current state-of-the-art models in terms of the consistency of appearance, structure, and semantics. Code will be available at https://github.com/weihaox/DREAM.
Weihao Xia 0001, Raoul de Charette, A. Cengiz Öztireli, Jing-Hao Xue
WACV4
2024 A simple scheme to amplify inter-class discrepancy for improving few-shot fine-grained image classification
abstract
Few-shot image classification is a challenging topic in pattern recognition and computer vision. Few-shot fine-grained image classification is even more challenging, due to not only the few shots of labelled samples but also the subtle differences to distinguish subcategories in fine-grained images. A recent method called task discrepancy maximisation (TDM) can be embedded into the feature map reconstruction network (FRN) to generate discriminative features, by preserving the appearance details through reconstructing the query image and then assigning higher weights to more discriminative channels, producing the state-of-the-art performance for few-shot fine-grained image classification. However, due to the small inter-class discrepancy in fine-grained images and the small training set in few-shot learning, the training of FRN+TDM can result in excessively flexible boundaries between subcategories and hence overfitting. To resolve this problem, we propose a simple scheme to amplify inter-class discrepancy and thus improve FRN+TDM. To achieve this aim, instead of developing new modules, our scheme only involves two simple amendments to FRN+TDM: relaxing the inter-class score in TDM, and adding a centre loss to FRN. Extensive experiments on five benchmark datasets showcase that, although embarrassingly simple, our scheme is quite effective to improve the performance of few-shot fine-grained image classification. The code is available at https://github.com/Airgods/AFRN.git.
Zijie Guo, Rui Zhu 0006, Zhanyu Ma, Jun Guo 0002, Jing-Hao Xue
Pattern Recognit.6
2024 Self-reconstruction network for fine-grained few-shot classification
abstract
Metric-based methods are one of the most common methods to solve the problem of few-shot image classification. However, traditional metric-based few-shot methods suffer from overfitting and local feature misalignment. The recently proposed feature reconstruction-based approach, which reconstructs query image features from the support set features of a given class and compares the distance between the original query features and the reconstructed query features as the classification criterion, effectively solves the feature misalignment problem. However, the issue of overfitting still has not been considered. To this end, we propose a self-reconstruction metric module for diversifying query features and a restrained cross-entropy loss for avoiding over-confident predictions. By introducing them, the proposed self-reconstruction network can effectively alleviate overfitting. Extensive experiments on five benchmark fine-grained datasets demonstrate that our proposed method achieves state-of-the-art performance on both 5-way 1-shot and 5-way 5-shot classification tasks. Code is available at https://github.com/liz-lut/SRM-main.
Zhen Li 0026, Jiyang Xie 0001, Jing-Hao Xue, Zhanyu Ma
Pattern Recognit.5
2024 Query-centric distance modulator for few-shot classification
Wenxiao Wu, Yuanjie Shao, Changxin Gao, Jing-Hao Xue, Nong Sang
Pattern Recognit.4
2024 VPCFormer: A transformer-based multi-view finger vein recognition model and a new benchmark
Pengyang Zhao, Yizhuo Song, Jing-Hao Xue, Shuping Zhao, Qingmin Liao, Wenming Yang
Pattern Recognit.4
2024 Dual Correlation Network for Efficient Video Semantic Segmentation
abstract
Video data bring a big challenge to semantic segmentation due to the large volume of data and strong inter-frame redundancy. In this paper, we propose a dual local and global correlation network tailored for efficient video semantic segmentation. It consists of three modules: 1) a local attention based module, which measures correlation and achieves feature aggregation in a local region between key frame and non-key frame; 2) a consistent constraint module, which considers long-range correlation among pixels from a global view for promoting intra-frame semantic consistency of non-key frame; and 3) a key frame decision module, which selects key frames adaptively based on the ability of feature transferring. Extensive experiments on the Cityscapes and Camvid video datasets demonstrate that our proposed method could reduce inference time significantly while maintaining high accuracy. The implementation is available at https://github.com/An01168/DCNVSS.
Shumin An, Qingmin Liao, Zongqing Lu 0001, Jing-Hao Xue
IEEE Trans. Circuits Syst. Video Technol.4
2024 Interpretable Heterogeneous Teacher-Student Learning Framework for Hybrid-Supervised Pulmonary Nodule Detection
abstract
Existing pulmonary nodule detection methods often train models in a fully-supervised setting that requires strong labels (i.e., bounding box labels) as label information. However, manual annotation of bounding boxes in CT images is very time-consuming and labor-intensive. To alleviate the annotation burden, in this paper, we investigate pulmonary nodule detection by leveraging both strong labels and weak labels (i.e., center point labels) for training, and propose a novel hybrid-supervised pulmonary nodule detection (HND) method. The training of HND involves a heterogeneous teacher-student learning framework in two stages. In the first stage, we design a point-based consistency calibration network (PCC-Net) as a teacher, which is pre-trained to generate high-quality pseudo bounding box labels given point-augmented CT images as inputs. In the second stage, we develop an information bottleneck-guided pulmonary nodule detection network (IBD-Net) as a student to perform pulmonary nodule detection. In particular, we introduce information bottleneck to learn reliable pulmonary nodule-specific heatmaps under the guidance of PCC-Net, largely enhancing the model’s interpretability and improving the final detection performance. Based on the above designs, our method can effectively detect pulmonary nodule regions with only a limited number of bounding box labels. Experimental results on the public pulmonary nodule detection dataset LUNA16 show that our HND method achieves an excellent balance between the annotation cost and the detection performance.
Guangyu Huang, Yan Yan 0001, Jing-Hao Xue, Wentao Zhu 0002, Xióngbiao Luó
IEEE Trans. Circuits Syst. Video Technol.3
2024 Unpaired Caricature-Visual Face Recognition via Feature Decomposition-Restoration-Decomposition
abstract
Existing caricature-visual face recognition methods train the models based on caricature-visual image pairs from the same identities. Unfortunately, in many real-world applications, facial caricatures and visual facial images are usually unpaired in the training set due to the difficulty of collecting facial caricatures drawn by artists. In this paper, we study caricature-visual face recognition under the practical setting that only unpaired facial caricature and visual facial images are available as training samples, and define this setting as unpaired caricature-visual face recognition. To this end, we develop a novel feature decomposition-restoration-decomposition method (FDRD), which mainly consists of a backbone network, an identity-oriented feature decomposition module, and a modality-oriented feature restoration module, to extract modality-irrelevant identity features. To effectively train FDRD in the case of limited facial caricature training samples, we develop a two-stage learning framework. In the first stage, we perform single-modality restoration, enabling the model to have the basic ability of feature decomposition and restoration for each modality. In the second stage, we perform cross-modality recognition by exchanging new modality features between the two modalities, facilitating the model to focus on the decoupling of identity features and modality features. Experimental results demonstrate that our method performs favorably against several state-of-the-art face recognition methods and cross-modality methods. Our code is available at https://github.com/Capricorn-Karma/FDRD.
Yan Yan 0001, Jing-Hao Xue, Yang Hua 0001, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.3
2024 Dual-Mode Learning for Multi-Dataset X-Ray Security Image Detection
abstract
With the recent advance of deep learning, a large number of methods have been developed for prohibited item detection in X-ray security images. Generally, these methods train models on a single X-ray image dataset that may contain only limited categories of prohibited items. To detect more prohibited items, it is desirable to train a model on the multi-dataset that is constructed by combining multiple datasets. However, directly applying existing methods to the multi-dataset cannot guarantee good performance because of the large domain discrepancy between datasets and the occlusion in images. To address the above problems, we propose a novel Dual-Mode Learning Network (DML-Net) to effectively detect all the prohibited items in the multi-dataset. In particular, we develop an enhanced RetinaNet as the architecture of DML-Net, where we introduce a lattice appearance enhanced sub-net to enhance appearance representations. Such a way benefits the detection of occluded prohibited items. Based on the enhanced RetinaNet, the learning process of DML-Net involves both common mode learning (detecting the common prohibited items across datasets) and unique mode learning (detecting the unique prohibited items in each dataset). For common mode learning, we introduce an adversarial prototype alignment module to align the feature prototypes from different datasets in the domain-invariant feature space. For unique mode learning, we take advantage of feature distillation to enforce the student model to mimic the features extracted by multiple pre-trained teacher models. By tightly combining and jointly training the dual modes, our DML-Net method successfully eliminates the domain discrepancy and exhibits superior model capacity on the multi-dataset. Extensive experimental results on several combined X-ray image datasets demonstrate the effectiveness of our method against several state-of-the-art methods. Our code is available at https://github.com/vampirename/dmlnet.
Fenghong Yang, Runqing Jiang, Yan Yan 0001, Jing-Hao Xue, Hanzi Wang
IEEE Trans. Inf. Forensics Secur.4
2024 Relationship-Guided Knowledge Transfer for Class-Incremental Facial Expression Recognition
abstract
Human emotions contain both basic and compound facial expressions. In many practical scenarios, it is difficult to access all the compound expression categories at one time. In this paper, we investigate comprehensive facial expression recognition (FER) in the class-incremental learning paradigm, where we define well-studied and easily-accessible basic expressions as initial classes and learn new compound expressions incrementally. To alleviate the stability-plasticity dilemma in our incremental task, we propose a novel Relationship-Guided Knowledge Transfer (RGKT) method for class-incremental FER. Specifically, we develop a multi-region feature learning (MFL) module to extract fine-grained features for capturing subtle differences in expressions. Based on the MFL module, we further design a basic expression-oriented knowledge transfer (BET) module and a compound expression-oriented knowledge transfer (CET) module, by effectively exploiting the relationship across expressions. The BET module initializes the new compound expression classifiers based on expression relevance between basic and compound expressions, improving the plasticity of our model to learn new classes. The CET module transfers expression-generic knowledge learned from new compound expressions to enrich the feature set of old expressions, facilitating the stability of our model against forgetting old classes. Extensive experiments on three facial expression databases show that our method achieves superior performance in comparison with several state-of-the-art methods.
Yuanling Lv, Yan Yan 0001, Jing-Hao Xue, Si Chen 0002, Hanzi Wang
IEEE Trans. Image Process.3
2024 DeGCN: Deformable Graph Convolutional Networks for Skeleton-Based Action Recognition
abstract
Graph convolutional networks (GCN) have recently been studied to exploit the graph topology of the human body for skeleton-based action recognition. However, most of these methods unfortunately aggregate messages via an inflexible pattern for various action samples, lacking the awareness of intra-class variety and the suitableness for skeleton sequences, which often contain redundant or even detrimental connections. In this paper, we propose a novel Deformable Graph Convolutional Network (DeGCN) to adaptively capture the most informative joints. The proposed DeGCN learns the deformable sampling locations on both spatial and temporal graphs, enabling the model to perceive discriminative receptive fields. Notably, considering human action is inherently continuous, the corresponding temporal features are defined in a continuous latent space. Furthermore, we design an innovative multi-branch framework, which not only strikes a better trade-off between accuracy and model size, but also elevates the effect of ensemble between the joint and bone modalities remarkably. Extensive experiments show that our proposed method achieves state-of-the-art performances on three widely used datasets, NTU RGB+D, NTU RGB+D 120, and NW-UCLA.
Woomin Myung, Jing-Hao Xue, Guijin Wang
IEEE Trans. Image Process.3
2024 Visual-Textual Attribute Learning for Class-Incremental Facial Expression Recognition
abstract
In this paper, we study facial expression recognition (FER) in the class-incremental learning (CIL) setting, which defines the classification of well-studied and easily-accessible basic expressions as an initial task while learning new compound expressions gradually. Motivated by the fact that compound expressions are meaningful combinations of basic expressions, we treat basic expressions as attributes (i.e., semantic descriptors), and thus compound expressions are represented in terms of attributes. To this end, we propose a novel visual-textual attribute learning network (VTA-Net), mainly consisting of a textual-guided visual module (TVM) and a textual compositional module (TCM), for class-incremental FER. Specifically, TVM extracts textual-aware visual features and classifies expressions by incorporating the textual information into visual attribute learning. Meanwhile, TCM generates visual-aware textual features and predicts expressions by exploiting the dependency between textual attributes and category names of old and new expressions based on a textual compositional graph. In particular, a visual-textual distillation loss is introduced to calibrate TVM and TCM during incremental learning. Finally, the outputs from TVM and TCM are fused to make a final prediction. On the one hand, at each incremental task, the representations of visual attributes are enhanced since visual attributes are shared across old and new expressions. This increases the stability of our method. On the other hand, the textual modality, which involves rich prior knowledge of the relevance between expressions, facilitates our model to identify subtle visual distinctions between compound expressions, improving the plasticity of our method. Experimental results on both in-the-lab and in-the-wild facial expression databases show the superiority of our method against several state-of-the-art methods for class-incremental FER.
Yuanling Lv, Guangyu Huang, Yan Yan 0001, Jing-Hao Xue, Si Chen 0002, Hanzi Wang
IEEE Trans. Multim.4
2024 Exploit the Best of Both End-to-End and Map-Based Methods for Multi-Focus Image Fusion
abstract
Multi-focus image fusion is a technique to fuse the images focused on different depth ranges to generate an all-in-focus image. Existing deep learning approaches to multi-focus image fusion can be categorized as end-to-end methods and decision map based methods. End-to-end methods can generate natural fusion near the focus-defocus boundaries (FDB), but the output is often inconsistent with the input in the areas far from the boundaries (FFB). On the contrary, decision map based methods can preserve original images in the FFB areas, but often generate artifacts near the FDB. In this paper, we propose a dual-branch network for multi-focus image fusion (DB-MFIF) to exploit the best of both worlds, achieving better results in both FDB and FFB areas, i.e. with naturally sharper FDB areas and more consistent FFB areas with the inputs. In our DB-MFIF, an end-to-end branch and a decision map based branch are proposed to mutually assist each other. In addition, to this end, two map-based loss functions are also proposed. Experiments show that our method surpasses existing algorithms on multiple datasets, both qualitatively and quantitatively, and achieves the state-of-the-art performance. The code and model is available on GitHub:https://github.com/Zancelot/DB-MFIF.
Juncheng Zhang, Qingmin Liao, Jing-Hao Xue, Wenming Yang
IEEE Trans. Multim.4
2024 When Sparse Neural Network Meets Label Noise Learning: A Multistage Learning Framework
abstract
Recent methods in network pruning have indicated that a dense neural network involves a sparse subnetwork (called a winning ticket), which can achieve similar test accuracy to its dense counterpart with much fewer network parameters. Generally, these methods search for the winning tickets on well-labeled data. Unfortunately, in many real-world applications, the training data are unavoidably contaminated with noisy labels, thereby leading to performance deterioration of these methods. To address the above-mentioned problem, we propose a novel two-stream sample selection network (TS3-Net), which consists of a sparse subnetwork and a dense subnetwork, to effectively identify the winning ticket with noisy labels. The training of TS3-Net contains an iterative procedure that switches between training both subnetworks and pruning the smallest magnitude weights of the sparse subnetwork. In particular, we develop a multistage learning framework including a warm-up stage, a semisupervised alternate learning stage, and a label refinement stage, to progressively train the two subnetworks. In this way, the classification capability of the sparse subnetwork can be gradually improved at a high sparsity level. Extensive experimental results on both synthetic and real-world noisy datasets (including MNIST, CIFAR-10, CIFAR-100, ANIMAL-10N, Clothing1M, and WebVision) demonstrate that our proposed method achieves state-of-the-art performance with very small memory consumption for label noise learning. Code is available at https://github.com/Runqing-forMost/TS3-Net/tree/master.
Runqing Jiang, Yan Yan 0001, Jing-Hao Xue, Hanzi Wang
IEEE Trans. Neural Networks Learn. Syst.3
2024 Survival Analysis of High-Dimensional Data With Graph Convolutional Networks and Geometric Graphs
abstract
This article proposes a survival model based on graph convolutional networks (GCNs) with geometric graphs directly constructed from high-dimensional features. First, we clarify that the graphs used in GCNs play an important role in processing the relational information of samples, and the graphs that align well with the underlying data structure could be beneficial for survival analysis. Second, we show that sparse geometric graphs derived from high-dimensional data are more favorable compared with dense graphs when used in GCNs for survival analysis. Third, from this insight, we propose a model for survival analysis based on GCNs. By using multiple sparse geometric graphs and a proposed sequential forward floating selection algorithm, the new model is able to simultaneously perform survival analysis and unveil the local neighborhoods of samples. The experimental results on real-world datasets show that the proposed survival analysis approach based on GCNs outperforms a variety of existing methods and indicate that geometric graphs can aid survival analysis of high-dimensional data.
Yurong Ling, Zijing Liu, Jing-Hao Xue
IEEE Trans. Neural Networks Learn. Syst.3
2024 Toward Certified Robustness of Distance Metric Learning
abstract
Metric learning aims to learn a distance metric such that semantically similar instances are pulled together while dissimilar instances are pushed away. Many existing methods consider maximizing or at least constraining a distance margin in the feature space that separates similar and dissimilar pairs of instances to guarantee their generalization ability. In this article, we advocate imposing an adversarial margin in the input space so as to improve the generalization and robustness of metric learning algorithms. We first show that the adversarial margin, defined as the distance between training instances and their closest adversarial examples in the input space, takes account of both the distance margin in the feature space and the correlation between the metric and triplet constraints. Next, to enhance robustness to instance perturbation, we propose to enlarge the adversarial margin through minimizing a derived novel loss function termed the perturbation loss. The proposed loss can be viewed as a data-dependent regularizer and easily plugged into any existing metric learning methods. Finally, we show that the enlarged margin is beneficial to the generalization ability by using the theoretical technique of algorithmic robustness. Experimental results on 16 datasets demonstrate the superiority of the proposed method over existing state-of-the-art methods in both discrimination accuracy and robustness against possible noise.
Yiwen Guo, Mingzhi Dong, Jing-Hao Xue
IEEE Trans. Neural Networks Learn. Syst.4
2023 SPL-Net: Spatial-Semantic Patch Learning Network for Facial Attribute Recognition with Limited Labeled Data
Yan Yan 0001, Ying Shu, Si Chen 0002, Jing-Hao Xue, Chunhua Shen, Hanzi Wang
Int. J. Comput. Vis.4
2023 ReNAP: Relation network with adaptiveprototypical learning for few-shot classification
Yalan Li, Yixiao Zheng, Rui Zhu 0006, Zhanyu Ma, Jing-Hao Xue, Jie Cao 0014
Neurocomputing6
2023 Discriminant Feature Extraction by Generalized Difference Subspace
abstract
In this paper, we reveal the discriminant capacity of orthogonal data projection onto the generalized difference subspace (GDS), both theoretically and experimentally. In our previous work, we demonstrated that the GDS projection works as a quasi-orthogonalization of class subspaces, which is an effective feature extraction for subspace based classifiers. Here, we further show that GDS projection also works as a discriminant feature extraction through a similar mechanism to the Fisher discriminant analysis (FDA). A direct proof of the connection between GDS projection and FDA is difficult due to the significant difference in their formulations. To circumvent the complication, we first introduce geometrical Fisher discriminant analysis (gFDA) based on a simplified Fisher criterion. It is derived from a heuristic yet practically plausible assumption: the direction of the sample mean vector of a class is largely aligned to the first principal component vector of the class, given that the principal component analysis (PCA) is applied without data centering. gFDA works stably even under few samples, bypassing the small sample size (SSS) problem of FDA. We then prove that gFDA is equivalent to GDS projection with a small correction term. This equivalence ensures GDS projection to inherit the discriminant ability from FDA via gFDA. Furthermore, we discuss two useful extensions of these methods, 1) a nonlinear extension by kernel trick, 2) a combination with CNN features. The equivalence and the effectiveness of the extensions have been verified through extensive experiments on the extended Yale B+, CMU face database, ALOI, ETH80, MNIST, and CIFAR10, mainly focusing on image recognition under small samples.
Kazuhiro Fukui, Naoya Sogi, Takumi Kobayashi 0001, Jing-Hao Xue, Atsuto Maki
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 GAN Inversion: A Survey
abstract
GAN inversion aims to invert a given image back into the latent space of a pretrained GAN model so that the image can be faithfully reconstructed from the inverted code by the generator. As an emerging technique to bridge the real and fake image domains, GAN inversion plays an essential role in enabling pretrained GAN models, such as StyleGAN and BigGAN, for applications of real image editing. Moreover, GAN inversion interprets GAN's latent space and examines how realistic images can be generated. In this paper, we provide a survey of GAN inversion with a focus on its representative algorithms and its applications in image restoration and image manipulation. We further discuss the trends and challenges for future research. A curated list of GAN inversion methods, datasets, and other related information can be found at https://github.com/weihaox/awesome-gan-inversion.
Weihao Xia 0001, Yulun Zhang 0001, Yujiu Yang 0001, Jing-Hao Xue, Bolei Zhou, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Deep metric learning for few-shot image classification: A Review of recent developments
Zhanyu Ma, Jing-Hao Xue
Pattern Recognit.4
2023 The neglected background cues can facilitate finger vein recognition
Pengyang Zhao, Shuping Zhao, Jing-Hao Xue, Wenming Yang, Qingmin Liao
Pattern Recognit.3
2023 Statistical hypothesis testing as a novel perspective of pooling for image quality assessment
abstract
Image quality assessment is usually achieved by pooling local quality scores. However, commonly used pooling strategies, based on simple sample statistics, are not always sensitive to distortions. In this short communication, we propose a novel perspective of pooling: reliable pooling through statistical hypothesis testing, which enables effective detection of subtle changes of population parameters when the underlying distribution of local quality scores is affected by distortions. To illustrate the significance of this novel perspective, we design a new pooling strategy utilising simple one-sided one-sample t-test. The experiments on benchmark databases show the reliability of hypothesis testing-based pooling, compared with state-of-the-art pooling strategies.
Rui Zhu 0006, Fei Zhou 0001, Wenming Yang, Jing-Hao Xue
Signal Process. Image Commun.4
2023 Locally-Enriched Cross-Reconstruction for Few-Shot Fine-Grained Image Classification
abstract
Few-shot fine-grained image classification has attracted considerable attention in recent years for its realistic setting to imitate how humans conduct recognition tasks. Metric-based few-shot classifiers have achieved high accuracies. However, their metric function usually requires two arguments of vectors, while transforming or reshaping three-dimensional feature maps to vectors can result in loss of spatial information. Image reconstruction is thus involved to retain more appearance details: the test images are reconstructed by different classes and then classified to the one with the smallest reconstruction error. However, discriminative local information, vital to distinguish sub-categories in fine-grained images with high similarities, is not well elaborated when only the base features from a usual embedding module are adopted for reconstruction. Hence, we propose the novel local content-enriched cross-reconstruction network (LCCRN) for few-shot fine-grained classification. In LCCRN, we design two new modules: the local content-enriched module (LCEM) to learn the discriminative local features, and the cross-reconstruction module (CRM) to fully engage the local features with the appearance details obtained from a separate embedding module. The classification score is calculated based on the weighted sum of reconstruction errors of the cross-reconstruction tasks, with weights learnt from the training process. Extensive experiments on four fine-grained datasets showcase the superior classification performance of LCCRN compared with the state-of-the-art few-shot classification methods. Codes are available at:https://github.com/lutsong/LCCRN.
Jijie Wu, Rui Zhu 0006, Zhanyu Ma, Jing-Hao Xue
IEEE Trans. Circuits Syst. Video Technol.6
2023 TROSD: A New RGB-D Dataset for Transparent and Reflective Object Segmentation in Practice
abstract
Transparent and reflective objects are omnipresent in our daily life, but their unique visual and optical characteristics are notoriously challenging even for state-of-the-art deep networks of semantic segmentation. To alleviate this challenge, we construct a new large-scale real-world RGB-D dataset called TROSD, which is more comprehensive than existing datasets for transparent and reflective object segmentation. Our TROSD dataset contains 11,060 RGB-D images with three semantic classes in terms of transparent objects, reflective objects, and others, covering a variety of daily scenes. Together with the dataset, we also introduce a novel network (TROSNet) as a high-standard baseline to assist other researchers to develop and benchmark their algorithms of transparent and reflective object segmentation. Moreover, extensive experiments also clearly show that the proposed TROSD dataset has an excellent capacity to facilitate the development of semantic segmentation algorithms with strong generalizability.
Guodong Zhang 0004, Wenming Yang, Jing-Hao Xue, Guijin Wang
IEEE Trans. Circuits Syst. Video Technol.4
2023 Drop Loss for Person Attribute Recognition With Imbalanced Noisy-Labeled Samples
abstract
Person attribute recognition (PAR) aims to simultaneously predict multiple attributes of a person. Existing deep learning-based PAR methods have achieved impressive performance. Unfortunately, these methods usually ignore the fact that different attributes have an imbalance in the number of noisy-labeled samples in the PAR training datasets, thus leading to suboptimal performance. To address the above problem of imbalanced noisy-labeled samples, we propose a novel and effective loss called drop loss for PAR. In the drop loss, the attributes are treated differently in an easy-to-hard way. In particular, the noisy-labeled candidates, which are identified according to their gradient norms, are dropped with a higher drop rate for the harder attribute. Such a manner adaptively alleviates the adverse effect of imbalanced noisy-labeled samples on model learning. To illustrate the effectiveness of the proposed loss, we train a simple ResNet-50 model based on the drop loss and term it DropNet. Experimental results on two representative PAR tasks (including facial attribute recognition and pedestrian attribute recognition) demonstrate that the proposed DropNet achieves comparable or better performance in terms of both balanced accuracy and classification accuracy over several state-of-the-art PAR methods.
Yan Yan 0001, Youze Xu, Jing-Hao Xue, Yang Lu 0009, Hanzi Wang, Wentao Zhu 0002
IEEE Trans. Cybern.3
2023 APANet: Adaptive Prototypes Alignment Network for Few-Shot Semantic Segmentation
abstract
Few-shot semantic segmentation aims to segment novel-class objects in a given query image with only a few labeled support images. Most advanced solutions exploit a metric learning framework that performs segmentation through matching each query feature to a learned class-specific prototype. However, this framework suffers from biased classification due to incomplete feature comparisons. To address this issue, we present an adaptive prototype representation by introducing class-specific and class-agnostic prototypes and thus construct complete sample pairs for learning semantic alignment with query features. The complementary features learning manner effectively enriches feature comparison and helps yield an unbiased segmentation model in the few-shot setting. It is implemented with a two-branch end-to-end network (i.e., a class-specific branch and a class-agnostic branch), which generates prototypes and then combines query features to perform comparisons. In addition, the proposed class-agnostic branch is simple yet effective. In practice, it can adaptively generate multiple class-agnostic prototypes for query images and learn feature alignment in a self-contrastive manner. Extensive experiments on PASCAL-5$^{i}$and COCO-20$^{i}$demonstrate the superiority of our method. At no expense of inference efficiency, our model achieves state-of-the-art results in both 1-shot and 5-shot settings for semantic segmentation.
Bin-Bin Gao, Zongqing Lu 0001, Jing-Hao Xue, Chengjie Wang 0001, Qingmin Liao
IEEE Trans. Multim.4
2023 On the Comparisons of Decorrelation Approaches for Non-Gaussian Neutral Vector Variables
abstract
-norm equals one. In addition, its neutral properties make it significantly different from the commonly studied vector variables (e.g., the Gaussian vector variables). Due to the aforementioned properties, the conventionally applied linear transformation approaches [e.g., principal component analysis (PCA) and independent component analysis (ICA)] are not suitable for neutral vector variables, as PCA cannot transform a neutral vector variable, which is highly negatively correlated, into a set of mutually independent scalar variables and ICA cannot preserve the bounded property after transformation. In recent work, we proposed an efficient nonlinear transformation approach, i.e., the parallel nonlinear transformation (PNT), for decorrelating neutral vector variables. In this article, we extensively compare PNT with PCA and ICA through both theoretical analysis and experimental evaluations. The results of our investigations demonstrate the superiority of PNT for decorrelating the neutral vector variables.
Zhanyu Ma, Xiaoou Lu, Jiyang Xie 0001, Zhen Yang 0004, Jing-Hao Xue, Zheng-Hua Tan, Bo Xiao 0006, Jun Guo 0002
IEEE Trans. Neural Networks Learn. Syst.5
2022 When Facial Expression Recognition Meets Few-Shot Learning: A Joint and Alternate Learning Framework
abstract
Human emotions involve basic and compound facial expressions. However, current research on facial expression recognition (FER) mainly focuses on basic expressions, and thus fails to address the diversity of human emotions in practical scenarios. Meanwhile, existing work on compound FER relies heavily on abundant labeled compound expression training data, which are often laboriously collected under the professional instruction of psychology. In this paper, we study compound FER in the cross-domain few-shot learning setting, where only a few images of novel classes from the target domain are required as a reference. In particular, we aim to identify unseen compound expressions with the model trained on easily accessible basic expression datasets. To alleviate the problem of limited base classes in our FER task, we propose a novel Emotion Guided Similarity Network (EGS-Net), consisting of an emotion branch and a similarity branch, based on a two-stage learning framework. Specifically, in the first stage, the similarity branch is jointly trained with the emotion branch in a multi-task fashion. With the regularization of the emotion branch, we prevent the similarity branch from overfitting to sampled base classes that are highly overlapped across different episodes. In the second stage, the emotion branch and the similarity branch play a “two-student game” to alternately learn from each other, thereby further improving the inference ability of the similarity branch on unseen compound expressions. Experimental results on both in-the-lab and in-the-wild compound expression datasets demonstrate the superiority of our proposed method against several state-of-the-art methods.
Xinyi Zou, Yan Yan 0001, Jing-Hao Xue, Si Chen 0002, Hanzi Wang
AAAI3
2022 Learn-to-Decompose: Cascaded Decomposition Network for Cross-Domain Few-Shot Facial Expression Recognition
Xinyi Zou, Yan Yan 0001, Jing-Hao Xue, Si Chen 0002, Hanzi Wang
ECCV (19)3
2022 Dimension reduction for high-dimensional small counts with KL divergence
abstract
Dimension reduction for high-dimensional count data with a large proportion of zeros is an important task in various applications. As a large number of dimension reduction methods rely on the proximity measure, we develop a dissimilarity measure that is well-suited for small counts based on the Kullback-Leibler divergence. We compare the proposed measure with other widely used dissimilarity measures and show that the proposed one has superior discriminative ability when applied to high-dimensional count data having an excess of zeros. Extensive empirical results, on both simulated and publicly-available real-world datasets that contain many zeros, demonstrate that the proposed dissimilarity measure can improve a wide range of dimension reduction methods.
Yurong Ling, Jing-Hao Xue
UAI2
2022 Adaptive Deep Disturbance-Disentangled Learning for Facial Expression Recognition
Delian Ruan, Rongyun Mo, Yan Yan 0001, Si Chen 0002, Jing-Hao Xue, Hanzi Wang
Int. J. Comput. Vis.5
2022 ASSP: An adaptive sample statistics-based pooling for full-reference image quality assessment
Yurong Ling, Fei Zhou 0001, Kun Guo 0004, Jing-Hao Xue
Neurocomputing4
2022 A Spectral-Spatial Jointed Spectral Super-Resolution and Its Application to HJ-1A Satellite Images
abstract
To generate a high-spatial-resolution hyperspectral (HHS) image from a high-spatial-resolution multispectral (HMS) image, both spatial information and spectral information should be considered simultaneously if we want to build a more accurate mapping from HMS to HHS. To this end, a spectral and spatial jointed spectral super-resolution method is proposed in this letter using an end-to-end learning strategy for each subspace with the cluster-based multibranch backpropagation neural network (BPNN). More specifically, in addition to the spectra similarity, a modified superpixel segmentation is introduced to jointly take spatial contextual information into account, and a new framework with it is given. Comparisons on the Columbia University Automated Vision Environment (CAVE) data set show that our proposed method outperforms other relative state-of-the-art methods more than 0.3 in the root mean squared error (RMSE) and more than 1.0 in the spectral angle mapper (SAM) index. Especially, an exemplary application is demonstrated using the synchronized observation data collected by the multispectral and hyperspectral sensors mounted on the HJ-1A satellite at the same time.
Xiaolin Han 0001, Huan Zhang 0013, Jing-Hao Xue
IEEE Geosci. Remote. Sens. Lett.3
2022 Advanced Dropout: A Model-Free Methodology for Bayesian Dropout Optimization
abstract
Due to lack of data, overfitting ubiquitously exists in real-world applications of deep neural networks (DNNs). We propose advanced dropout, a model-free methodology, to mitigate overfitting and improve the performance of DNNs. The advanced dropout technique applies a model-free and easily implemented distribution with parametric prior, and adaptively adjusts dropout rate. Specifically, the distribution parameters are optimized by stochastic gradient variational Bayes in order to carry out an end-to-end training. We evaluate the effectiveness of the advanced dropout against nine dropout techniques on seven computer vision datasets (five small-scale datasets and two large-scale datasets) with various base models. The advanced dropout outperforms all the referred techniques on all the datasets. We further compare the effectiveness ratios and find that advanced dropout achieves the highest one on most cases. Next, we conduct a set of analysis of dropout rate characteristics, including convergence of the adaptive dropout rate, the learned distributions of dropout masks, and a comparison with dropout rate generation without an explicit distribution. In addition, the ability of overfitting prevention is evaluated and confirmed. Finally, we extend the application of the advanced dropout to uncertainty inference, network pruning, text classification, and regression. The proposed advanced dropout is also superior to the corresponding referred methods. Codes are available at https://github.com/PRIS-CV/AdvancedDropout.
Jiyang Xie 0001, Zhanyu Ma, Jianjun Lei 0001, Guoqiang Zhang 0003, Jing-Hao Xue, Zheng-Hua Tan, Jun Guo 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Constrained mutual convex cone method for image set based recognition
Naoya Sogi, Rui Zhu 0006, Jing-Hao Xue, Kazuhiro Fukui
Pattern Recognit.3
2022 Deep Multi-Task Multi-Label CNN for Effective Facial Attribute Classification
abstract
Facial Attribute Classification (FAC) has attracted increasing attention in computer vision and pattern recognition. However, state-of-the-art FAC methods perform face detection/alignment and FAC independently. The inherent dependencies between these tasks are not fully exploited. In addition, most methods predict all facial attributes using the same CNN network architecture, which ignores the different learning complexities of facial attributes. To address the above problems, we propose a novel deep multi-task multi-label CNN, termed DMM-CNN, for effective FAC. Specifically, DMM-CNN jointly optimizes two closely-related tasks (i.e., facial landmark detection and FAC) to improve the performance of FAC by taking advantage of multi-task learning. To deal with the diverse learning complexities of facial attributes, we divide the attributes into two groups: objective attributes and subjective attributes. Two different network architectures are respectively designed to extract features for two groups of attributes, and a novel dynamic weighting scheme is proposed to automatically assign the loss weight to each facial attribute during training. Furthermore, an adaptive thresholding strategy is developed to effectively alleviate the problem of class imbalance for multi-label learning. Experimental results on the challenging CelebA and LFWA datasets show the superiority of the proposed DMM-CNN method compared with several state-of-the-art FAC methods.
Longbiao Mao, Yan Yan 0001, Jing-Hao Xue, Hanzi Wang
IEEE Trans. Affect. Comput.3
2022 Stage-Aware Feature Alignment Network for Real-Time Semantic Segmentation of Street Scenes
abstract
Over the past few years, deep convolutional neural network-based methods have made great progress in semantic segmentation of street scenes. Some recent methods align feature maps to alleviate the semantic gap between them and achieve high segmentation accuracy. However, they usually adopt the feature alignment modules with the same network configuration in the decoder and thus ignore the different roles of stages of the decoder during feature aggregation, leading to a complex decoder structure. Such a manner greatly affects the inference speed. In this paper, we present a novel Stage-aware Feature Alignment Network (SFANet) based on the encoder-decoder structure for real-time semantic segmentation of street scenes. Specifically, a Stage-aware Feature Alignment module (SFA) is proposed to align and aggregate two adjacent levels of feature maps effectively. In the SFA, by taking into account the unique role of each stage in the decoder, a novel stage-aware Feature Enhancement Block (FEB) is designed to enhance spatial details and contextual information of feature maps from the encoder. In this way, we are able to address the misalignment problem with a very simple and efficient multi-branch decoder structure. Moreover, an auxiliary training strategy is developed to explicitly alleviate the multi-scale object problem without bringing additional computational costs during the inference phase. Experimental results show that the proposed SFANet exhibits a good balance between accuracy and speed for real-time semantic segmentation of street scenes. In particular, based on ResNet-18, SFANet respectively obtains 78.1% and 74.7% mean of class-wise Intersection-over-Union (mIoU) at inference speeds of 37 FPS and 96 FPS on the challenging Cityscapes and CamVid test datasets by using only a single GTX 1080Ti GPU.
Xi Weng, Yan Yan 0001, Si Chen 0002, Jing-Hao Xue, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.4
2022 SSL++: Improving Self-Supervised Learning by Mitigating the Proxy Task-Specificity Problem
abstract
The success of deep convolutional networks (ConvNets) generally relies on a massive amount of well-labeled data, which is labor-intensive and time-consuming to collect and annotate in many scenarios. To eliminate such limitation, self-supervised learning (SSL) is recently proposed. Specifically, by solving a pre-designed proxy task, SSL is capable of capturing general-purpose features without requiring human supervision. Existing efforts focus obsessively on designing a particular proxy task but ignore the semanticity of samples that are advantageous to downstream tasks, resulting in the inherent limitation that the learned features are specific to the proxy task, namely the proxy task-specificity of features. In this work, to improve the generalizability of features learned by existing SSL methods, we present a novel self-supervised framework SSL++ to incorporate the proxy task-independent semanticity of samples into the representation learning process. Technically, SSL++ aims to leverage the complementarity, between the low-level generic features learned by a proxy task and the high-level semantic features newly learned by the generated semantic pseudo-labels, to mitigate the task-specificity and improve the generalizability of features. Extensive experiments show that SSL++ performs favorably against the state-of-the-art approaches on the established and latest SSL benchmarks.
Jing-Hao Xue, Jianlong Chang, Jianzhong Zhang 0003, Jufeng Yang, Qi Tian 0001
IEEE Trans. Image Process.2
2022 Defocus Image Deblurring Network With Defocus Map Estimation as Auxiliary Task
abstract
Different from the object motion blur, the defocus blur is caused by the limitation of the cameras' depth of field. The defocus amount can be characterized by the parameter of point spread function and thus forms a defocus map. In this paper, we propose a new network architecture called Defocus Image Deblurring Auxiliary Learning Net (DID-ANet), which is specifically designed for single image defocus deblurring by using defocus map estimation as auxiliary task to improve the deblurring result. To facilitate the training of the network, we build a novel and large-scale dataset for single image defocus deblurring, which contains the defocus images, the defocus maps and the all-sharp images. To the best of our knowledge, the new dataset is the first large-scale defocus deblurring dataset for training deep networks. Moreover, the experimental results demonstrate that the proposed DID-ANet outperforms the state-of-the-art methods for both tasks of defocus image deblurring and defocus map estimation, both quantitatively and qualitatively. The dataset, code, and model is available on GitHub: https://github.com/xytmhy/DID-ANet-Defocus-Deblurring.
Qingmin Liao, Juncheng Zhang, Jing-Hao Xue
IEEE Trans. Image Process.5
2022 Efficient Semantic Segmentation via Self-Attention and Self-Distillation
abstract
Lightweight models are pivotal in efficient semantic segmentation, but they often suffer from insufficient context information due to limited convolution and small receptive field. To address this problem, we propose a tailored approach to efficient semantic segmentation by leveraging two complementary distillation schemes for supplementing context information to small networks: 1) a self-attention distillation scheme, which transfers long-range context knowledge adaptively from large teacher networks to small student networks; and 2) a layer-wise context distillation scheme, which transfers structured context from deep layers to shallow layers within student networks for promoting semantic consistency of the shallow layers. Extensive experiments on the ADE20K, Cityscapes, and Camvid datasets well demonstrate the effectiveness of our proposal.
Shumin An, Qingmin Liao, Zongqing Lu 0001, Jing-Hao Xue
IEEE Trans. Intell. Transp. Syst.4
2022 Deep Learning in Lane Marking Detection: A Survey
abstract
Lane marking detection is a fundamental but crucial step in intelligent driving systems. It can not only provide relevant road condition information to prevent lane departure but also assist vehicle positioning and forehead car detection. However, lane marking detection faces many challenges, including extreme lighting, missing lane markings, and obstacle obstructions. Recently, deep learning-based algorithms draw much attention in intelligent driving society because of their excellent performance. In this paper, we review deep learning methods for lane marking detection, focusing on their network structures and optimization objectives, the two key determinants of their success. Besides, we summarize existing lane-related datasets, evaluation criteria, and common data processing techniques. We also compare the detection performance and running time of various methods, and conclude with some current challenges and future trends for deep learning-based lane marking detection algorithm.
Youcheng Zhang, Zongqing Lu 0001, Xuechen Zhang 0003, Jing-Hao Xue, Qingmin Liao
IEEE Trans. Intell. Transp. Syst.4
2022 GenDet: Meta Learning to Generate Detectors From Few Shots
abstract
Object detection has made enormous progress and has been widely used in many applications. However, it performs poorly when only limited training data is available for novel classes that the model has never seen before. Most existing approaches solve few-shot detection tasks implicitly without directly modeling the detectors for novel classes. In this article, we propose GenDet, a new meta-learning-based framework that can effectively generate object detectors for novel classes from few shots and, thus, conducts few-shot detection tasks explicitly. The detector generator is trained by numerous few-shot detection tasks sampled from base classes each with sufficient samples, and thus, it is expected to generalize well on novel classes. An adaptive pooling module is further introduced to suppress distracting samples and aggregate the detectors generated from multiple shots. Moreover, we propose to train a reference detector for each base class in the conventional way, with which to guide the training of the detector generator. The reference detectors and the detector generator can be trained simultaneously. Finally, the generated detectors of different classes are encouraged to be orthogonal to each other for better generalization. The proposed approach is extensively evaluated on the ImageNet, VOC, and COCO data sets under various few-shot detection settings, and it achieves new state-of-the-art results.
Liyang Liu, Bochao Wang, Zhanghui Kuang, Jing-Hao Xue, Wenming Yang, Qingmin Liao, Wayne Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2021 Amortized Bayesian Prototype Meta-learning: A New Probabilistic Meta-learning Approach to Few-shot Image Classification
abstract
Probabilistic meta-learning methods recently have achieved impressive success in few-shot image classification. However, they introduce a huge number of random variables for neural network weights and thus severe computational and inferential challenges. In this paper, we propose a novel probabilistic meta-learning method called amortized Bayesian prototype meta-learning. In contrast to previous methods, we introduce only a small number of random variables for latent class prototypes rather than a huge number for network weights; we learn to learn the posterior distributions of these latent prototypes in an amortized inference way with no need for an extra amortization network, such that we can easily approximate their posteriors conditional on few labeled samples, whenever at meta-training or meta-testing stage. The proposed method can be trained end-to-end without any pre-training. Compared with other probabilistic meta-learning methods, our proposed approach is more interpretable with much less random variables, while still be able to achieve competitive performance for few-shot image classification problems on various benchmark datasets. Its excellent robustness and predictive uncertainty are also demonstrated through ablation studies.
Jijie Wu, Wenming Yang, Jing-Hao Xue
AISTATS5
2021 Lane Line Detection based on Parallel Spatial Separation Convolution
Xile Shen, Zongqing Lu 0001, Youcheng Zhang, Jing-Hao Xue
BMVC4
2021 FETNet: Feature Exchange Transformer Network for RGB-D Object Detection
Jing-Hao Xue, Pengwei Xie, Guijin Wang
BMVC2
2021 Learning Spatial-Semantic Relationship for Facial Attribute Recognition With Limited Labeled Data
abstract
Recent advances in deep learning have demonstrated excellent results for Facial Attribute Recognition (FAR), typically trained with large-scale labeled data. However, in many real-world FAR applications, only limited labeled data are available, leading to remarkable deterioration in performance for most existing deep learning-based FAR methods. To address this problem, here we propose a method termed Spatial-Semantic Patch Learning (SSPL). The training of SSPL involves two stages. First, three auxiliary tasks, consisting of a Patch Rotation Task (PRT), a Patch Segmentation Task (PST), and a Patch Classification Task (PCT), are jointly developed to learn the spatial-semantic relationship from large-scale unlabeled facial data. We thus obtain a powerful pre-trained model. In particular, PRT exploits the spatial information of facial images in a self-supervised learning manner. PST and PCT respectively capture the pixel-level and image-level semantic information of facial images based on a facial parsing model. Second, the spatial-semantic knowledge learned from auxiliary tasks is transferred to the FAR task. By doing so, it enables that only a limited number of labeled data are required to fine-tune the pre-trained model. We achieve superior performance compared with state-of-the-art methods, as substantiated by extensive experiments and studies.
Ying Shu, Yan Yan 0001, Si Chen 0002, Jing-Hao Xue, Chunhua Shen, Hanzi Wang
CVPR4
2021 TediGAN: Text-Guided Diverse Face Image Generation and Manipulation
abstract
In this work, we propose TediGAN, a novel framework for multi-modal image generation and manipulation with textual descriptions. The proposed method consists of three components: StyleGAN inversion module, visual-linguistic similarity learning, and instance-level optimization. The inversion module maps real images to the latent space of a well-trained StyleGAN. The visual-linguistic similarity learns the text-image matching by mapping the image and text into a common embedding space. The instancelevel optimization is for identity preservation in manipulation. Our model can produce diverse and high-quality images with an unprecedented resolution at 10242. Using a control mechanism based on style-mixing, our TediGAN inherently supports image synthesis with multi-modal inputs, such as sketches or semantic labels, with or without instance guidance. To facilitate text-guided multi-modal synthesis, we propose the Multi-Modal CelebA-HQ, a large-scale dataset consisting of real face images and corresponding semantic segmentation map, sketch, and textual descriptions. Extensive experiments on the introduced dataset demonstrate the superior performance of our proposed method. Code and data are available at https://github.com/weihaox/TediGAN.
Weihao Xia 0001, Yujiu Yang 0001, Jing-Hao Xue, Baoyuan Wu
CVPR3
2021 Enhance Via Decoupling: Improving Multi-Label Classifiers With Variational Feature Augmentation
abstract
Multi-label classification remains a challenging problem due to the inherent label imbalance issue, which brings overfitting of minor categories to modern deep models. In this paper, to tackle this issue, we propose a novel method named Variational Feature Augmentation (VFA) to enhance the deep neural networks for multi-label classification. Our method decouples the feature vectors extracted by the backbone network into multiple low-dimensional spaces via a novely proposed Variational Feature Decoupling Module. The decoupled feature vectors are then re-combined with a shuffle operation and a Feature Augmentation Layer to enrich the minor co-occurrence relations, mitigating the label imbalance. Different from most other methods, VFA does not modify the network architecture or introduce extra computation cost in inference phase. We conduct comprehensive experiments on four benchmarks of two visual multi-label classification tasks, pedestrian attribute recognition and multi-label image recognition, and the results demonstrate the effectiveness and generality of the proposed VFA.
Guijin Wang, Jing-Hao Xue, Zijian Ding
ICIP3
2021 Parallax Contextual Representations For Stereo Matching
abstract
In this work, we study the context aggregation in stereo matching from a new parallax perspective. Unlike previous works, we propose to characterize and augment a pixel with its parallax contextual representation (PCR), which has not been explored before. We also propose a new concept called disparity prototype to describe the overall representation of a disparity plane. Our proposed PCR module consists of three steps: 1) divide disparity planes for a rough estimation of disparity; 2) estimate the disparity prototypes for each disparity plane; 3) derive PCR-augmented representations with disparity prototypes. Extensive experiments on various datasets using different networks validate the effectiveness of our proposal.
Qingmin Liao, Zongqing Lu 0001, Jing-Hao Xue
ICIP4
2021 CLUSAC: Clustering Sample Consensus for Fundamental Matrix Estimation
abstract
In the process of model fitting for fundamental matrix estimation, RANSAC and its variants disregard and fail to reduce the interference of outliers. These methods select correspondences and calculate the model scores from the original dataset. In this work, we propose an inlier filtering method that can filter inliers from the original dataset. Using the filtered inliers can substantially reduce the interference of outliers. Based on the filtered inliers, we propose a new algorithm called CLUSAC, which calculates model quality scores on all filtered inliers. Our approach is evaluated through estimating the fundamental matrix in the dataset kusvod2, and it shows superior performance to other compared RANSAC variants in terms of precision.
Xuanyu Xiao, Zongqing Lu 0001, Jing-Hao Xue
ICIP3
2021 Towards Impartial Multi-task Learning
Liyang Liu, Yi Li 0050, Zhanghui Kuang, Jing-Hao Xue, Wenming Yang, Qingmin Liao, Wayne Zhang 0001
ICLR4
2021 Disparity Estimation with Scene Depth Cues
abstract
The cost volume plays a pivotal role in stereo matching, usually working as an optimization object. However, we find it also can provide effective scene prior to guide the disparity learning, as it reflects well the depth relationship between scenario objects. Inspired by this new perspective, we propose the CSA module, which consists of a new correlation and selection (CS) layer and a new aggregation layer. The CS layer can regulate the matching costs and re-encode the feature information into the correlation volume. The aggregation layer can preserve better the depth cues of the refined cost volume, through a convolution network and a unimodalization operation. The proposed module can be trained in a supervised manner, making the extraction of scene depth cues more accurate. Extensive experiments on the Sceneflow and KITTI datasets have demonstrated that with our module embedded, SOTA networks can achieve substantially better performance.
Zongqing Lu 0001, Qingmin Liao, Jing-Hao Xue
ICME5
2021 Better Stereo Matching From Simple Yet Effective Wrangling of Deep Features
abstract
Cost volume plays a pivotal role in stereo matching. Most recent works focused on deep feature extraction and cost refinement for a more accurate cost volume. Unlike them, we probe from a different perspective: feature wrangling. We find that simple wrangling of deep features can effectively improve the construction of cost volume and thus the performance of stereo matching. Specifically, we develop two simple yet effective wrangling techniques of deep features, spatially a differentiable feature transformation and channel-wise a memory-economical feature expansion, for better cost construction. Exploiting the local ordering information provided by a differentiable rank transform, we achieve an enhancement of the search for correspondence; with the help of disparity division, our feature expansion allows for more features into the cost volume with no extra memory required. Equipped with these two feature wrangling techniques, our simple network can perform outstandingly on the widely used KITTI and Sceneflow datasets.
Zongqing Lu 0001, Qingmin Liao, Jing-Hao Xue
ICME4
2021 RGB Guided Depth Map Super-Resolution with Coupled U-Net
abstract
The depth maps captured by RGB-D cameras usually are of low resolution, entailing recent efforts to develop depth super-resolution (DSR) methods. However, several problems remain in existing DSR methods. First, conventional DSR methods often suffer from unexpected artifacts. Secondly, high-resolution (HR) RGB features and low-resolution (LR) depth features are often fused in shallow layers only. Thirdly, only the last layer of features is used for reconstruction. To address the above problems, we propose Coupled U-Net (CU-Net), a new color image guided DSR method built on two U-Net branches for HR color images and LR depth maps, respectively. The CU-Net embeds a dual skip connection structure to leverage the feature interaction of the two branches, and a multi-scale fusion to fuse the deeper and multi-scale features of two branch decoders for more effective feature reconstruction. Moreover, a channel attention module is proposed to eliminate artifacts. Extensive experiments show that the proposed CU-Net outperforms state-of-the-art methods.
Yingjie Cui, Qingmin Liao, Wenming Yang, Jing-Hao Xue
ICME4
2021 EFRNet: A Lightweight Network with Efficient Feature Fusion and Refinement for Real-Time Semantic Segmentation
abstract
To pursue high accuracy, most image semantic segmentation methods are computationally costly and thus not suitable to real-time applications. Existing lightweight methods either adopt a single branch without feature fusion, which dam-ages accuracy, or introduce extra branches for feature fusion, which harms efficiency. In this paper, we propose a lightweight network named EFRNet, with feature fusion and refinement in a single branch to achieve better balance between accuracy and efficiency in real-time semantic segmentation. Specifically, in EFRNet, we design a novel Feature Fusion Module to fuse multi-stage features in a single CNN efficiently, and we propose a lightweight Channel Attention Refinement Module to refine features with few extra parameters. Extensive experiments show that our EFRNet achieves decent accuracy with an extremely small model size and high inference speed. It achieves the best accuracy of 70.02% mIoU compared with state-of-the-art lightweight methods on CamVid with only 0.48M parameters.
Kuayue Zhang, Qingmin Liao, Juncheng Zhang, Jing-Hao Xue
ICME6
2021 Group Fisher Pruning for Practical Network Compression
abstract
Network compression has been widely studied since it is able to reduce the memory and computation cost during inference. However, previous methods seldom deal with complicated structures like residual connections, group/depth-wise convolution and feature pyramid network, where channels of multiple layers are coupled and need to be pruned simultaneously. In this paper, we present a general channel pruning approach that can be applied to various complicated structures. Particularly, we propose a layer grouping algorithm to find coupled channels automatically. Then we derive a unified metric based on Fisher information to evaluate the importance of a single channel and coupled channels. Moreover, we find that inference speedup on GPUs is more correlated with the reduction of memory rather than FLOPs, and thus we employ the memory reduction of each channel to normalize the importance. Our method can be used to prune any structures including those with coupled channels. We conduct extensive experiments on various backbones, including the classic ResNet and ResNeXt, mobile-friendly MobileNetV2, and the NAS-based RegNet, both on image classification and object detection which is under-explored. Experimental results validate that our method can effectively prune sophisticated networks, boosting inference speed without sacrificing accuracy.
Liyang Liu, Zhanghui Kuang, Aojun Zhou, Jing-Hao Xue, Xinjiang Wang, Wenming Yang, Qingmin Liao, Wayne Zhang 0001
ICML5
2021 Hourglass Face Detector for Hard Face
abstract
Face detection is an upstream task of facial image analysis. In many real-world scenarios, we need to detect small, occluded or dense faces that are hard to detect, but hard face detection is a challenging task in particular considering the balance between accuracy and inference speed for real-world applications. This paper proposes an Hourglass Face Detector (HFD) for hard face by developing a deep one-stage fully-convolutional hourglass network, which achieves an excellent balance between accuracy and inference speed. To this end, the HFD firstly shrinks a feature map by a series of stridden convolutional layers rather than pooling layers, so that useful subtle information is preserved better. Secondly, it exploits context information by merging fine-grained shallow feature maps with deep ones full of semantic information, making a better fusion of detailed information and semantic information to achieve a better detection of small faces. Moreover, the HFD exploits prior and multiscale information from the training data to enhance its scale-invariance and adaptability of anchor scales. Compared with the SSH and S3FD methods, the HFD can achieve a better performance in average precision on detecting hard faces as well as a quicker inference. Experiments on the WIDER FACE and FDDB datasets demonstrate the superior performance of our proposed method.
Zijun Yu, Jian Yin 0016, Wenming Yang, Jing-Hao Xue, Qingmin Liao
IJCNN5
2021 D³Net: Dual-Branch Disturbance Disentangling Network for Facial Expression Recognition
abstract
One of the main challenges in facial expression recognition (FER) is to address the disturbance caused by various disturbing factors, including common ones (such as identity, pose, and illumination) and potential ones (such as hairstyle, accessory, and occlusion). Recently, a number of FER methods have been developed to explicitly or implicitly alleviate the disturbance involved in facial images. However, these methods either consider only a few common disturbing factors or neglect the prior information of these disturbing factors, thus resulting in inferior recognition performance. In this paper, we propose a novel Dual-branch Disturbance Disentangling Network (D3Net), mainly consisting of an expression branch and a disturbance branch, to perform effective FER. In the disturbance branch, a label-aware sub-branch (LAS) and a label-free sub-branch (LFS) are elaborately designed to cope with different types of disturbing factors. On the one hand, LAS explicitly captures the disturbance due to some common disturbing factors by transfer learning on a pretrained model. On the other hand, LFS implicitly encodes the information of potential disturbing factors in an unsupervised manner. In particular, we introduce an Indian buffet process (IBP) prior to model the distribution of potential disturbing factors in LFS. Moreover, we leverage adversarial training to increase the differences between disturbance features and expression features, thereby enhancing the disentanglement of disturbing factors. By disentangling the disturbance from facial images, we are able to extract discriminative expression features. Extensive experiments demonstrate that our proposed method performs favorably against several state-of-the-art FER methods on both in-the-lab and in-the-wild databases.
Rongyun Mo, Yan Yan 0001, Jing-Hao Xue, Si Chen 0002, Hanzi Wang
ACM Multimedia3
2021 How Video Super-Resolution and Frame Interpolation Mutually Benefit
abstract
Video super-resolution (VSR) and video frame interpolation (VFI) are inter-dependent for enhancing videos of low resolution and low frame rate. However, most studies treat VSR and temporal VFI as independent tasks. In this work, we design a spatial-temporal super-resolution network based on exploring the interaction between VSR and VFI. The main idea is to improve the middle frame of VFI by the super-resolution (SR) frames and feature maps from VSR. In the meantime, VFI also provides extra information for VSR and thus, through interacting, the SR of consecutive frames of the original video can also be improved by the feedback from the generated middle frame. Drawing on this, our approach leverages a simple interaction of VSR and VFI and achieves state-of-the-art performance on various datasets. Due to such a simple strategy, our approach is universally applicable to any existing VSR or VFI networks for effectively improving their video enhancement performance.
Chengcheng Zhou, Zongqing Lu 0001, Linge Li, Qiangyu Yan, Jing-Hao Xue
ACM Multimedia5
2021 Small-Vote Sample Selection for Label-Noise Learning
Youze Xu, Yan Yan 0001, Jing-Hao Xue, Yang Lu 0009, Hanzi Wang
ECML/PKDD (3)3
2021 Dimension reduction for data with heterogeneous missingness
abstract
Dimension reduction plays a pivotal role in analysing high-dimensional data. However, observations with missing values present serious difficulties in directly applying standard dimension reduction techniques. As a large number of dimension reduction approaches are based on the Gram matrix, we first investigate the effects of missingness on dimension reduction by studying the statistical properties of the Gram matrix with or without missingness, and then we present a bias-corrected Gram matrix with nice statistical properties under heterogeneous missingness. Extensive empirical results, on both simulated and publicly available real datasets, show that the proposed unbiased Gram matrix can significantly improve a broad spectrum of representative dimension reduction approaches.
Yurong Ling, Zijing Liu, Jing-Hao Xue
UAI3
2021 RFRN: A recurrent feature refinement network for accurate and efficient scene text detection
Guanyu Deng, Yue Ming 0001, Jing-Hao Xue
Neurocomputing3
2021 Deep InterBoost networks for small-sample image classification
Dongliang Chang, Zhanyu Ma, Zheng-Hua Tan, Jing-Hao Xue, Jie Cao 0014, Jun Guo 0002
Neurocomputing5
2021 Mutual-learning sequence-level knowledge distillation for automatic speech recognition
Yue Ming 0001, Lei Yang 0034, Jing-Hao Xue
Neurocomputing4
2021 A concise review of recent few-shot meta-learning methods
Jing-Hao Xue, Zhanyu Ma
Neurocomputing3
2021 TSingNet: Scale-aware and context-rich feature learning for traffic sign detection and recognition in the wild
Yuanyuan Liu 0004, Jiyao Peng, Jing-Hao Xue, Yongquan Chen, Zhang-Hua Fu
Neurocomputing3
2021 3D-TDC: A 3D temporal dilation convolution framework for video action recognition
Yue Ming 0001, Chao Li 0026, Jing-Hao Xue
Neurocomputing4
2021 Cali-sketch: Stroke calibration and completion for high-quality face image generation from human-like sketches
Weihao Xia 0001, Yujiu Yang 0001, Jing-Hao Xue
Neurocomputing3
2021 Guest Editorial: Special issue on deep learning with small samples
Jing-Hao Xue, Jufeng Yang, Yan Yan 0001, Yujiu Yang 0001, Zongqing Lu 0001, Zhanyu Ma
Neurocomputing1
2021 Generalisations of stochastic supervision models
Xiaoou Lu, Yangqi Qiao, Rui Zhu 0006, Guijin Wang, Zhanyu Ma, Jing-Hao Xue
Pattern Recognit.6
2021 Non-Local Aggregation for RGB-D Semantic Segmentation
abstract
Exploiting both RGB (2D appearance) and Depth (3D geometry) information can improve the performance of semantic segmentation. However, due to the inherent difference between the RGB and Depth information, it remains a challenging problem in how to integrate RGB-D features effectively. In this letter, to address this issue, we propose a Non-local Aggregation Network (NANet), with a well-designed Multi-modality Non-local Aggregation Module (MNAM), to better exploit the non-local context of RGB-D features at multi-stage. Compared with most existing RGB-D semantic segmentation schemes, which only exploit local RGB-D features, the MNAM enables the aggregation of non-local RGB-D information along both spatial and channel dimensions. The proposed NANet achieves comparable performances with state-of-the-art methods on popular RGB-D benchmarks, NYUDv2 and SUN-RGBD.
Guodong Zhang 0004, Jing-Hao Xue, Pengwei Xie, Sifan Yang, Guijin Wang
IEEE Signal Process. Lett.2
2021 ReMarNet: Conjoint Relation and Margin Learning for Small-Sample Image Classification
abstract
Despite achieving state-of-the-art performance, deep learning methods generally require a large amount of labeled data during training and may suffer from overfitting when the sample size is small. To ensure good generalizability of deep networks under small sample sizes, learning discriminative features is crucial. To this end, several loss functions have been proposed to encourage large intra-class compactness and inter-class separability. In this paper, we propose to enhance the discriminative power of features from a new perspective by introducing a novel neural network termed Relation-and-Margin learning Network (ReMarNet). Our method assembles two networks of different backbones so as to learn the features that can perform excellently in both of the aforementioned two classification mechanisms. Specifically, a relation network is used to learn the features that can support classification based on the similarity between a sample and a class prototype; at the meantime, a fully connected network with the cross entropy loss is used for classification via the decision boundary. Experiments on four image datasets demonstrate that our approach is effective in learning discriminative features from a small set of labeled samples and achieves competitive performance against state-of-the-art methods. Code is available at https://github.com/liyunyu08/ReMarNet.
Liyun Yu, Zhanyu Ma, Jing-Hao Xue, Jie Cao 0014, Jun Guo 0002
IEEE Trans. Circuits Syst. Video Technol.5
2021 Domain Fingerprints for No-Reference Image Quality Assessment
abstract
Human fingerprints are detailed and nearly unique markers of human identity. Such a unique and stable fingerprint is also left on each acquired image. It can reveal how an image was degraded during the image acquisition procedure and thus is closely related to the quality of an image. In this work, we propose a new no-reference image quality assessment (NR-IQA) approach called domain-aware IQA (DA-IQA), which for the first time introduces the concept of domain fingerprint to the NR-IQA field. The domain fingerprint of an image is learned from image collections of different degradations and then used as the unique characteristics to identify the degradation sources and assess the quality of the image. To this end, we design a new domain-aware architecture, which enables simultaneous determination of both the distortion sources and the quality of an image. With the distortion in an image better characterized, the image quality can be more accurately assessed, as verified by extensive experiments, which show that the proposed DA-IQA performs better than almost all the compared state-of-the-art NR-IQA methods.
Weihao Xia 0001, Yujiu Yang 0001, Jing-Hao Xue, Jing Xiao 0006
IEEE Trans. Circuits Syst. Video Technol.3
2021 Semantic-Aware Occlusion-Robust Network for Occluded Person Re-Identification
abstract
In recent years, deep learning-based person re-identification (Re-ID) methods have made significant progress. However, the performance of these methods substantially decreases when dealing with occlusion, which is ubiquitous in realistic scenarios. In this article, we propose a novel semantic-aware occlusion-robust network (SORN) that effectively exploits the intrinsic relationship between the tasks of person Re-ID and semantic segmentation for occluded person Re-ID. Specifically, the SORN is composed of three branches, including a local branch, a global branch, and a semantic branch. In particular, the local branch extracts part-based local features, and the global branch leverages a novel spatial-patch contrastive loss (SPC) to extract occlusion-robust global features. Meanwhile, the semantic branch generates a foreground-background mask for a pedestrian image, which indicates the non-occluded areas of the human body. The three branches are jointly trained in a unified multi-task learning network. Finally, pedestrian matching is performed based on the local features extracted from the non-occluded areas and the global features extracted from the whole pedestrian image. Extensive experimental results on a large-scale occluded person Re-ID dataset (i.e., Occluded-DukeMTMC) and two partial person Re-ID datasets (i.e., Partial-REID and Partial-iLIDS) show the superiority of the proposed method compared with several state-of-the-art methods for occluded and partial person Re-ID. We also demonstrate the effectiveness of the proposed method on two general person Re-ID datasets (i.e., Market-1501 and DukeMTMC-reID).
Yan Yan 0001, Jing-Hao Xue, Yang Hua 0001, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.3
2021 BSNet: Bi-Similarity Network for Few-shot Fine-grained Image Classification
abstract
Few-shot learning for fine-grained image classification has gained recent attention in computer vision. Among the approaches for few-shot learning, due to the simplicity and effectiveness, metric-based methods are favorably state-of-the-art on many tasks. Most of the metric-based methods assume a single similarity measure and thus obtain a single feature space. However, if samples can simultaneously be well classified via two distinct similarity measures, the samples within a class can distribute more compactly in a smaller feature space, producing more discriminative feature maps. Motivated by this, we propose a so-called Bi-Similarity Network (BSNet) that consists of a single embedding module and a bi-similarity module of two similarity measures. After the support images and the query images pass through the convolution-based embedding module, the bi-similarity module learns feature maps according to two similarity measures of diverse characteristics. In this way, the model is enabled to learn more discriminative and less similarity-biased features from few shots of fine-grained images, such that the model generalization ability can be significantly improved. Through extensive experiments by slightly modifying established metric/similarity based networks, we show that the proposed approach produces a substantial improvement on several fine-grained image benchmark datasets. Codes are available at: https://github.com/PRIS-CV/BSNet.
Jijie Wu, Zhanyu Ma, Jie Cao 0014, Jing-Hao Xue
IEEE Trans. Image Process.6
2021 DS-UI: Dual-Supervised Mixture of Gaussian Mixture Models for Uncertainty Inference in Image Recognition
abstract
This paper proposes a dual-supervised uncertainty inference (DS-UI) framework for improving Bayesian estimation-based UI in DNN-based image recognition. In the DS-UI, we combine the classifier of a DNN, i.e., the last fully-connected (FC) layer, with a mixture of Gaussian mixture models (MoGMM) to obtain an MoGMM-FC layer. Unlike existing UI methods for DNNs, which only calculate the means or modes of the DNN outputs' distributions, the proposed MoGMM-FC layer acts as a probabilistic interpreter for the features that are inputs of the classifier to directly calculate the probabilities of them for the DS-UI. In addition, we propose a dual-supervised stochastic gradient-based variational Bayes (DS-SGVB) algorithm for the MoGMM-FC layer optimization. Unlike conventional SGVB and optimization algorithms in other UI methods, the DS-SGVB not only models the samples in the specific class for each Gaussian mixture model (GMM) in the MoGMM, but also considers the negative samples from other classes for the GMM to reduce the intra-class distances and enlarge the inter-class margins simultaneously for enhancing the learning ability of the MoGMM-FC layer in the DS-UI. Experimental results show the DS-UI outperforms the state-of-the-art UI methods in misclassification detection. We further evaluate the DS-UI in open-set out-of-domain/-distribution detection and find statistically significant improvements. Visualizations of the feature spaces demonstrate the superiority of the DS-UI. Codes are available at https://github.com/PRIS-CV/DS-UI.
Jiyang Xie 0001, Zhanyu Ma, Jing-Hao Xue, Guoqiang Zhang 0003, Yinhe Zheng, Jun Guo 0002
IEEE Trans. Image Process.3
2021 Ripple-GAN: Lane Line Detection With Ripple Lane Line Detection Network and Wasserstein GAN
abstract
With artificial intelligence technology being advanced by leaps and bounds, intelligent driving has attracted a huge amount of attention recently in research and development. In intelligent driving, lane line detection is a fundamental but challenging task particularly under complex road conditions. In this paper, we propose a simple yet appealing network called Ripple Lane Line Detection Network (RiLLD-Net), to exploit quick connections and gradient maps for effective learning of lane line features. RiLLD-Net can handle most common scenes of lane line detection. Then, in order to address challenging scenarios such as occluded or complex lane lines, we propose a more powerful network called Ripple-GAN, by integrating RiLLD-Net, confrontation training of Wasserstein generative adversarial networks, and multi-target semantic segmentation. Experiments show that, especially for complex or obscured lane lines, Ripple-GAN can produce a superior detection performance to other state-of-the-art methods.
Youcheng Zhang, Zongqing Lu 0001, Jing-Hao Xue, Qingmin Liao
IEEE Trans. Intell. Transp. Syst.4
2021 A Polarization-Imaging-Based Machine Learning Framework for Quantitative Pathological Diagnosis of Cervical Precancerous Lesions
abstract
Polarization images encode high resolution microstructural information even at low resolution. We propose a framework combining polarization imaging and traditional microscopy imaging, constructing a dual-modality machine learning framework that is not only accurate but also generalizable and interpretable. We demonstrate the viability of our proposed framework using the cervical intraepithelial neoplasia grading task, providing a polarimetry feature parameter to quantitatively characterize microstructural variations with lesion progression in hematoxylin-eosin-stained pathological sections of cervical precancerous tissues. By taking advantages of polarization imaging techniques and machine learning methods, the model enables interpretable and quantitative diagnosis of cervical precancerous lesion cases with improved sensitivity and accuracy in a low-resolution and wide-field system. The proposed framework applies routine image-analysis technology to identify the macro-structure and segment the target region in H&E-stained pathological images, and then employs emerging polarization method to extract the micro-structure information of the target region, which intends to expand the boundary of the current image-heavy digital pathology, bringing new possibilities for quantitative medical diagnosis.
Jiachen Wan, Jing-Hao Xue, Jibin Zou, Honghui He, Anli Hou, Hui Ma 0003
IEEE Trans. Medical Imaging4
2021 IncDet: In Defense of Elastic Weight Consolidation for Incremental Object Detection
abstract
Elastic weight consolidation (EWC) has been successfully applied for general incremental learning to overcome the catastrophic forgetting issue. It adaptively constrains each parameter of the new model not to deviate much from its counterpart in the old model during fine-tuning on new class data sets, according to its importance weight for old tasks. However, the previous study demonstrates that it still suffers from catastrophic forgetting when directly used in object detection. In this article, we show EWC is effective for incremental object detection if with critical adaptations. First, we conduct controlled experiments to identify two core issues why EWC fails if trivially applied to incremental detection: 1) the absence of old class annotations in new class images makes EWC misclassify objects of old classes in these images as background and 2) the quadratic regularization loss in EWC easily leads to gradient explosion when balancing old and new classes. Then, based on the abovementioned findings, we propose the corresponding solutions to tackle these issues: 1) utilize pseudobounding box annotations of old classes on new data sets to compensate for the absence of old class annotations and 2) adopt a novel Huber regularization instead of the original quadratic loss to prevent from unstable training. Finally, we propose a general EWC-based incremental object detection framework and implement it under both Fast R-CNN and Faster R-CNN, showing its flexibility and versatility. In terms of either the final performance or the performance drop with respect to the upper bound of joint training on all seen classes, evaluations on the PASCAL VOC and COCO data sets show that our method achieves a new state of the art.
Liyang Liu, Zhanghui Kuang, Jing-Hao Xue, Wenming Yang, Wayne Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2021 Class-Variant Margin Normalized Softmax Loss for Deep Face Recognition
abstract
In deep face recognition, the commonly used softmax loss and its newly proposed variations are not yet sufficiently effective to handle the class imbalance and softmax saturation issues during the training process while extracting discriminative features. In this brief, to address both issues, we propose a class-variant margin (CVM) normalized softmax loss, by introducing a true-class margin and a false-class margin into the cosine space of the angle between the feature vector and the class-weight vector. The true-class margin alleviates the class imbalance problem, and the false-class margin postpones the early individual saturation of softmax. With negligible computational complexity increment during training, the new loss function is easy to implement in the common deep learning frameworks. Comprehensive experiments on the LFW, YTF, and MegaFace protocols demonstrate the effectiveness of the proposed CVM loss function.
Wanping Zhang, Yongru Chen, Wenming Yang, Guijin Wang, Jing-Hao Xue, Qingmin Liao
IEEE Trans. Neural Networks Learn. Syst.5
2020 Lightweight Single Image Super-Resolution Through Efficient Second-Order Attention Spindle Network
abstract
Recent years have witnessed great success of applying deep convolutional neural networks (CNNs) to single image super-resolution (SISR). However, most of these algorithms focus on increasing modeling capability through developing deeper and wider networks, improving the performance but at a cost of huge computation. Targeting at a better trade-off between efficiency and effectiveness, we propose ESASN, an efficient second-order attention spindle network for lightweight SISR. ESASN is built upon efficient second-order attention spindle (ESAS) blocks, each of which contains two well-designed new modules, efficient multi-scale (EMS) module and second-order attention (SOA) module. EMS reduces a considerable number of parameters while retaining the multi-scale structure to explore rich features. SOA further rescales the multi-scale feature maps, capturing the inter-dependencies among channels pixel-wisely with little additional cost. Both qualitative and quantitative experimental results demonstrate that the combination of EMS and SOA works out favorably for SISR, lifting the performance with fewer parameters. Code is available at https://github.com/yiyunchen/ESASN.
Jing-Hao Xue, Wenming Yang, Qingmin Liao
ICME3
2020 Deep Disturbance-Disentangled Learning for Facial Expression Recognition
abstract
To achieve effective facial expression recognition (FER), it is of great importance to address various disturbing factors, including pose, illumination, identity, and so on. However, a number of FER databases merely provide the labels of facial expression, identity, and pose, but lack the label information for other disturbing factors. As a result, many methods are only able to cope with one or two disturbing factors, ignoring the heavy entanglement between facial expression and multiple disturbing factors. In this paper, we propose a novel Deep Disturbance-disentangled Learning (DDL) method for FER. DDL is capable of simultaneously and explicitly disentangling multiple disturbing factors by taking advantage of multi-task learning and adversarial transfer learning. The training of DDL involves two stages. First, a Disturbance Feature Extraction Model (DFEM) is pre-trained to perform multi-task learning for classifying multiple disturbing factors on the large-scale face database (which has the label information for various disturbing factors). Second, a Disturbance-Disentangled Model (DDM), which contains a global shared sub-network and two task-specific (i.e., expression and disturbance) sub-networks, is learned to encode the disturbance-disentangled information for expression recognition. The expression sub-network adopts a multi-level attention mechanism to extract expression-specific features, while the disturbance sub-network leverages adversarial transfer learning to extract disturbance-specific features based on the pre-trained DFEM. Experimental results on both the in-the-lab FER databases (including CK+, MMI, and Oulu-CASIA) and the in-the-wild FER databases (including RAF-DB and SFEW) demonstrate the superiority of our proposed method compared with several state-of-the-art methods.
Delian Ruan, Yan Yan 0001, Si Chen 0002, Jing-Hao Xue, Hanzi Wang
ACM Multimedia4
2020 Controllable Continuous Gaze Redirection
abstract
In this work, we present interpGaze, a novel framework for controllable gaze redirection that achieves both precise redirection and continuous interpolation. Given two gaze images with different attributes, our goal is to redirect the eye gaze of one person into any gaze direction depicted in the reference image or to generate continuous intermediate results. To accomplish this, we design a model including three cooperative components: an encoder, a controller and a decoder. The encoder maps images into a well-disentangled and hierarchically-organized latent space. The controller adjusts the magnitudes of latent vectors to the desired strength of corresponding attributes by altering a control vector. The decoder converts the desired representations from the attribute space to the image space. To facilitate covering the full space of gaze directions, we introduce a high-quality gaze image dataset with a large range of directions, which also benefits researchers in related areas. Extensive experimental validation and comparisons to several baseline methods show that the proposed interpGaze outperforms state-of-the-art methods in terms of image quality and redirection precision.
Weihao Xia 0001, Yujiu Yang 0001, Jing-Hao Xue, Wensen Feng
ACM Multimedia3
2020 Generalization Bound of Gradient Descent for Non-Convex Metric Learning
abstract
Metric learning aims to learn a distance measure that can benefit distance-based methods such as the nearest neighbour (NN) classifier. While considerable efforts have been made to improve its empirical performance and analyze its generalization ability by focusing on the data structure and model complexity, an unresolved question is how choices of algorithmic parameters, such as the number of training iterations, affect metric learning as it is typically formulated as an optimization problem and nowadays more often as a non-convex problem. In this paper, we theoretically address this question and prove the agnostic Probably Approximately Correct (PAC) learnability for metric learning algorithms with non-convex objective functions optimized via gradient descent (GD); in particular, our theoretical guarantee takes the iteration number into account. We first show that the generalization PAC bound is a sufficient condition for agnostic PAC learnability and this bound can be obtained by ensuring the uniform convergence on a densely concentrated subset of the parameter space. We then show that, for classifiers optimized via GD, their generalizability can be guaranteed if the classifier and loss function are both Lipschitz smooth, and further improved by using fewer iterations. To illustrate and exploit the theoretical findings, we finally propose a novel metric learning method called Smooth Metric and representative Instance LEarning (SMILE), designed to satisfy the Lipschitz smoothness property and learned via GD with an early stopping mechanism for better discriminability and less computational cost of NN.
Mingzhi Dong, Rui Zhu 0006, Yujiang Wang 0001, Jing-Hao Xue
NeurIPS5
2020 Metric Learning for Categorical and Ambiguous Features: An Adversarial Method
Mingzhi Dong, Yiwen Guo, Jing-Hao Xue
ECML/PKDD (2)4
2020 Deep learning for image super-resolution
Wenming Yang, Fei Zhou 0001, Rui Zhu 0006, Kazuhiro Fukui, Guijin Wang, Jing-Hao Xue
Neurocomputing6
2020 Unsupervised multi-domain multimodal image-to-image translation with explicit domain-constrained disentanglement
Weihao Xia 0001, Yujiu Yang 0001, Jing-Hao Xue
Neural Networks3
2020 Learning Local Metrics and Influential Regions for Classification
abstract
The performance of distance-based classifiers heavily depends on the underlying distance metric, so it is valuable to learn a suitable metric from the data. To address the problem of multimodality, it is desirable to learn local metrics. In this short paper, we define a new intuitive distance with local metrics and influential regions, and subsequently propose a novel local metric learning algorithm called LMLIR for distance-based classification. Our key intuition is to partition the metric space into influential regions and a background region, and then regulate the effectiveness of each local metric to be within the related influential regions. We learn multiple local metrics and influential regions to reduce the empirical hinge loss, and regularize the parameters on the basis of a resultant learning bound. Encouraging experimental results are obtained from various public and popular data sets.
Mingzhi Dong, Yujiang Wang 0001, Jing-Hao Xue
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Defocus map estimation from a single image using improved likelihood feature and edge-based basis
Qingmin Liao, Jing-Hao Xue, Fei Zhou 0001
Pattern Recognit.3
2020 Enhanced Grassmann discriminant analysis with randomized time warping for motion recognition
Lincon Sales de Souza, Bernardo Bentes Gatto, Jing-Hao Xue, Kazuhiro Fukui
Pattern Recognit.3
2020 Data-augmented matched subspace detector for hyperspectral subpixel target detection
Mingzhi Dong, Ziyu Wang 0003, Lianru Gao, Lefei Zhang, Jing-Hao Xue
Pattern Recognit.6
2020 Real-MFF: A large realistic multi-focus image dataset with ground truth
Juncheng Zhang, Qingmin Liao, Wenming Yang, Jing-Hao Xue
Pattern Recognit. Lett.6
2020 Adjusting the imbalance ratio by the dimensionality of imbalanced data
Rui Zhu 0006, Yiwen Guo, Jing-Hao Xue
Pattern Recognit. Lett.3
2020 Hyperspectral image denoising based on low-rank coefficients and orthonormal dictionary
Fanlong Zhang, Guowei Yang 0002, Jing-Hao Xue
Signal Process.3
2020 Inter-class angular margin loss for face recognition
Jingna Sun, Wenming Yang, Riqiang Gao, Jing-Hao Xue, Qingmin Liao
Signal Process. Image Commun.4
2020 Special Issue on Advances in Statistical Methods-based Visual Quality Assessment
Fei Zhou 0001, Wenming Yang, Xinbo Gao 0001, Hantao Liu, Rui Zhu 0006, Jing-Hao Xue
Signal Process. Image Commun.6
2020 Hyperspectral and Multispectral Image Fusion Using Optimized Twin Dictionaries
abstract
Spectral or spatial dictionary has been widely used in fusing low-spatial-resolution hyperspectral (LH) images and high-spatial-resolution multispectral (HM) images. However, only using spectral dictionary is insufficient for preserving spatial information, and vice versa. To address this problem, a new LH and HM image fusion method termed OTD using optimized twin dictionaries is proposed in this paper. The fusion problem of OTD is formulated analytically in the framework of sparse representation, as an optimization of twin spectral-spatial dictionaries and their corresponding sparse coefficients. More specifically, the spectral dictionary representing the generalized spectrums and its spectral sparse coefficients are optimized by utilizing the observed LH and HM images in the spectral domain; and the spatial dictionary representing the spatial information and its spatial sparse coefficients are optimized by modeling the rest of high-frequency information in the spatial domain. In addition, without non-negative constraints, the alternating direction methods of multipliers (ADMM) are employed to implement the above optimization process. Comparison results with the related state-of-the-art fusion methods on various datasets demonstrate that our proposed OTD method achieves a better fusion performance in both spatial and spectral domains.
Xiaolin Han 0001, Jing Yu 0005, Jing-Hao Xue
IEEE Trans. Image Process.3
2020 OSLNet: Deep Small-Sample Classification With an Orthogonal Softmax Layer
abstract
A deep neural network of multiple nonlinear layers forms a large function space, which can easily lead to overfitting when it encounters small-sample data. To mitigate overfitting in small-sample classification, learning more discriminative features from small-sample data is becoming a new trend. To this end, this paper aims to find a subspace of neural networks that can facilitate a large decision margin. Specifically, we propose the Orthogonal Softmax Layer (OSL), which makes the weight vectors in the classification layer remain orthogonal during both the training and test processes. The Rademacher complexity of a network using the OSL is only 1/K, where K is the number of classes, of that of a network using the fully connected classification layer, leading to a tighter generalization error bound. Experimental results demonstrate that the proposed OSL has better performance than the methods used for comparison on four small-sample benchmark datasets, as well as its applicability to large-sample datasets. Codes are available at: https://github.com/dongliangchang/OSLNet.
Dongliang Chang, Zhanyu Ma, Zheng-Hua Tan, Jing-Hao Xue, Jie Cao 0014, Jingyi Yu 0001, Jun Guo 0002
IEEE Trans. Image Process.5
2020 An α-Matte Boundary Defocus Model-Based Cascaded Network for Multi-Focus Image Fusion
abstract
Capturing an all-in-focus image with a single camera is difficult since the depth of field of the camera is usually limited. An alternative method to obtain the all-in-focus image is to fuse several images that are focused at different depths. However, existing multi-focus image fusion methods cannot obtain clear results for areas near the focused/defocused boundary (FDB). In this paper, a novel α-matte boundary defocus model is proposed to generate realistic training data with the defocus spread effect precisely modeled, especially for areas near the FDB. Based on this α-matte defocus model and the generated data, a cascaded boundary-aware convolutional network termed MMF-Net is proposed and trained, aiming to achieve clearer fusion results around the FDB. Specifically, the MMF-Net consists of two cascaded subnets for initial fusion and boundary fusion. These two subnets are designed to first obtain a guidance map of FDB and then refine the fusion near the FDB. Experiments demonstrate that with the help of the new α-matte boundary defocus model, the proposed MMF-Net outperforms the state-of-the-art methods both qualitatively and quantitatively.
Qingmin Liao, Juncheng Zhang, Jing-Hao Xue
IEEE Trans. Image Process.5
2020 LCSCNet: Linear Compressing-Based Skip-Connecting Network for Image Super-Resolution
abstract
In this paper, we develop a concise but efficient network architecture called linear compressing based skipconnecting network (LCSCNet) for image super-resolution. Compared with two representative network architectures with skip connections, ResNet and DenseNet, a linear compressing layer is designed in LCSCNet for skip connection, which connects former feature maps and distinguishes them from newly-explored feature maps. In this way, the proposed LCSCNet enjoys the merits of the distinguish feature treatment of DenseNet and the parametereconomic form of ResNet. Moreover, to better exploit hierarchical information from both low and high levels of various receptive fields in deep models, inspired by gate units in LSTM, we also propose an adaptive element-wise fusion strategy with multisupervised training. Experimental results in comparison with state-of-the-art algorithms validate the effectiveness of LCSCNet.
Wenming Yang, Xuechen Zhang 0003, Yapeng Tian, Wei Wang 0194, Jing-Hao Xue, Qingmin Liao
IEEE Trans. Image Process.5
2020 An Equalized Margin Loss for Face Recognition
abstract
In this paper, we propose a new loss function, termed the equalized margin (EqM) loss, which is designed to make both intra-class scopes and inter-class margins similar over all classes, such that all the classes can be evenly distributed on the hypersphere of the feature space. The EqM loss controls both the lower limit of intra-class similarity by exploiting hard-sample mining and the upper limit of inter-class similarity by assuring equalized margins. Therefore, using the EqM loss, we can not only obtain more discriminative features, but also overcome the negative impacts from the data imbalance on the inter-class margins. We also observe that the EqM loss is stable with the variation of the scale in normalized Softmax. Furthermore, by conducting extensive experiments on LFW, YTF, CFP, MegaFace and IJB-B, we are able to verify the effectiveness and superiority of the EqM loss, compared with other state-of-the-art loss functions for face recognition.
Jingna Sun, Wenming Yang, Jing-Hao Xue, Qingmin Liao
IEEE Trans. Multim.3
2020 Insights Into Multiple/Single Lower Bound Approximation for Extended Variational Inference in Non-Gaussian Structured Data Modeling
abstract
For most of the non-Gaussian statistical models, the data being modeled represent strongly structured properties, such as scalar data with bounded support (e.g., beta distribution), vector data with unit length (e.g., Dirichlet distribution), and vector data with positive elements (e.g., generalized inverted Dirichlet distribution). In practical implementations of non-Gaussian statistical models, it is infeasible to find an analytically tractable solution to estimating the posterior distributions of the parameters. Variational inference (VI) is a widely used framework in Bayesian estimation. Recently, an improved framework, namely, the extended VI (EVI), has been introduced and applied successfully to a number of non-Gaussian statistical models. EVI derives analytically tractable solutions by introducing lower bound approximations to the variational objective function. In this paper, we compare two approximation strategies, namely, the multiple lower bounds (MLBs) approximation and the single lower bound (SLB) approximation, which can be applied to carry out the EVI. For implementation, two different conditions, the weak and the strong conditions, are discussed. Convergence of the EVI depends on the selection of the lower bound, regardless of the choice of weak or strong condition. We also discuss the convergence properties to clarify the differences between MLB and SLB. Extensive comparisons are made based on some EVI-based non-Gaussian statistical models. Theoretical analysis is conducted to demonstrate the differences between the weak and strong conditions. Experimental results based on real data show advantages of the SLB approximation over the MLB approximation.
Zhanyu Ma, Jiyang Xie 0001, Yuping Lai, Jalil Taghia, Jing-Hao Xue, Jun Guo 0002
IEEE Trans. Neural Networks Learn. Syst.5
2020 Bi-Stream Pose-Guided Region Ensemble Network for Fingertip Localization From Stereo Images
abstract
In human-computer interaction, it is important to accurately estimate the hand pose, especially fingertips. However, traditional approaches to fingertip localization mainly rely on depth images and thus suffer considerably from noise and missing values. Instead of depth images, stereo images can also provide 3-D information of hands. There are nevertheless limitations on the dataset size, global viewpoints, hand articulations, and hand shapes in publicly available stereo-based hand pose datasets. To mitigate these limitations and promote further research on hand pose estimation from stereo images, we build a new large-scale binocular hand pose dataset called THU-Bi-Hand, offering a new perspective for fingertip localization. In the THU-Bi-Hand dataset, there are 447k pairs of stereo images of different hand shapes from ten subjects with accurate 3-D location annotations of the wrist and five fingertips. Captured with minimal restriction on the range of hand motion, the dataset covers a large global viewpoint space and hand articulation space. To better present the performance of fingertip localization on THU-Bi-Hand, we propose a novel scheme termed bi-stream pose-guided region ensemble network (Bi-Pose-REN). It extracts more representative feature regions around joints in the feature maps under the guidance of the previously estimated pose. The feature regions are integrated hierarchically according to the topology of hand joints to regress a refined hand pose. Bi-Pose-REN and several existing methods are evaluated on THU-Bi-Hand so that benchmarks are provided for further research. Experimental results show that our Bi-Pose-REN has achieved the best performance on THU-Bi-Hand.
Guijin Wang, Cairong Zhang, Xinghao Chen 0001, Xiangyang Ji, Jing-Hao Xue
IEEE Trans. Neural Networks Learn. Syst.5
2020 Guest Editorial Special Issue on Recent Advances in Theory, Methodology, and Applications of Imbalanced Learning
Jing-Hao Xue, Zhanyu Ma, Manuel Roveri, Nathalie Japkowicz
IEEE Trans. Neural Networks Learn. Syst.1
2020 A Novel Separating Hyperplane Classification Framework to Unify Nearest-Class-Model Methods for High-Dimensional Data
abstract
In this article, we establish a novel separating hyperplane classification (SHC) framework to unify three nearest-class-model methods for high-dimensional data: the nearest subspace method (NSM), the nearest convex hull method (NCHM), and the nearest convex cone method (NCCM). Nearest-class-model methods are an important paradigm for the classification of high-dimensional data. We first introduce the three nearest-class-model methods and then conduct dual analysis for theoretically investigating them, to understand deeply their underlying classification mechanisms. A new theorem for the dual analysis of NCCM is proposed in this article by discovering the relationship between a convex cone and its polar cone. We then establish the new SHC framework to unify the nearest-class-model methods based on the theoretical results. One important application of this new SHC framework is to help explain empirical classification results: why one class model has a better performance than others on certain data sets. Finally, we propose a new nearest-class-model method, the soft NCCM, under the novel SHC framework to solve the overlapping class model problem. For illustrative purposes, we empirically demonstrate the significance of our SHC framework and the soft NCCM through two types of typical real-world high-dimensional data: the spectroscopic data and the face image data.
Rui Zhu 0006, Ziyu Wang 0003, Naoya Sogi, Kazuhiro Fukui, Jing-Hao Xue
IEEE Trans. Neural Networks Learn. Syst.5
2019 Residual Dilated Network with Attention for Image Blind Denoising
abstract
Image denoising has recently witnessed substantial progress. However, many existing methods remain suboptimal for texture restoration due to treating different image regions and channels indiscriminately. Also they need to specify the noise level in advance, which largely hinders their use in blind denoising. Therefore, we introduce both attention mechanism and automatic noise level estimation into image denoising. Specifically, we propose a new, effective end-to-end attention-embedded neural network for image denoising, named as Residual Dilated Attention Network (RDAN). Our RDAN is composed of a series of tailored Residual Dilated Attention Blocks (RDAB) and Residual Conv Attention Blocks (RCAB). The RDAB and RCAB incorporates both non-local and local operations, which enable a comprehensive capture of structural information. In addition, we incorporate a Gaussian-based noise level estimation into RDAN to accomplish blind denoising. Experimental results have demonstrated that our RDAN can substantially outperforms the state-of-the-art denoising methods as well as promisingly preserve image texture.
Guanqun Hou, Yujiu Yang 0001, Jing-Hao Xue
ICME3
2019 A New Object Scene Flow Algorithm Based on Support Points Selection and Robust Moving Object Proposal
abstract
Recent algorithms of object scene flow estimation suffer from low computational efficiency or unstable moving object proposals. To tackle these two problems simultaneously, in this paper we propose a new, efficient and robust algorithm for object scene flow estimation, through making two technical contributions. Firstly to improve the efficiency, we propose to select only a few pixels termed support points for matching cost calculation rather than using all pixels. The support points are defined as those pixels with high confidence in feature matching. Secondly to attain stable moving object proposals, we propose a motion magnitude-adaptive thresholding scheme for ego-motion outlier detection, after patch matching on CNN-extracted high quality features. These two contributions, though simple, ensure a remarkable improvement in both efficiency and accuracy from the original object scene flow method, as well as making the proposed algorithm a strong practicable alternative to much more sophisticated state-of-the-art competitors.
Zhengyang Sun, Zongqing Lu 0001, Jing-Hao Xue, Qingmin Liao
ICME3
2019 A New Approach to Automatic Clothing Matting from Mannequins
abstract
It is crucial to extract retail clothes from images of mannequins when building a database of clothing images for virtual try-on systems. However, clothes often have complex texture and translucent material, such as holes and laces. It is thus difficult to extract clothes as foreground by existing generic natural image matting methods. Hence in this paper, we present a novel approach to automatic clothing matting from mannequins, with auxiliary information from a rough background image of the mannequin only. Experiments show that we can achieve remarkable improvement on the alpha matte near challenging regions of complex texture and translucent material of clothes. Moreover, our approach can automatically generate trimaps to facilitate the development and evaluation of other image matting algorithms.
Zongqing Lu 0001, Jing-Hao Xue, Qingmin Liao
ICME3
2019 A New Rotation-Invariant Deep Network for 3D Object Recognition
abstract
When inputs are rotated, most 3D convolutional neural networks (CNNs) will have their performance much dropped, especially for those models with voxelized input of 3D objects. The newly proposed Spherical CNNS, with the concept of the rotation-equivariant spherical correlation, aims to achieve rotation invariance. Inspired by this, we propose a new rotation-invariant deep network to recognize rotated 3D objects. Specifically, we adopt the spherical representation and the spherical correlation S^2 layer of Spherical CNNs, for their capacity of representing 3D objects and rotation equivariance. In the meantime, we improve the computational efficiency and expressiveness of Spherical CNNs, by replacing its time-consuming and depth-limited SO(3) layer with a PointNet-style network architecture. Hence our proposed network can maintain the equivariance as the network grows deeper while substantially reducing its runtime, leading to a much better efficiency and expressiveness of rotation-invariant representation. Experimental results show that our network performs better than or comparable to the state-of-the-art methods in the ModelNet40 classification challenge.
Yachi Zhang, Zongqing Lu 0001, Jing-Hao Xue, Qingmin Liao
ICME3
2019 A hybrid finger identification pattern using Polarized depth-Weighted Binary Direction Coding
Wenming Yang, Wenyang Ji, Jing-Hao Xue, Qingmin Liao
Neurocomputing3
2019 Image-text dual neural network with decision strategy for small-sample image classification
Fangyi Zhu, Zhanyu Ma, Guang Chen 0003, Jen-Tzung Chien, Jing-Hao Xue, Jun Guo 0002
Neurocomputing6
2019 Learning distance to subspace for the nearest subspace methods in high-dimensional data classification
Rui Zhu 0006, Mingzhi Dong, Jing-Hao Xue
Inf. Sci.3
2019 MSDH: Matched subspace detector with heterogeneous noise
Lefei Zhang, Lianru Gao, Jing-Hao Xue
Pattern Recognit. Lett.4
2019 FV-GAN: Finger Vein Representation Using Generative Adversarial Networks
abstract
In finger vein verification, the most important and challenging part is to robustly extract finger vein patterns from low-contrast infrared finger images with limited a priori knowledge. Although recent convolutional neural network (CNN)-based methods for finger vein verification have shown powerful capacity for feature representation and promising perspective in this area, they still have two critical issues to address. First, these CNN-based methods unexceptionally utilize fully connected layers, which restrict the size of finger vein images to process and increase the processing time. Second, the capacity of CNN for feature representation generally suffers from the low quality of finger vein ground-truth pattern maps for training, particularly due to outliers and vessel breaks. To address these issues, in this paper, we propose a novel approach termed FV-GAN to finger vein extraction and verification, based on generative adversarial network (GAN), as the first attempt in this area. Unlike the CNN-based methods, FV-GAN learns from the joint distribution of finger vein images and pattern maps rather than the direct mapping between them, with the aim at achieving stronger robustness against outliers and vessel breaks. Moreover, FV-GAN adopts fully convolutional networks as the basic architecture and discards fully connected layers, which relaxes the constraint on the input image size and reduces the computational expenditure for feature extraction. Furthermore, we design an adversarial training strategy and propose a hybrid loss function for FV-GAN. The experimental results on two public databases show significant improvement by FV-GAN in finger vein verification in terms of both verification accuracy and equal error rate.
Wenming Yang, Changqing Hui, Zhiquan Chen, Jing-Hao Xue, Qingmin Liao
IEEE Trans. Inf. Forensics Secur.4
2019 Deep Learning for Single Image Super-Resolution: A Brief Review
abstract
Single image super-resolution (SISR) is a notoriously challenging ill-posed problem that aims to obtain a high-resolution output from one of its low-resolution versions. Recently, powerful deep learning algorithms have been applied to SISR and have achieved state-of-the-art performance. In this survey, we review representative deep learning-based SISR methods and group them into two categories according to their contributions to two essential aspects of SISR: The exploration of efficient neural network architectures for SISR and the development of effective optimization objectives for deep SISR learning. For each category, a baseline is first established, and several critical limitations of the baseline are summarized. Then, representative works on overcoming these limitations are presented based on their original content, as well as our critical exposition and analyses, and relevant comparisons are conducted from a variety of perspectives. Finally, we conclude this review with some current challenges and future trends in SISR that leverage deep learning algorithms.
Wenming Yang, Xuechen Zhang 0003, Yapeng Tian, Wei Wang 0194, Jing-Hao Xue, Qingmin Liao
IEEE Trans. Multim.5
2018 Discriminatively guided filtering (DGF) for hyperspectral image classification
Ziyu Wang 0003, Huafeng Hu, Lefei Zhang, Jing-Hao Xue
Neurocomputing4
2018 Matched shrunken subspace detectors for hyperspectral target detection
Ziyu Wang 0003, Jing-Hao Xue
Neurocomputing2
2018 No-reference image quality assessment for photographic images based on robust statistics
Zhengda Zeng, Wenming Yang, Jing-Hao Xue, Qingmin Liao
Neurocomputing4
2018 MvSSIM: A quality assessment index for hyperspectral images
Rui Zhu 0006, Fei Zhou 0001, Jing-Hao Xue
Neurocomputing3
2018 LRID: A new metric of multi-class imbalance degree based on likelihood-ratio test
abstract
In this paper, we introduce a new likelihood ratio imbalance degree (LRID) to measure the class-imbalance extent of multi-class data. Imbalance ratio (IR) is usually used to measure class-imbalance extent in imbalanced learning problems. However, IR cannot capture the detailed information in the class distribution of multi-class data, because it only utilises the information of the largest majority class and the smallest minority class. Imbalance degree (ID) has been proposed to solve the problem of IR for multi-class data. However, we note that improper use of distance metric in ID can have harmful effect on the results. In addition, ID assumes that data with more minority classes are more imbalanced than data with less minority classes, which is not always true in practice. Thus ID cannot provide reliable measurement when the assumption is violated. In this paper, we propose a new metric based on the likelihood-ratio test, LRID, to provide a more reliable measurement of class-imbalance extent for multi-class data. Experiments on both simulated and real data show that LRID is competitive with IR and ID, and can reduce the negative correlation with F1 scores by up to 0.55.
Rui Zhu 0006, Ziyu Wang 0003, Zhanyu Ma, Guijin Wang, Jing-Hao Xue
Pattern Recognit. Lett.5
2018 Cone-based joint sparse modelling for hyperspectral image classification
Ziyu Wang 0003, Rui Zhu 0006, Kazuhiro Fukui, Jing-Hao Xue
Signal Process.4
2018 SPSIM: A Superpixel-Based Similarity Index for Full-Reference Image Quality Assessment
abstract
Full-reference image quality assessment algorithms usually perform comparisons of features extracted from square patches. These patches do not have any visual meanings. On the contrary, a superpixel is a set of image pixels that share similar visual characteristics and is thus perceptually meaningful. Features from superpixels may improve the performance of image quality assessment. Inspired by this, we propose a new superpixel-based similarity index by extracting perceptually meaningful features and revising similarity measures. The proposed method evaluates image quality on the basis of three measurements, namely, superpixel luminance similarity, superpixel chrominance similarity, and pixel gradient similarity. The first two measurements assess the overall visual impression on local images. The third measurement quantifies structural variations. The impact of superpixel-based regional gradient consistency on image quality is also analyzed. Distorted images showing high regional gradient consistency with the corresponding reference images are visually appreciated. Therefore, the three measurements are further revised by incorporating the regional gradient consistency into their computations. A weighting function that indicates superpixel-based texture complexity is utilized in the pooling stage to obtain the final quality score. Experiments on several benchmark databases demonstrate that the proposed method is competitive with the state-of-the-art metrics.
Qingmin Liao, Jing-Hao Xue, Fei Zhou 0001
IEEE Trans. Image Process.3
2018 Decorrelation of Neutral Vector Variables: Theory and Applications
abstract
In this paper, we propose novel strategies for neutral vector variable decorrelation. Two fundamental invertible transformations, namely, serial nonlinear transformation and parallel nonlinear transformation, are proposed to carry out the decorrelation. For a neutral vector variable, which is not multivariate-Gaussian distributed, the conventional principal component analysis cannot yield mutually independent scalar variables. With the two proposed transformations, a highly negatively correlated neutral vector can be transformed to a set of mutually independent scalar variables with the same degrees of freedom. We also evaluate the decorrelation performances for the vectors generated from a single Dirichlet distribution and a mixture of Dirichlet distributions. The mutual independence is verified with the distance correlation measurement. The advantages of the proposed decorrelation strategies are intensively studied and demonstrated with synthesized data and practical application evaluations.
Zhanyu Ma, Jing-Hao Xue, Arne Leijon, Zheng-Hua Tan, Zhen Yang 0004, Jun Guo 0002
IEEE Trans. Neural Networks Learn. Syst.2
2017 Transferring CNNS to multi-instance multi-label classification on small datasets
abstract
Image tagging is a well known challenge in image processing. It is typically addressed through multi-instance multi-label (MIML) classification methodologies. Convolutional Neural Networks (CNNs) possess great potential to perform well on MIML tasks, since multi-level convolution and max pooling coincide with the multi-instance setting and the sharing of hidden representation may benefit multi-label modeling. However, CNNs usually require a large amount of carefully labeled data for training, which is hard to obtain in many real applications. In this paper, we propose a new approach for transferring pre-trained deep networks such as VGG16 on Imagenet to small MIML tasks. We extract features from each group of the network layers and apply multiple binary classifiers to them for multi-label prediction. Moreover, we adopt an L1-norm regularized Logistic Regression (L1LR) to find the most effective features for learning the multi-label classifiers. The experiment results on two most-widely used and relatively small benchmark MIML image datasets demonstrate that the proposed approach can substantially outperform the state-of-the-art algorithms, in terms of all popular performance metrics.
Mingzhi Dong, Kunkun Pang, Yang Wu 0001, Jing-Hao Xue, Timothy M. Hospedales, Tsukasa Ogasawara
ICIP4
2017 What is the probability of replicating a statistically significant association in genome-wide association studies?
abstract
The goal of genome-wide association studies (GWASs) is to discover genetic variants associated with diseases/traits. Replication is a common validation method in GWASs. We regard an association as true finding when it shows significance in both primary and replication studies. A question worth pondering is what is the probability of a primary association (i.e. a statistically significant association in the primary study) being validated in the replication study? This article systematically reviews the answers to this question from different points of view. As Bayesian methods can help us integrate out the uncertainty about the underlying effect of the primary association, we will mainly focus on the Bayesian view in this article. We refer the Bayesian replication probability as the replication rate (RR). We further describe an estimation method for RR, which makes use of the summary statistics from the primary study. We can use the estimated RR to determine the sample size of the replication study and to check the consistency between the results of the primary study and those of the replication study. We describe an R-package to estimate and apply RR in GWASs. Simulation and real data experiments show that the estimated RR has good prediction and calibration performance. We also use these data to demonstrate the usefulness of RR. The R-package is available at http://bioinformatics.ust.hk/RRate.html.
Wei Jiang 0019, Jing-Hao Xue, Weichuan Yu
Briefings Bioinform.2
2017 Building a discriminatively ordered subspace on the generating matrix to classify high-dimensional spectral data
Rui Zhu 0006, Kazuhiro Fukui, Jing-Hao Xue
Inf. Sci.3
2017 On the orthogonal distance to class subspaces for high-dimensional data classification
Rui Zhu 0006, Jing-Hao Xue
Inf. Sci.2
2017 The matched subspace detector with interaction effects
abstract
This paper aims to propose a new hyperspectral target-detection method termed the matched subspace detector with interaction effects (MSDinter). The MSDinter introduces “interaction effects” terms into the popular matched subspace detector (MSD), from regression analysis in multivariate statistics and the bilinear mixing model in hyperspectral unmixing. In this way, the interaction between the target and the surrounding background, which should have but not yet been considered by the MSD, is modelled and estimated, such that superior performance of target detection can be achieved. Besides deriving the MSDinter methodologically, we also demonstrate its superiority empirically using two hyperspectral imaging datasets.
Ziyu Wang 0003, Jing-Hao Xue
Pattern Recognit.2
2017 Joint sparse model-based discriminative K-SVD for hyperspectral image classification
abstract
Sparse representation classification (SRC) is being widely investigated on hyperspectral images (HSI). For SRC methods to achieve high classification performance, not only is the development of sparse representation models essential, the designing and learning of quality dictionaries also plays an important role. That is, a redundant dictionary with well-designated atoms is required in order to ensure low reconstruction error, high discriminative power, and stable sparsity. In this paper, we propose a new method to learn such dictionaries for HSI classification. We borrow the concept of joint sparse model (JSM) from SRC to dictionary learning. JSM assumes local smoothness and joint sparsity and was initially proposed for classification of HSI. We leverage JSM to develop an extension of discriminative K-SVD for learning a promising discriminative dictionary for HSI. Through a semi-supervised strategy, the new dictionary learning method, termed JSM-DKSVD, utilises all spectrums over the local neighbourhoods of labelled training pixels for discriminative dictionary learning. It can produce a redundant dictionary with rich spectral and spatial information as well as high discriminative power. The learned dictionary can then be compatibly used in conjunction with the established SRC methods, and can significantly improve their performance for HSI classification.
Ziyu Wang 0003, Jianxiong Liu, Jing-Hao Xue
Signal Process.3
2017 Segmenting Multi-Source Images Using Hidden Markov Fields With Copula-Based Multivariate Statistical Distributions
abstract
Nowadays, multi-source image acquisition attracts an increasing interest in many fields, such as multi-modal medical image segmentation. Such acquisition aims at considering complementary information to perform image segmentation, since the same scene has been observed by various types of images. However, strong dependence often exists between multi-source images. This dependence should be taken into account when we try to extract joint information for precisely making a decision. In order to statistically model this dependence between multiple sources, we propose a novel multi-source fusion method based on the Gaussian copula. The proposed fusion model is integrated in a statistical framework with the hidden Markov field inference in order to delineate a target volume from multi-source images. Estimation of parameters of the models and segmentation of the images are jointly performed by an iterative algorithm based on Gibbs sampling. Experiments are performed on multi-sequence MRI to segment tumors. The results show that the proposed method based on the Gaussian copula is effective to accomplish multi-source image segmentation.
Jérôme Lapuyade-Lahorgue, Jing-Hao Xue, Su Ruan
IEEE Trans. Image Process.2
2017 Matched Shrunken Cone Detector (MSCD): Bayesian Derivations and Case Studies for Hyperspectral Target Detection
abstract
Hyperspectral images (HSIs) possess non-negative properties for both hyperspectral signatures and abundance coefficients, which can be naturally modeled using cone-based representation. However, in hyperspectral target detection, cone-based methods are barely studied. In this paper, we propose a new regularized cone-based representation approach to hyperspectral target detection, as well as its two working models by incorporating into the cone representation l2-norm and l1-norm regularizations, respectively. We call the new approach the matched shrunken cone detector (MSCD). Also important, we provide principled derivations of the proposed MSCD from the Bayesian perspective: we show that MSCD can be derived by assuming a multivariate half-Gaussian distribution or a multivariate half-Laplace distribution as the prior distribution of the coefficients of the models. In the experimental studies, we compare the proposed MSCD with the subspace methods and the sparse representation-based methods for HSI target detection. Two real hyperspectral data sets are used for evaluating the detection performances on sub-pixel targets and full-pixel targets, respectively. Results show that the proposed MSCD can outperform other methods in both cases, demonstrating the competitiveness of the regularized cone-based representation.
Ziyu Wang 0003, Rui Zhu 0006, Kazuhiro Fukui, Jing-Hao Xue
IEEE Trans. Image Process.4
2016 Latent variable pictorial structure for human pose estimation on depth images
Guijin Wang, Qingmin Liao, Jing-Hao Xue
Neurocomputing4
2016 Combining multiple biometric traits with an order-preserving score fusion algorithm
Yicong Liang, Xiaoqing Ding, Changsong Liu, Jing-Hao Xue
Neurocomputing4
2016 Randomized time warping for motion recognition
Chendra Hadi Suryanto, Jing-Hao Xue, Kazuhiro Fukui
Image Vis. Comput.2
2016 A novel hierarchical framework for human action recognition
Hongzhao Chen, Guijin Wang, Jing-Hao Xue
Pattern Recognit.3
2016 Sufficient Canonical Correlation Analysis
abstract
Canonical correlation analysis (CCA) is an effective way to find two appropriate subspaces in which Pearson's correlation coefficients are maximized between projected random vectors. Due to its well-established theoretical support and relatively efficient computation, CCA is widely used as a joint dimension reduction tool and has been successfully applied to many image processing and computer vision tasks. However, as reported, the traditional CCA suffers from overfitting in many practical cases. In this paper, we propose sufficient CCA (S-CCA) to relieve CCA's overfitting problem, which is inspired by the theory of sufficient dimension reduction. The effectiveness of S-CCA is verified both theoretically and experimentally. Experimental results also demonstrate that our S-CCA outperforms some of CCA's popular extensions during the prediction phase, especially when severe overfitting occurs.
Yiwen Guo, Xiaoqing Ding, Changsong Liu, Jing-Hao Xue
IEEE Trans. Image Process.4
2015 MiLDA: A graph embedding approach to multi-view face recognition
Yiwen Guo, Xiaoqing Ding, Jing-Hao Xue
Neurocomputing3
2015 Depth-images-based pose estimation using regression forests and graphical models
Guijin Wang, Qingmin Liao, Jing-Hao Xue
Neurocomputing4
2015 Embedding metric learning into set-based face recognition for video surveillance
Guijin Wang, Chenbo Shi, Jing-Hao Xue
Neurocomputing4
2015 Why Does Rebalancing Class-Unbalanced Data Improve AUC for Linear Discriminant Analysis?
abstract
Many established classifiers fail to identify the minority class when it is much smaller than the majority class. To tackle this problem, researchers often first rebalance the class sizes in the training dataset, through oversampling the minority class or undersampling the majority class, and then use the rebalanced data to train the classifiers. This leads to interesting empirical patterns. In particular, using the rebalanced training data can often improve the area under the receiver operating characteristic curve (AUC) for the original, unbalanced test data. The AUC is a widely-used quantitative measure of classification performance, but the property that it increases with rebalancing has, as yet, no theoretical explanation. In this note, using Gaussian-based linear discriminant analysis (LDA) as the classifier, we demonstrate that, at least for LDA, there is an intrinsic, positive relationship between the rebalancing of class sizes and the improvement of AUC. We show that the largest improvement of AUC is achieved, asymptotically, when the two classes are fully rebalanced to be of equal sizes.
Jing-Hao Xue, Peter Hall 0003
IEEE Trans. Pattern Anal. Mach. Intell.1
2015 Subcategory Clustering with Latent Feature Alignment and Filtering for Object Detection
abstract
For objects with large appearance variations, it has been proved that their detection performance can be effectively improved by clustering positive training instances into subcategories and learning multi-component models for the subcategories. However, it is not trivial to generate subcategories of high quality, due to the difficulty in measuring the similarity between positive instances. In this letter we propose a new weakly supervised clustering method to achieve better sub-categorization. Our method provides a more precise measurement of the similarity by aligning the positive instances through latent variables and filtering the aligned features. As a better alternative to the initialization step of the latent-SVM algorithm for the learning of the multi-component models, our method can lead to a superior performance gain for object detection. We demonstrate this on various real-world datasets.
Zhiwei Ruan, Guijin Wang, Jing-Hao Xue, Xinggang Lin
IEEE Signal Process. Lett.3
2015 Advanced Joint Bayesian Method for Face Verification
abstract
Generative Bayesian models have recently become the most promising framework in classifier design for face verification. However, we report in this paper that the joint Bayesian method, a successful classifier in this framework, suffers performance degradation due to its underuse of the expectation-maximization algorithm in its training phase. To rectify the underuse, we propose a new method termed advanced joint Bayesian (AJB). AJB has a good convergence property and achieves a higher verification rate than both the Joint Bayesian method and other state-of-the-art classifiers on the labeled faces in the wild face database.
Yicong Liang, Xiaoqing Ding, Jing-Hao Xue
IEEE Trans. Inf. Forensics Secur.3
2014 Fisher's linear discriminant embedded metric learning
Yiwen Guo, Xiaoqing Ding, Chi Fang, Jing-Hao Xue
Neurocomputing4
2014 Detection of user-registered dog faces
Zhiwei Ruan, Guijin Wang, Jing-Hao Xue, Xinggang Lin
Neurocomputing3
2012 Ridler and Calvard's, Kittler and Illingworth's and Otsu's methods for image thresholding
Jing-Hao Xue, Yu-Jin Zhang
Pattern Recognit. Lett.1
2011 Median-based image thresholding
Jing-Hao Xue, D. M. Titterington
Image Vis. Comput.1
2011 t -Tests, F -Tests and Otsu's Methods for Image Thresholding
abstract
Otsu's binarization method is one of the most popular image-thresholding methods; Student's t -test is one of the most widely-used statistical tests to compare two groups. This paper aims to stress the equivalence between Otsu's binarization method and the search for an optimal threshold that provides the largest absolute Student's t-statistic. It is then naturally demonstrated that the extension of Otsu's binarization method to multi-level thresholding is equivalent to the search for optimal thresholds that provide the largest F -statistic through one-way analysis of variance (ANOVA). Furthermore, general equivalences between some parametric image-thresholding methods and the search for optimal thresholds with the largest likelihood-ratio test statistics are briefly discussed.
Jing-Hao Xue, D. M. Titterington
IEEE Trans. Image Process.1
2010 Joint discriminative-generative modelling based on statistical tests for classification
Jing-Hao Xue, D. M. Titterington
Pattern Recognit. Lett.1
2009 Interpretation of hybrid generative/discriminative algorithms
Jing-Hao Xue, D. M. Titterington
Neurocomputing1
2008 Comment on "On Discriminative vs. Generative Classifiers: A Comparison of Logistic Regression and Naive Bayes"
Jing-Hao Xue, D. M. Titterington
Neural Process. Lett.1
2008 Do unbalanced data have a negative effect on LDA?
Jing-Hao Xue, D. M. Titterington
Pattern Recognit.1
2008 Short note on two output-dependent hidden Markov models
Jing-Hao Xue, D. M. Titterington
Pattern Recognit. Lett.1
2005 Making XML document markup international
abstract
In name and in practice, the World-Wide Web (hereafter Web) is used around the World beyond English-speaking areas. This creates a tremendous need to internationalize standard terminology used in the technologies that make the Web possible. Existing efforts on XML internationalization (i18n) and localization (i10n) have focused on the content of XML documents instead of the terms used in markup (annotations) such as elements and attributes. The SGML standard ISO 8879 supports the use of Unicode (ISO 10646) throughout a document, including markups. However, most elements and attributes of XML documents are still defined in English, thereby limiting their use among non-English speakers. This paper presents an XSLT-based method that can completely localize the markup of XML documents into different natural languages. We also describe how the proposed technique can be applied to translation problems in programming (e.g. C and Java) or documentation (e.g. LATEX or other formatting languages) so that a program or a document can be converted to and from an XML format. Copyright © 2004 John Wiley & Sons, Ltd.
Yijun Yu 0001, Jianguo Lu, John Mylopoulos, Weiwei Sun 0008, Jing-Hao Xue, Erik H. D'Hollander
Softw. Pract. Exp.5
2003 An integrated method of adaptive enhancement for unsupervised segmentation of MRI brain images
Jing-Hao Xue, Aleksandra Pizurica, Wilfried Philips, Etienne E. Kerre, Rik Van de Walle, Ignace Lemahieu
Pattern Recognit. Lett.1
2001 A novel method for adaptive enhancement and unsupervised segmentation of MRI brain image
abstract
This paper describes a novel global-to-local method for the adaptive enhancement and unsupervised segmentation of brain tissues in MRI (magnetic resonance imaging) images. Three brain tissues are of interest: CSF (cerebrospinal fluid), GM (gray matter), WM (white matter). Firstly, we de-noise the image using wavelet thresholding, and segment the image with minimum error thresholding. Both the thresholdings are global-wise. Subsequently, we combine locally adaptive weighted median and weighted average filters with FCM (fuzzy C-means) clustering to achieve a local-wise segmentation. The performance of the proposed method is quantitatively validated by four indices with respect to a MRI brain phantom.
Jing-Hao Xue, Wilfried Philips, Aleksandra Pizurica, Ignace Lemahieu
ICASSP1
2001 Knowledge-based segmentation and labeling of brain structures from MRI images
Jing-Hao Xue, Su Ruan, Bruno Moretti, Marinette Revenu, Daniel Bloyet
Pattern Recognit. Lett.1
2000 Fuzzy Modeling of Knowledge for MRI Brain Structure Segmentation
abstract
In this paper, we propose a novel automatic method based on fuzzy modeling of knowledge to segment brain structures in MRI (magnetic resonance imaging) images. The segmentation is achieved by the region-wise classification using GAs (genetic algorithms), followed by voxel-wise refinement using parallel region growing. To improve the accuracy of the labeling, we introduce a fuzzy model of ROI (regions of interest) by analogy with the electrostatic potential distribution, to represent more appropriately knowledge of shape, distance and reaction between structures, and to estimate more reliably the statistical moments. This modeling is also used in the design of the fitness function of GAs, and the criteria of region growing. The performance of our proposed method is quantitatively validated by 4 indexes with respect to manually segmented images.
Jing-Hao Xue, Su Ruan, Bruno Moretti, Marinette Revenu, Daniel Bloyet, Wilfried Philips
ICIP1
2000 Unsupervised Segmentation of Three-Dimensional Brain Images
abstract
This paper presents an unsupervised segmentation method applied to classify brain tissues in 3D for magnetic resonance (MR) images. An MR image volume may be composed of a mixture of several tissue types due to partial volume effects. The statistical model of the mixtures is proposed and studied by means of simulations. It is shown that it can be approximated by a Gaussian function under some conditions. The D'Agostino-Pearson normality test is used to calculate the risk /spl alpha/ of the approximation. In order to classify a brain into three brain tissues and deal with the problem of partial volume effects, the proposed algorithm classifies firstly the brain into pure classes and mix-classes, it then re-classifies the mix-classes into pure classes by adding the knowledge about the topology of the brain, based on the multifractal dimension. Both steps use Markov random field models. The algorithm is evaluated using both simulated images and real MR images.
Su Ruan, Mohamed-Jalal Fadili, Daniel Bloyet, Jing-Hao Xue
ICPR4
2000 Brain Tissue Classification of Magnetic Resonance Images Using Partial Volume Modeling
abstract
This paper presents a fully automatic three-dimensional classification of brain tissues for Magnetic Resonance (MR) images. An MR image volume may be composed of a mixture of several tissue types due to partial volume effects. Therefore, we consider that in a brain dataset there are not only the three main types of brain tissue: gray matter, white matter, and cerebro spinal fluid, called pure classes, but also mixtures, called mixclasses. A statistical model of the mixtures is proposed and studied by means of simulations. It is shown that it can be approximated by a Gaussian function under some conditions. The D'Agostino-Pearson normality test is used to assess the risk alpha of the approximation. In order to classify a brain into three types of brain tissue and deal with the problem of partial volume effects, the proposed algorithm uses two steps: 1) segmentation of the brain into pure and mixclasses using the mixture model; 2) reclassification of the mixclasses into the pure classes using knowledge about the obtained pure classes. Both steps use Markov random field (MRF) models. The multifractal dimension, describing the topology of the brain, is added to the MRFs to improve discrimination of the mixclasses. The algorithm is evaluated using both simulated images and real MR images with different T1-weighted acquisition sequences.
Su Ruan, Cyril Jaggi, Jing-Hao Xue, Mohamed-Jalal Fadili, Daniel Bloyet
IEEE Trans. Medical Imaging3