Xiaohua Xie

dblp:22/5763 · DBLP profile ↗
← Back
149ranked-venue papers
9as first author
101since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 105 · 7 first-author · 66 since 2021Artificial intelligence and machine learning · 80 · 4 first-author · 59 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Security and privacy · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2027 Uncertainty-aware prototype consolidation for unsupervised aerial person re-identification under varying heights
Zhizhi Lu, Yuli Huang, Xiaohua Xie, Jian-Huang Lai
Expert Syst. Appl.4
2026 Rethinking security of diffusion-based generative steganography
Jiahao Zhu 0005, Lingxiao Yang, Weiqi Luo 0001, Xiaohua Xie
Inf. Sci.7
2026 Neural Prediction Errors as a Unified Cue for Abstract Visual Reasoning
abstract
Humans exhibit remarkable abilities in recognizing relationships and performing complex reasoning. In contrast, deep neural networks have long been critiqued for their limitations in abstract visual reasoning (AVR), a key challenge in achieving artificial general intelligence. Drawing on the well-known concept of prediction errors from neuroscience, we propose that prediction errors can serve as a unified mechanism for both supervised and self-supervised learning in AVR. In our novel supervised learning model, AVR is framed as a prediction-and-matching process, where the central component is the discrepancy (i.e., prediction error) between a predicted feature based on abstract rules and candidate features within a reasoning context. In the self-supervised model, prediction errors as a key component unify the learning and inference processes. Both supervised and self-supervised prediction-based models achieve state-of-the-art performance on a broad range of AVR datasets and task conditions. Most notably, hierarchical prediction errors in the supervised model automatically decrease during training, an emergent phenomenon closely resembling the decrease of dopamine signals observed in biological learning. These findings underscore the critical role of prediction errors in AVR and highlight the potential of leveraging neuroscience theories to advance computational models for high-level cognition in artificial intelligence.
Lingxiao Yang, Xiaohua Xie, Wei-Shi Zheng 0001, Ru-Yuan Zhang
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 MMA++: Effective Multi-Modal Adaptation for Vision-Language Models
abstract
Large scale pre-trained Vision-Language Models (VLMs) have shown good generalization capabilities across diverse downstream tasks. However, adapting such large-scale models to few-shot generalization scenarios remains challenging due to the trade-off between preserving general knowledge and incorporating task-specific information. In this paper, we propose MMA++, an advanced and effective Multi-Modal Adapter framework for parameter-efficient VLM adaptation. Unlike prior works that independently inject adapters into each modality or uniformly across layers, MMA++ performs a dataset-level analysis to identify discriminative and generalizable features, and selectively applies adapters to the higher layers of both vision and text encoders. To bridge the modality gap, we further propose a shared feature projection space that enhances alignment between modalities. Beyond architecture design, we identify the fusion scale $\alpha$α-which controls the strength of adapter integration-as a key factor in few-shot generalization. We empirically and theoretically demonstrate that $\alpha$α should not be static, but adapted based on training data size. To reduce the effort of tuning this value across different datasets, we propose the $\alpha$α-consistency framework, consisting of: (1) a consistency training strategy under varying fusion scales; and (2) an $\alpha$α-decoupling strategy that uses a larger fusion scale during training and a smaller one at inference to account for sample size mismatch. We evaluate MMA++ on a wide range of few-shot generalization tasks, including base-to-novel generalization, cross-dataset transfer, and domain generalization. Our method consistently achieves leading performance.
Lingxiao Yang, Ru-Yuan Zhang, Yanchen Wang, Xiaohua Xie
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 CLIP -powered modality centering with spiral training for visible-infrared person re-identification
Jianghao Xiong, Xiaohua Xie, Qinyu Feng, Jian-Huang Lai
Pattern Recognit.2
2026 Dual-level aggregation network for video-based visible-infrared group re-identification
Jianghao Xiong, Xiaohua Xie, Jian-Huang Lai
Pattern Recognit.2
2026 Progressive NeRF Training With Dynamic Frequency Allocation for Sparse-View Synthesis
abstract
NeRF's performance degrades under sparse view conditions due to overfitting to limited input data. Recent works have alleviated this issue by linearly expanding the visible frequency bands of NeRF's inputs as training progresses, which means they allocate equal time to low frequency and high frequency components. However, we observe that learning low frequency components requires less training time, whereas learning high frequency components demands more. Based on this insight, we propose a Dynamic Position Encoding (DPE) mask that allocates additional optimization time to higher frequency bands as they are gradually activated. To further improve the efficiency of NeRF training, we introduce a Multi-Stage Training (MST) paradigm: we generate four RGB image scales via three$2\times$downsamplings of training images and NeRF is optimized from the lowest resolution to the original resolution in the final stage for high rendering quality. Extensive experiments on the LLFF and DTU benchmarks demonstrate that our approach not only synthesizes high-fidelity renderings with only three input views but also achieves a remarkable$3.1\times$reduction in training time.
Meng Pan, Xiaohua Xie
IEEE Signal Process. Lett.2
2026 Toward Multi-Source Sky-Ground Re-Identification: A New Benchmark and an Innovative Approach
abstract
Person re-identification (Re-ID) aims to accurately match pairs of person images across different cameras. Existing Re-ID methods primarily focus on associations within single-type camera networks (e.g., ground-ground or sky-sky matching), which are ineffective in addressing the significant viewpoint discrepancies in multi-type camera networks. One key reason for this is the absence of suitable large-scale datasets for algorithm evaluation, which limits the applicability of Re-ID across more diverse scenarios, despite its critical importance. To expand the scope of visual coverage and facilitate search operations in the special locations, we construct a novel benchmark: Multi-Source Sky-Land person Re-ID dataset (MSSL), including 66,928 images from 2,099 volunteers in nearly 20 unique scenes. Additionally, we observe that existing Re-ID systems struggle with drastic viewpoint variations in sky-ground Re-ID, especially on MSSL. To address these issues, we propose Multi-Source Prompts (MSP), separately learning finer cross-modal features of pedestrians from both sky and ground perspectives. These refined features better represent the true appearance of pedestrians from different viewpoints. Subsequently, we employ Multi-Source Alignment Loss to mitigate the impact of drastic viewpoint changes. Extensive experiments demonstrate that our approach achieves state-of-the-art performance on our MSSL, as well as on other benchmarks such as AG-ReID dataset. Our MSSL dataset and the code will be available athttps://github.com/sysuchx/SkyGroundReID.
Zhizhi Lu, Nailong Zhao, Yuli Huang, Lingxiao Yang, Xiaohua Xie, Jian-Huang Lai
IEEE Trans. Multim.6
2025 Mind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive Jailbreaking
abstract
Recently, the Image Prompt Adapter (IP-Adapter) has been increasingly integrated into text-to-image diffusion models (T2I-DMs) to improve controllability. However, in this paper, we reveal that T2I-DMs equipped with the IP-Adapter (T2I-IP-DMs) enable a new jailbreak attack named the hijacking attack. We demonstrate that, by uploading imperceptible image-space adversarial examples (AEs), the adversary can hijack massive benign users to jailbreak an Image Generation Service (IGS) driven by T2I-IP-DMs and mislead the public to discredit the service provider. Worse still, the IP-Adapter’s dependency on open-source image encoders reduces the knowledge required to craft AEs. Extensive experiments verify the technical feasibility of the hijacking attack. In light of the revealed threat, we investigate several existing defenses and explore combining the IP-Adapter with adversarially trained models to overcome existing defenses’ limitations. Our code is available at https://github.com/fhdnskfbeuv/attackIPA.
Junxi Chen, Junhao Dong 0001, Xiaohua Xie
CVPR3
2025 GuardSplat: Efficient and Robust Watermarking for 3D Gaussian Splatting
abstract
3D Gaussian Splatting (3DGS) has recently created impressive 3D assets for various applications. However, considering security, capacity, invisibility, and training efficiency, the copyright of 3DGS assets is not well protected as existing watermarking methods are unsuited for its rendering pipeline. In this paper, we propose GuardSplat, an innovative and efficient framework for watermarking 3DGS assets. Specifically, 1) We propose a CLIP-guided pipeline for optimizing the message decoder with minimal costs. The key objective is to achieve high-accuracy extraction by leveraging CLIP’s aligning capability and rich representations, demonstrating exceptional capacity and efficiency. 2) We tailor a Spherical-Harmonic-aware (SH-aware) Message Embedding module for 3DGS, seamlessly embedding messages into the SH features of each 3D Gaussian while preserving the original 3D structure. This enables watermarking 3DGS assets with minimal fidelity trade-offs and prevents malicious users from removing the watermarks from the model files, meeting the demands for invisibility and security. 3) We present an Anti-distortion Message Extraction module to improve robustness against various distortions. Experiments demonstrate that GuardSplat outperforms state-of-the-art and achieves fast optimization speed.
Guangcong Wang, Jiahao Zhu 0005, Jian-Huang Lai, Xiaohua Xie
CVPR5
2025 LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
abstract
Recent open-vocabulary detectors achieve promising performance with abundant region-level annotated data. In this work, we show that an open-vocabulary detector co-training with a large language model by generating image-level detailed captions for each image can further improve performance. To achieve the goal, we first collect a dataset, GroundingCap-1M, wherein each image is accompanied by associated grounding labels and an image-level detailed caption. With this dataset, we finetune an open-vocabulary detector with training objectives including a standard grounding loss and a caption generation loss. We take advantage of a large language model to generate both region-level short captions for each region of interest and image-level long captions for the whole image. Under the supervision of the large language model, the resulting detector, LLMDet, outperforms the baseline by a clear margin, enjoying superior open-vocabulary ability. Further, we show that the improved LLMDet can in turn build a stronger large multi-modal model, achieving mutual benefits. The code, model, and dataset are available at https://github.com/iSEE-Laboratory/LLMDet.
Shenghao Fu, Qize Yang, Qijie Mo, Junkai Yan, Xihan Wei, Jingke Meng, Xiaohua Xie, Wei-Shi Zheng 0001
CVPR7
2025 Training-Free Class Purification for Open-Vocabulary Semantic Segmentation
abstract
Fine-tuning pre-trained vision-language models has emerged as a powerful approach for enhancing open-vocabulary semantic segmentation (OVSS). However, the substantial computational and resource demands associated with training on large datasets have prompted interest in training-free methods for OVSS. Existing training-free approaches primarily focus on modifying model architectures and generating prototypes to improve segmentation performance. However, they often neglect the challenges posed by class redundancy, where multiple categories are not present in the current test image, and visual-language ambiguity, where semantic similarities among categories create confusion in class activation. These issues can lead to suboptimal class activation maps and affinity-refined activation maps. Motivated by these observations, we propose FreeCP, a novel training-free class purification framework designed to address these challenges. FreeCP focuses on purifying semantic categories and rectifying errors caused by redundancy and ambiguity. The purified class representations are then leveraged to produce final segmentation predictions. We conduct extensive experiments across eight benchmarks to validate FreeCP's effectiveness. Results demonstrate that FreeCP, as a plug-and-play module, significantly boosts segmentation performance when combined with other OVSS methods.
Qi Chen 0013, Lingxiao Yang, Nailong Zhao, Jian-Huang Lai, Xiaohua Xie
ICCV7
2025 ViSpeak: Visual Instruction Feedback in Streaming Videos
Shenghao Fu, Qize Yang, Yuan-Ming Li, Yi-Xing Peng, Kun-Yu Lin, Xihan Wei, Jianfang Hu, Xiaohua Xie, Wei-Shi Zheng 0001
ICCV8
2025 Adversarial Distribution Matching for Diffusion Distillation Towards Efficient Image and Video Synthesis
abstract
Distribution Matching Distillation (DMD) is a promising score distillation technique that compresses pre-trained teacher diffusion models into efficient one-step or multi-step student generators. Nevertheless, its reliance on the reverse Kullback-Leibler (KL) divergence minimization potentially induces mode collapse (or mode-seeking) in certain applications. To circumvent this inherent drawback, we propose Adversarial Distribution Matching (ADM), a novel framework that leverages diffusion-based discriminators to align the latent predictions between real and fake score estimators for score distillation in an adversarial manner. In the context of extremely challenging one-step distillation, we further improve the pre-trained generator by adversarial distillation with hybrid discriminators in both latent and pixel spaces. Different from the mean squared error used in DMD2 pre-training, our method incorporates the distributional loss on ODE pairs collected from the teacher model, and thus providing a better initialization for score distillation fine-tuning in the next stage. By combining the adversarial distillation pre-training with ADM fine-tuning into a unified pipeline termed DMDX, our proposed method achieves superior one-step performance on SDXL compared to DMD2 while consuming less GPU time. Additional experiments that apply multi-step ADM distillation on SD3-Medium, SD3.5-Large, and CogVideoX set a new benchmark towards efficient image and video synthesis.
Yanzuo Lu, Yuxi Ren, Xin Xia 0005, Shanchuan Lin, Xuefeng Xiao 0001, Andy Jinhua Ma, Xiaohua Xie, Jian-Huang Lai
ICCV8
2025 Dense2MoE: Restructuring Diffusion Transformer to MoE for Efficient Text-to-Image Generation
abstract
Diffusion Transformer (DiT) has demonstrated remarkable performance in text-to-image generation; however, its large parameter size results in substantial inference overhead. Existing parameter compression methods primarily focus on pruning, but aggressive pruning often leads to severe performance degradation due to reduced model capacity. To address this limitation, we pioneer the transformation of a dense DiT into a Mixture of Experts (MoE) for structured sparsification, reducing the number of activated parameters while preserving model capacity. Specifically, we replace the Feed-Forward Networks (FFNs) in DiT Blocks with MoE layers, reducing the number of activated parameters in the FFNs by 62.5\%. Furthermore, we propose the Mixture of Blocks (MoB) to selectively activate DiT blocks, thereby further enhancing sparsity. To ensure an effective dense-to-MoE conversion, we design a multi-step distillation pipeline, incorporating Taylor metric-based expert initialization, knowledge distillation with load balancing, and group feature loss for MoB optimization. We transform large diffusion transformers (e.g., FLUX.1 [dev]) into an MoE structure, reducing activated parameters by 60\% while maintaining original performance and surpassing pruning-based approaches in extensive experiments. Overall, Dense2MoE establishes a new paradigm for efficient text-to-image generation.
Youwei Zheng, Yuxi Ren, Xin Xia 0005, Xuefeng Xiao 0001, Xiaohua Xie
ICCV5
2025 SegmentDreamer: Towards High-Fidelity Text-to-3D Synthesis with Segmented Consistency Trajectory Distillation
abstract
Recent advancements in text-to-3D generation improve the visual quality of Score Distillation Sampling (SDS) and its variants by directly connecting Consistency Distillation (CD) to score distillation. However, due to the imbalance between self-consistency and cross-consistency, these CD-based methods inherently suffer from improper conditional guidance, leading to sub-optimal generation results. To address this issue, we present SegmentDreamer, a novel framework designed to fully unleash the potential of consistency models for high-fidelity text-to-3D generation. Specifically, we reformulate SDS through the proposed Segmented Consistency Trajectory Distillation (SCTD), effectively mitigating the imbalance issues by explicitly defining the relationship between self- and cross-consistency. Moreover, SCTD partitions the Probability Flow Ordinary Differential Equation (PF-ODE) trajectory into multiple sub-trajectories and ensures consistency within each segment, which can theoretically provide a significantly tighter upper bound on distillation error. Additionally, we propose a distillation pipeline for a more swift and stable generation. Extensive experiments demonstrate that our SegmentDreamer outperforms state-of-the-art methods in visual quality, enabling high-fidelity 3D asset creation through 3D Gaussian Splatting (3DGS).
Guangcong Wang, Xiaohua Xie, Yi Zhou 0005
ICCV4
2025 Imperceptible and Robust Adversarial Perturbation: Attention-Guided Watermark Vaccine Against Watermark Removal
abstract
Visible watermarks are generally embedded into digital images to claim their ownership for copyright protection. Unfortunately, the watermark removal models based on Deep Neural Networks (DNNs) are able to remove the watermarks from watermarked images, posing a great threat to image copyright protection. To prevent the watermark from being removed, watermark vaccines, i.e., adversarial perturbations, are usually added to the watermarked images to attack the target models, making them unable to remove the watermarks. However, the existing approaches indiscriminately add the watermark vaccine to the whole image region, and have not considered the vaccine failure caused by image noises, thereby still suffering from the issues of low imperceptibility and robustness. To address the above issues, we propose an Attention-Guided Watermark Vaccine (AGWV) scheme. Specifically, we propose pixel-feature attention (PFA) to identify the proper region for adding watermark vaccine, so as to achieve high imperceptibility for the added watermark vaccine. Then, we adopt image noises to perturb the vaccinated images and further optimize the watermark vaccine to correct the attention bias caused by image noise, thereby enhancing the robustness of watermark vaccines. Moreover, we design a vaccine evaluation model to intuitively evaluate the protective performances of watermark vaccines. Extensive experiments demonstrate that the proposed AGWV outperforms the state-of-the-arts in the aspects of both imperceptibility and robustness for defending against watermark removal models. Supplementary Material is available at https://github.com/YujiangLi0v0/ICME25.git
Yujiang Li, Zhili Zhou 0001, Zhongliang Yang, Baowei Wang, Tao Qi 0001, Xiaohua Xie, Jiantao Zhou 0001
ICME6
2025 Open-World Drone Active Tracking with Goal-Centered Rewards
abstract
Drone Visual Active Tracking aims to autonomously follow a target object by controlling the motion system based on visual observations, providing a more practical solution for effective tracking in dynamic environments. However, accurate Drone Visual Active Tracking using reinforcement learning remains challenging due to the absence of a unified benchmark and the complexity of open-world environments with frequent interference. To address these issues, we pioneer a systematic solution. First, we propose DAT, the first open-world drone active air-to-ground tracking benchmark. It encompasses 24 city-scale scenes, featuring targets with human-like behaviors and high-fidelity dynamics simulation. DAT also provides a digital twin tool for unlimited scene generation. Additionally, we propose a novel reinforcement learning method called GC-VAT, which aims to improve the performance of drone tracking targets in complex scenarios. Specifically, we design a Goal-Centered Reward to provide precise feedback across viewpoints to the agent, enabling it to expand perception and movement range through unrestricted perspectives. Inspired by curriculum learning, we introduce a Curriculum-Based Training strategy that progressively enhances the tracking performance in complex environments. Besides, experiments on simulator and real-world images demonstrate the superior performance of GC-VAT, achieving a Tracking Success Rate of approximately 72% on the simulator. The benchmark and code are available at https://github.com/SHWplus/DAT_Benchmark.
Haowei Sun, Jinwu Hu, Zhirui Zhang, Haoyuan Tian, Xinze Xie, Yufeng Wang 0004, Xiaohua Xie, Zhu Liang Yu, Mingkui Tan
NeurIPS7
2025 Global-Local Prompts-Driven Semantic Guidance for Aerial-Ground Person Re-identification
Ronghong Zhu, Xiaohua Xie, Jian-Huang Lai
PRCV (16)3
2025 Hard-Normal Example-Aware Template Mutual Matching for Industrial Anomaly Detection
Xiaohua Xie, Lingxiao Yang, Jian-Huang Lai
Int. J. Comput. Vis.2
2025 Learning with Enriched Inductive Biases for Vision-Language Models
Lingxiao Yang, Ru-Yuan Zhang, Xiaohua Xie
Int. J. Comput. Vis.4
2025 Imperceptible diffusion modification for facial privacy protection
Junhao Dong 0001, Jian-Huang Lai, Xiaohua Xie
Neurocomputing4
2025 Tracking dynamic community evolution based on Social Relevance and Strong Events
Xiaohua Xie, Chengkai Chen, Yingjie Yang
Knowl. Inf. Syst.2
2025 Unsupervised group re-identification from aerial perspective via strategic member harmonization
Xiaohua Xie, Jian-Huang Lai
Pattern Recognit.3
2025 Teacher-Student Collaboration: Effective Semi-Supervised Model for Defect Instance Segmentation
abstract
Recent defect instance segmentation methods heavily rely on pixel-level annotated images. However, acquiring labeled defect data from modern manufacturing industries takes significant time and effort. In this paper, we propose a novel semi-supervised approach for defect instance segmentation via Teacher-Student model Collaboration (TSC) to address the challenges of small defect dataset sizes and the blurring boundaries of defects. Specifically, we propose a generalized distribution fusion module (GDFM) to improve the quality of pseudo-labels. This module constructs a Gaussian mixture model to estimate the feature distributions from the student model. Leveraging Bayes’ theorem, we calculate the posterior probability, which significantly enhances the accuracy of classification pseudo-labels and refines the ambiguous regions in segmentation pseudo-labels produced by the teacher model. To manage the blurring boundaries of defects, we propose a cross-supervision contrastive learning module (CSCL). By combining the idea of online hard example mining with contrastive learning, we propose a simple yet effective method to distinguish the easy/hard and positive/negative areas of defect instances of unlabeled and labeled images. Extensive experiments demonstrate that our TSC model achieves state-of-the-art performance across three semi-supervised defect instance segmentation datasets with low annotation ratios. Note to Practitioners—The defect instance segmentation task aims to accurately locate each defect with a corresponding mask. Recent CNN models for defect instance segmentation heavily rely on pixel-level annotations, which demand significant time and effort within modern manufacturing industries. Therefore, we expand the semi-supervised framework to encompass defect instance segmentation and propose a semi-supervised approach for defect instance segmentation. We use labeled images to train the model and utilize the highly confident output of unlabeled images as pseudo-labels to improve the model instance segmentation performance. Our approach is tailored to address two prominent characteristics in detection inspection: small dataset sizes and blurring boundaries, thereby undergoing corresponding improvements. Meanwhile, our method demonstrates promising defect detection performance in real-world industrial settings with minimal annotation requirements. Extensive experiments demonstrate that our TSC model achieves state-of-the-art performance on three defect instance datasets.
Biaohua Ye, Jian-Huang Lai, Xiaohua Xie
IEEE Trans Autom. Sci. Eng.3
2025 Releasing Inequality Phenomenon in ℓ∞-Norm Adversarial Training via Input Gradient Distillation
abstract
Adversarial training (AT) is considered the most effective defense against adversarial attacks. However, a recent study revealed that ℓ∞-norm adversarial training ( ℓ∞-AT) will also induce unevenly distributed input gradients, which is called the inequality phenomenon. This phenomenon makes the ℓ∞ -norm adversarially trained model more vulnerable than the standard-trained model when high-attribution or randomly selected pixels are perturbed, enabling robust and practical closed-box attacks against ℓ∞ -adversarially trained models. In this paper, we propose a simple yet effective method called Input Gradient Distillation (IGD) to release the inequality phenomenon in ℓ∞-AT. IGD distills the standard-trained teacher model’s equal decision pattern into the ℓ∞-adversarially trained student model by aligning input gradients of the student model and the standard-trained model with the Cosine Similarity. Experiments show that IGD can mitigate the inequality phenomenon and its threats while preserving adversarial robustness. Compared to vanilla ℓ∞-AT, IGD reduces error rates against inductive noise, inductive occlusion, random noise, and noisy images in ImageNet-C by up to 60%, 16%, 50%, and 21%, respectively. Other than empirical experiments, we also conduct a theoretical analysis to explain why releasing the inequality phenomenon can improve such robustness and discuss why the severity of the inequality phenomenon varies according to the dataset’s image resolution.
Junxi Chen, Junhao Dong 0001, Xiaohua Xie, Jian-Huang Lai
IEEE Trans. Inf. Forensics Secur.3
2025 Adv-Inversion: Stealthy Adversarial Attacks via GAN-Inversion for Facial Privacy Protection
Weiqi Luo 0001, Xiaohua Xie, Peijia Zheng, Wenmin Huang, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.3
2025 Cross-Camera Pedestrian Trajectory Retrieval Based on Linear Trajectory Manifolds
abstract
The goal of pedestrian trajectory retrieval is to infer the multi-camera path of a targeted pedestrian using images or videos from a camera network, which is crucial for passenger flow analytics and individual pedestrian retrieval. Conventional approaches hinge on spatiotemporal modeling, necessitating the gathering of positional information for each camera and trajectory data between every camera pair for the training phase. To mitigate these stringent requirements, our proposed methodology employs solely temporal information for modeling. Specifically, we introduce an Implicit Trajectory Encoding scheme, dubbed Temporal Rotary Position Embedding (T-RoPE), which integrates the temporal aspects of within-camera tracklets directly into their visual representations, thereby shaping a novel feature space. Our analysis reveals that, within this refined feature space, the challenge of inter-camera trajectory extraction can be effectively addressed by delineating a linear trajectory manifold. The visual characteristics gleaned from each candidate trajectory are utilized to compare and rank against the query feature, culminating in the ultimate trajectory retrieval outcome. To validate our method, we collected a new pedestrian trajectory dataset from a multi-storey shopping mall, namely the Mall Trajectory Dataset. Extensive experimentation across diverse datasets has demonstrated the versatility of our T-RoPE module as a plug-and-play enhancement to various network architectures, significantly enhancing the precision of pedestrian trajectory retrieval tasks. The dataset and code are released at https://github.com/zhangxin1995/MTD.
Xin Zhang 0113, Xiaohua Xie, Jian-Huang Lai
IEEE Trans. Image Process.2
2025 Angular Reconstructive Discrete Embedding With Fusion Similarity for Multi-View Clustering
abstract
Effectively and efficiently mining valuable clustering patterns is a challenging problem when handling large-scale data from diverse sources. Existing approaches adopt anchor graph learning or binary representation embedding to reduce computational complexity. Normally, anchor graph learning can not directly obtain the clustering assignment except adopt the post-processing stage, such as graph cut or k-means clustering. The binary representation embedding neglects the structure information in Hamming space. In order to overcome these limitations, this paper proposes a novel, effective, and efficient angular reconstructive discrete embedding method with fusion similarity for a multi-view clustering (AFMC) that can jointly learn the global and local structure preserving binary representation and clustering assignment. Specifically, we propose to use angular reconstructive error minimization to maintain the global similarity correlation of binary representations of heterogeneous features in a common Hamming space. Moreover, we design a multi-view discrete ridge regression with fusion similarity term to handle the out-of-sample problem and preserve the local manifold structure. In addition, we propose an efficient optimization algorithm with linear computational complexity to solve the non-convex and non-smooth objective function. The experimental results demonstrate that AFMC outperforms several state-of-the-art large-scale multi-view clustering methods.
Jintang Bian, Xiaohua Xie, Chang-Dong Wang 0001, Lingxiao Yang, Jian-Huang Lai, Feiping Nie 0001
IEEE Trans. Knowl. Data Eng.2
2025 A Hierarchical Semantic Distillation Framework for Open-Vocabulary Object Detection
abstract
Open-vocabulary object detection (OVD) aims to detect objects beyond the training annotations, where detectors are usually aligned to a pre-trained vision-language model, e.g., CLIP, to inherit its generalizable recognition ability so that detectors can recognize new or novel objects. However, previous works directly align the feature space with CLIP and fail to learn the semantic knowledge effectively. In this work, we propose a hierarchical semantic distillation framework named HD-OVD to construct a comprehensive distillation process, which exploits generalizable knowledge from the CLIP model in three aspects. In the first hierarchy of HD-OVD, the detector learns fine-grainedinstance-wise semanticsfrom the CLIP image encoder by modeling relations among single objects in the visual space. Besides, we introduce text space novel-class-aware classification to help the detector assimilate the highly generalizableclass-wise semanticsfrom the CLIP text encoder, representing the second hierarchy. Lastly, abundantimage-wise semanticscontaining multi-object and their contexts are also distilled by an image-wise contrastive distillation. Benefiting from the elaborated semantic distillation in triple hierarchies, our HD-OVD inherits generalizable recognition ability from CLIP in instance, class, and image levels. Thus, we boost the novel AP on the OV-COCO dataset to 46.4% with a ResNet50 backbone, which outperforms others by a clear margin. We also conduct extensive ablation studies to analyze how each component works.
Shenghao Fu, Junkai Yan, Qize Yang, Xihan Wei, Xiaohua Xie, Wei-Shi Zheng 0001
IEEE Trans. Multim.5
2025 Multilevel Contrastive Multiview Clustering With Dual Self-Supervised Learning
abstract
Multiview clustering (MVC) aims to integrate multiple related but different views of data to achieve more accurate clustering performance. Contrastive learning has found many applications in MVC due to its successful performance in unsupervised visual representation learning. However, existing MVC methods based on contrastive learning overlook the potential of high similarity nearest neighbors as positive pairs. In addition, these methods do not capture the multilevel (i.e., cluster, instance, and prototype levels) representational structure that naturally exists in multiview datasets. These limitations could further hinder the structural compactness of learned multiview representations. To address these issues, we propose a novel end-to-end deep MVC method called multilevel contrastive MVC (MCMC) with dual self-supervised learning (DSL). Specifically, we first treat the nearest neighbors of an object from the latent subspace as the positive pairs for multiview contrastive loss, which improves the compactness of the representation at the instance level. Second, we perform multilevel contrastive learning (MCL) on clusters, instances, and prototypes to capture the multilevel representational structure underlying the multiview data in the latent space. In addition, we learn consistent cluster assignments for MVC by adopting a DSL method to associate different level structural representations. The evaluation experiment showed that MCMC can achieve intracluster compactness, intercluster separability, and higher accuracy (ACC) in clustering performance. Our code is available at https://github.com/bianjt-morning/MCMC.
Jintang Bian, Yixiang Lin, Xiaohua Xie, Chang-Dong Wang 0001, Lingxiao Yang, Jian-Huang Lai, Feiping Nie 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 Generalizable and Discriminative Representations for Adversarially Robust Few-Shot Learning
abstract
Few-shot image classification (FSIC) is beneficial for a variety of real-world scenarios, aiming to construct a recognition system with limited training data. In this article, we extend the original FSIC task by incorporating defense against malicious adversarial examples. This can be an arduous challenge because numerous deep learning-based approaches remain susceptible to adversarial examples, even when trained with ample amounts of data. Previous studies on this problem have predominantly concentrated on the meta-learning framework, which involves sampling numerous few-shot tasks during the training stage. In contrast, we propose a straightforward but effective baseline via learning robust and discriminative representations without tedious meta-task sampling, which can further be generalized to unforeseen adversarial FSIC tasks. Specifically, we introduce an adversarial-aware (AA) mechanism that exploits feature-level distinctions between the legitimate and the adversarial domains to provide supplementary supervision. Moreover, we design a novel adversarial reweighting training strategy to ameliorate the imbalance among adversarial examples. To further enhance the adversarial robustness without compromising discriminative features, we propose the cyclic feature purifier during the postprocessing projection, which can reduce the interference of unforeseen adversarial examples. Furthermore, our method can obtain robust feature embeddings that maintain superior transferability, even when facing cross-domain adversarial examples. Extensive experiments and systematic analyses demonstrate that our method achieves state-of-the-art robustness as well as natural performance among adversarially robust FSIC algorithms on three standard benchmarks by a substantial margin.
Junhao Dong 0001, Yuan Wang 0030, Xiaohua Xie, Jian-Huang Lai, Yew-Soon Ong
IEEE Trans. Neural Networks Learn. Syst.3
2025 A Soft Iterative Receiver With Simplified EP Detection for Coded MIMO Systems
abstract
Expectation propagation (EP) achieves excellent performance with high-order modulation in massive multiple-input multiple-output (MIMO) detection. The soft output of the EP detector can be iteratively combined with turbo soft decoders to enhance error-correction performance. However, the implementation of EP-based iterative detection and decoding (IDD) receivers suffer from an exponential increase in computational complexity as the number of antennas and modulation order grows. In this brief, we propose a simplified EP approximation-based IDD (sEPA-IDD) scheme for hardware implementation. To alleviate the computational burden, a simplified message update scheme is proposed, reducing complexity by 68% without performance degradation. Additionally, a unified design for extrinsic message computation further improves hardware utilization. Finally, we introduce the first unfolded EP-based IDD architecture to boost throughput. Compared with state-of-the-art (SOA) IDD receivers, the sEPA-IDD receiver implemented on 65 nm CMOS delivers a throughput of 3.07 Gb/s with a maximum 0.5 dB gain, achieving 4.03× higher throughput and 6.04× greater area efficiency.
Xiaosi Tan, Xiaohua Xie, Houren Ji, Tiancan Xia, Yongming Huang 0001, Xiaohu You 0001, Chuan Zhang 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2024 Unsupervised Group Re-identification via Adaptive Clustering-Driven Progressive Learning
abstract
Group re-identification (G-ReID) aims to correctly associate groups with the same members captured by different cameras. However, supervised approaches for this task often suffer from the high cost of cross-camera sample labeling. Unsupervised methods based on clustering can avoid sample labeling, but the problem of member variations often makes clustering unstable, leading to incorrect pseudo-labels. To address these challenges, we propose an adaptive clustering-driven progressive learning approach (ACPL), which consists of a group adaptive clustering (GAC) module and a global dynamic prototype update (GDPU) module. Specifically, GAC designs the quasi-distance between groups, thus fully capitalizing on both individual-level and holistic information within groups. In the case of great uncertainty in intra-group members, GAC effectively minimizes the impact of non-discriminative features and reduces the noise in the model's pseudo-labels. Additionally, our GDPU devises a dynamic weight to update the prototypes and effectively mine the hard samples with complex member variations, which improves the model's robustness. Extensive experiments conducted on four popular G-ReID datasets demonstrate that our method not only achieves state-of-the-art performance on unsupervised G-ReID but also performs comparably to several fully supervised approaches.
Jian-Huang Lai, Xiaohua Xie
AAAI4
2024 MLNet: Mutual Learning Network with Neighborhood Invariance for Universal Domain Adaptation
abstract
Universal domain adaptation (UniDA) is a practical but challenging problem, in which information about the relation between the source and the target domains is not given for knowledge transfer. Existing UniDA methods may suffer from the problems of overlooking intra-domain variations in the target domain and difficulty in separating between the similar known and unknown class. To address these issues, we propose a novel Mutual Learning Network (MLNet) with neighborhood invariance for UniDA. In our method, confidence-guided invariant feature learning with self-adaptive neighbor selection is designed to reduce the intra-domain variations for more generalizable feature representation. By using the cross-domain mixup scheme for better unknown-class identification, the proposed method compensates for the misidentified known-class errors by mutual learning between the closed-set and open-set classifiers. Extensive experiments on three publicly available benchmarks demonstrate that our method achieves the best results compared to the state-of-the-arts in most cases and significantly outperforms the baseline across all the four settings in UniDA. Code is available at https://github.com/YanzuoLu/MLNet.
Yanzuo Lu, Andy Jinhua Ma, Xiaohua Xie, Jian-Huang Lai
AAAI4
2024 Adversarially Robust Few-shot Learning via Parameter Co-distillation of Similarity and Class Concept Learners
abstract
Few-shot learning (FSL) facilitates a variety of computer vision tasks yet remains vulnerable to adversarial attacks. Existing adversarially robust FSL methods rely on either visual similarity learning or class concept learning. Our analysis reveals that these two learning paradigms are complementary, exhibiting distinct robustness due to their unique decision boundary types (concepts clustering by the visual similarity label vs. classification by the class labels). To bridge this gap, we propose a novel framework unifying adversarially robust similarity learning and class concept learning. Specifically, we distill parameters from both network branches into a “unified embedding model” during robust optimization and redistribute them to individual network branches periodically. To capture generalizable robustness across diverse branches, we initialize adversaries in each episode with cross-branch class-wise “global adversarial perturbations” instead of less informative random initialization. We also propose a branch robustness harmonization to modulate the optimization of similarity and class concept learners via their relative adversarial robustness. Extensive experiments demonstrate the state-of-the-art performance of our method in diverse few-shot scenarios.
Junhao Dong 0001, Piotr Koniusz, Junxi Chen, Xiaohua Xie, Yew-Soon Ong
CVPR4
2024 Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis
abstract
Diffusion model is a promising approach to image generation and has been employed for Pose-Guided Person Image Synthesis (PGPIS) with competitive performance. While existing methods simply align the person appearance to the target pose, they are prone to overfitting due to the lack of a high-level semantic understanding on the source person image. In this paper, we propose a novel Coarse-to-Fine Latent Diffusion (CFLD) method for PGPIS. In the absence of image-caption pairs and textual prompts, we de-velop a novel training paradigm purely based on images to control the generation process of a pre-trained text-to-image diffusion model. A perception-refined decoder is designed to progressively refine a set of learnable queries and extract semantic understanding of person images as a coarse-grained prompt. This allows for the decoupling of fine-grained appearance and pose information controls at different stages, and thus circumventing the potential over-fitting problem. To generate more realistic texture details, a hybrid- granularity attention module is proposed to encode multi-scale fine-grained appearance features as bias terms to augment the coarse-grained prompt. Both quantitative and qualitative experimental results on the DeepFashion benchmark demonstrate the superiority of our method over the state of the arts for PGPIS. Code is available at https://github.com/YanzuoLu/CFLD.
Yanzuo Lu, Manlin Zhang, Andy Jinhua Ma, Xiaohua Xie, Jian-Huang Lai
CVPR4
2024 MMA: Multi-Modal Adapter for Vision-Language Models
abstract
Pretrained Vision-Language Models (VLMs) have served as excellent foundation models for transfer learning in diverse downstream tasks. However, tuning VLMs for few-shot generalization tasks faces a discrimination - generalization dilemma, i.e., general knowledge should be preserved and task-specific knowledge should be fine-tuned. How to precisely identify these two types of representations remains a challenge. In this paper, we propose a Multi-Modal Adapter (MMA) for VLMs to improve the alignment between representations from text and vision branches. MMA aggregates features from different branches into a shared feature space so that gradients can be communicated across branches. To determine how to incorporate MMA, we systematically analyze the discriminability and generalizability of features across diverse datasets in both the vision and language branches, and find that (1) higher lay-ers contain discriminable dataset-specific knowledge, while lower layers contain more generalizable knowledge, and (2) language features are more discriminable than visual features, and there are large semantic gaps between the features of the two modalities, especially in the lower layers. Therefore, we only incorporate MMA to a few higher lay-ers of transformers to achieve an optimal balance between discrimination and generalization. We evaluate the effectiveness of our approach on three tasks: generalization to novel classes, novel target datasets, and domain generalization. Compared to many state-of-the-art methods, our MMA achieves leading performance in all evaluations. Code is at https://github.com/ZjjConan/Multi-Modal-Adapter
Lingxiao Yang, Ru-Yuan Zhang, Yanchen Wang, Xiaohua Xie
CVPR4
2024 View-decoupled Transformer for Person Re-identification under Aerial-ground Camera Network
abstract
Existing person re-identification methods have achieved remarkable advances in appearance-based identity association across homogeneous cameras, such as ground-ground matching. However, as a more practical scenario, aerial-ground person re-identification (AGPReID) among heterogeneous cameras has received minimal attention. To alleviate the disruption of discriminative identity representation by dramatic view discrepancy as the most significant challenge in AGPReID, the view-decoupled transformer (VDT) is proposed as a simple yet effective framework. Two major components are designed in VDT to decouple view-related and view-unrelated features, namely hierarchical subtractive separation and orthogonal loss, where the former separates these two features inside the VDT, and the latter constrains these two to be independent. In addition, we contribute a large-scale AGPReID dataset called CARGO, consisting of five/eight aerial/ground cameras, 5,000 identities, and 108,563 images. Experiments on two datasets show that VDT is a feasible and effective solution for AGPReID, surpassing the previous method on mAP/Rank1 by up to 5.0%/2.7% on CARGO and 3.7%/5.2% on AG-ReID, keeping the same magnitude of computational complexity. Our project is available at https://github.com/LinlyAC/VDT-AGPReID.
Vishal M. Patel, Xiaohua Xie, Jian-Huang Lai
CVPR4
2024 Tackling the Singularities at the Endpoints of Time Intervals in Diffusion Models
abstract
Most diffusion models assume that the reverse process adheres to a Gaussian distribution. However, this approxi-mation has not been rigorously validated, especially at sin-gularities, where t = 0 and t = 1. Improperly dealing with such singularities leads to an average brightness is-sue in applications, and limits the generation of images with extreme brightness or darkness. We primarily focus on tackling singularities from both theoretical and practi-cal perspectives. Initially, we establish the error bounds for the reverse process approximation, and showcase its Gaussian characteristics at singularity time steps. Based on this theoretical insight, we confirm the singularity at t = 1 is conditionally removable while it at t = 0 is an inherent property. Upon these significant conclusions, we propose a novel plug-and-play method SingDiffusion to address the initial singular time step sampling, which not only effectively resolves the average brightness issue for a wide range of diffusion models without extra training efforts, but also enhances their generation capability in achieving notable lower FID scores.
Pengze Zhang, Hubery Yin, Chen Li 0031, Xiaohua Xie
CVPR4
2024 Spike-Temporal Latent Representation for Energy-Efficient Event-to-Video Reconstruction
Jianxiong Tang, Jian-Huang Lai, Lingxiao Yang, Xiaohua Xie
ECCV (42)4
2024 Visible-Infrared Person Search: A Novel Benchmark and Solution
Jianghao Xiong, Xiaohua Xie, Jian-Huang Lai
ICPR (14)4
2024 Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models
abstract
Recent vision foundation models can extract universal representations and show impressive abilities in various tasks. However, their application on object detection is largely overlooked, especially without fine-tuning them. In this work, we show that frozen foundation models can be a versatile feature enhancer, even though they are not pre-trained for object detection. Specifically, we explore directly transferring the high-level image understanding of foundation models to detectors in the following two ways. First, the class token in foundation models provides an in-depth understanding of the complex scene, which facilitates decoding object queries in the detector's decoder by providing a compact context. Additionally, the patch tokens in foundation models can enrich the features in the detector's encoder by providing semantic details. Utilizing frozen foundation models as plug-and-play modules rather than the commonly used backbone can significantly enhance the detector's performance while preventing the problems caused by the architecture discrepancy between the detector's backbone and the foundation model. With such a novel paradigm, we boost the SOTA query-based detector DINO from 49.0% AP to 51.9% AP (+2.9% AP) and further to 53.8% AP (+4.8% AP) by integrating one or two foundation models respectively, on the COCO validation set after training for 12 epochs with R50 as the detector's backbone. Code will be available.
Shenghao Fu, Junkai Yan, Qize Yang, Xihan Wei, Xiaohua Xie, Wei-Shi Zheng 0001
NeurIPS5
2024 Prototype-guided domain adaptive one-stage object detector for defect detection
Biaohua Ye, Jian-Huang Lai, Xiaohua Xie, Jun-Yong Zhu
Adv. Eng. Informatics3
2024 Uncertainty Modeling for Group Re-Identification
Jian-Huang Lai, Zhan-Xiang Feng, Xiaohua Xie
Int. J. Comput. Vis.4
2024 Differentiable gated autoencoders for unsupervised feature selection
Jintang Bian, Bo Qiao 0003, Xiaohua Xie
Neurocomputing4
2024 Person group detection with global trajectory extraction in a disjoint camera network
Xin Zhang 0113, Xiaohua Xie, Jian-Huang Lai
Neurocomputing2
2024 Separable Spatial-Temporal Residual Graph for Cloth-Changing Group Re-Identification
abstract
Group re-identification (GReID) aims to correctly associate group images belonging to the same group identity, which is a crucial task for video surveillance. Existing methods only model the member feature representations inside each image (regarded as spatial members), which leads to potential failures in long-term video surveillance due to cloth-changing behaviors. Therefore, we focus on a new task called cloth-changing group re-identification (CCGReID), which needs to consider group relationship modeling in GReID and robust group representation against cloth-changing members. In this paper, we propose the separable spatial-temporal residual graph (SSRG) for CCGReID. Unlike existing GReID methods, SSRG considers both spatial members inside each group image and temporal members among multiple group images with the same identity. Specifically, SSRG constructs full graphs for each group identity within the batched data, which will be completely and non-redundantly separated into the spatial member graph (SMG) and temporal member graph (TMG). SMG aims to extract group features from spatial members, and TMG improves the robustness of the cloth-changing members by feature propagation. The separability enables SSRG to be available in the inference rather than only assisting supervised training. The residual guarantees efficient SSRG learning for SMG and TMG. To expedite research in CCGReID, we develop two datasets, including GroupPRCC and GroupVC, based on the existing CCReID datasets. The experimental results show that SSRG achieves state-of-the-art performance, including the best accuracy and low degradation (only 2.15% on GroupVC). Moreover, SSRG can be well generalized to the GReID task. As a weakly supervised method, SSRG surpasses the performance of some supervised methods and even approaches the best performance on the CSG dataset.
Jian-Huang Lai, Xiaohua Xie, Xiaofeng Jin, Sien Huang
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 A Neuroinspired Contrast Mechanism enables Few-Shot Object Detection
Lingxiao Yang, Dapeng Chen, Yifei Chen 0010, Wei Peng 0011, Xiaohua Xie
Pattern Recognit.5
2024 Benchmarking deep models on salient object detection
Huajun Zhou, Lingxiao Yang, Jian-Huang Lai, Xiaohua Xie
Pattern Recognit.5
2024 Pose Guided Person Image Generation Via Dual-Task Correlation and Affinity Learning
abstract
Pose Guided Person Image Generation (PGPIG) is the task of transforming a person's image from the source pose to a target pose. Existing PGPIG methods often tend to learn an end-to-end transformation between the source image and the target image, but do not seriously consider two issues: 1) the PGPIG is an ill-posed problem, and 2) the texture mapping requires effective supervision. In order to alleviate these two challenges, we propose a novel method by incorporating Dual-task Pose Transformer Network and Texture Affinity learning mechanism (DPTN-TA). To assist the ill-posed source-to-target task learning, DPTN-TA introduces an auxiliary task, i.e., source-to-source task, by a Siamese structure and further explores the dual-task correlation. Specifically, the correlation is built by the proposed Pose Transformer Module (PTM), which can adaptively capture the fine-grained mapping between sources and targets and can promote the source texture transmission to enhance the details of the generated images. Moreover, we propose a novel texture affinity loss to better supervise the learning of texture mapping. In this way, the network is able to learn complex spatial transformations effectively. Extensive experiments show that our DPTN-TA can produce perceptually realistic person images under significant pose changes. Furthermore, our DPTN-TA is not limited to processing human bodies but can be flexibly extended to view synthesis of other objects, i.e., faces and chairs, outperforming the state-of-the-arts in terms of both LPIPS and FID. Our code is available at: https://github.com/PangzeCheung/Dual-task-Pose-Transformer-Network.
Pengze Zhang, Lingxiao Yang, Xiaohua Xie, Jian-Huang Lai
IEEE Trans. Vis. Comput. Graph.3
2023 3DP Code-Based Compression and AR Visualization for Cardiovascular Palpation Training
Kaifeng Gong, Yinan Hao, Xiaohua Xie
CGI (3)5
2023 The Enemy of My Enemy is My Friend: Exploring Inverse Adversaries for Improving Adversarial Training
abstract
Although current deep learning techniques have yielded superior performance on various computer vision tasks, yet they are still vulnerable to adversarial examples. Adversarial training and its variants have been shown to be the most effective approaches to defend against adversarial examples. A particular class of these methods regularize the difference between output probabilities for an adversarial and its corresponding natural example. However, it may have a negative impact if a natural example is misclassified. To circumvent this issue, we propose a novel adversarial training scheme that encourages the model to produce similar output probabilities for an adversarial example and its “inverse adversarial” counterpart. Particularly, the counterpart is generated by maximizing the likelihood in the neighborhood of the natural example. Extensive experiments on various vision datasets and architectures demonstrate that our training method achieves state-of-the-art robustness as well as natural accuracy among robust models. Furthermore, using a universal version of inverse adversarial examples, we improve the performance of single-step adversarial training techniques at a low computational cost.
Junhao Dong 0001, Seyed-Mohsen Moosavi-Dezfooli, Jian-Huang Lai, Xiaohua Xie
CVPR4
2023 Texture-Guided Saliency Distilling for Unsupervised Salient Object Detection
abstract
Deep Learning-based Unsupervised Salient Object Detection (USOD) mainly relies on the noisy saliency pseudo labels that have been generated from traditional handcraft methods or pre-trained networks. To cope with the noisy labels problem, a class of methods focus on only easy samples with reliable labels but ignore valuable knowledge in hard samples. In this paper, we propose a novel USOD method to mine rich and accurate saliency knowledge from both easy and hard samples. First, we propose a Confidence-aware Saliency Distilling (CSD) strategy that scores samples conditioned on samples' confidences, which guides the model to distill saliency knowledge from easy samples to hard samples progressively. Second, we propose a Boundary-aware Texture Matching (BTM) strategy to refine the boundaries of noisy labels by matching the textures around the predicted boundaries. Extensive experiments on RGB, RGB-D, RGB-T, and video SOD benchmarks prove that our method achieves state-of-the-art USOD performance. Code is available at www.github.com/moothes/A2S-v2.
Huajun Zhou, Bo Qiao 0003, Lingxiao Yang, Jian-Huang Lai, Xiaohua Xie
CVPR5
2023 CuNeRF: Cube-Based Neural Radiance Field for Zero-Shot Medical Image Arbitrary-Scale Super Resolution
abstract
Medical image arbitrary-scale super-resolution (MIASSR) has recently gained widespread attention, aiming to supersample medical volumes at arbitrary scales via a single model. However, existing MIASSR methods face two major limitations: (i) reliance on high-resolution (HR) volumes and (ii) limited generalization ability, which restricts their applications in various scenarios. To overcome these limitations, we propose Cube-based Neural Radiance Field (CuNeRF), a zero-shot MIASSR framework that is able to yield medical images at arbitrary scales and free viewpoints in a continuous domain. Unlike existing MISR methods that only fit the mapping between low-resolution (LR) and HR volumes, CuNeRF focuses on building a continuous volumetric representation from each LR volume without the knowledge of the corresponding HR one. This is achieved by the proposed differentiable modules: cube-based sampling, isotropic volume rendering, and cube-based hierarchical rendering. Through extensive experiments on magnetic resource imaging (MRI) and computed tomography (CT) modalities, we demonstrate that CuNeRF can synthesize high-quality SR medical images, which outperforms state-of-the-art MISR methods, achieving better visual verisimilitude and fewer objectionable artifacts. Compared to existing MISR methods, our CuNeRF is more applicable in practice.
Lingxiao Yang, Jian-Huang Lai, Xiaohua Xie
ICCV4
2023 ASAG: Building Strong One-Decoder-Layer Sparse Detectors via Adaptive Sparse Anchor Generation
abstract
Recent sparse detectors with multiple, e.g. six, decoder layers achieve promising performance but much inference time due to complex heads. Previous works have explored using dense priors as initialization and built one-decoder-layer detectors. Although they gain remarkable acceleration, their performance still lags behind their six-decoder-layer counterparts by a large margin. In this work, we aim to bridge this performance gap while retaining fast speed. We find that the architecture discrepancy between dense and sparse detectors leads to feature conflict, hampering the performance of one-decoder-layer detectors. Thus we propose Adaptive Sparse Anchor Generator (ASAG) which predicts dynamic anchors on patches rather than grids in a sparse way so that it alleviates the feature conflict problem. For each image, ASAG dynamically selects which feature maps and which locations to predict, forming a fully adaptive way to generate image-specific anchors. Further, a simple and effective Query Weighting method eases the training instability from adaptiveness. Extensive experiments show that our method outperforms dense-initialized ones and achieves a better speed-accuracy trade-off. The code is available at https://github.com/iSEE-Laboratory/ASAG.
Shenghao Fu, Junkai Yan, Yipeng Gao, Xiaohua Xie, Wei-Shi Zheng 0001
ICCV4
2023 Neural Prediction Errors enable Analogical Visual Reasoning in Human Standard Intelligence Tests
abstract
Deep neural networks have long been criticized for lacking the ability to perform analogical visual reasoning. Here, we propose a neural network model to solve Raven's Progressive Matrices (RPM) - one of the standard intelligence tests in human psychology. Specifically, we design a reasoning block based on the well-known concept of prediction error (PE) in neuroscience. Our reasoning block uses convolution to extract abstract rules from high-level visual features of the 8 context images and generates the features of a predicted answer. PEs are then calculated between the predicted features and those of the 8 candidate answers, and are then passed to the next stage. We further integrate our novel reasoning blocks into a residual network and build a new Predictive Reasoning Network (PredRNet). Extensive experiments show that our proposed PredRNet achieves state-of-the-art average performance on several important RPM benchmarks. PredRNet also shows good generalization abilities in a variety of out-of-distribution scenarios and other visual reasoning tasks. Most importantly, our PredRNet forms low-dimensional representations of abstract rules and minimizes hierarchical prediction errors during model training, supporting the critical role of PE minimization in visual reasoning. Our work highlights the potential of using neuroscience theories to solve abstract visual reasoning problems in artificial intelligence. The code is available at https://github.com/ZjjConan/AVR-PredRNet.
Lingxiao Yang, Hongzhi You, Zonglei Zhen, Dahui Wang, Xiaohong Wan, Xiaohua Xie, Ru-Yuan Zhang
ICML6
2023 Spike Count Maximization for Neuromorphic Vision Recognition
abstract
Spiking Neural Networks (SNNs) are the promising models of neuromorphic vision recognition. The mean square error (MSE) and cross-entropy (CE) losses are widely applied to supervise the training of SNNs on neuromorphic datasets. However, the relevance between the output spike counts and predictions is not well modeled by the existing loss functions. This paper proposes a Spike Count Maximization (SCM) training approach for the SNN-based neuromorphic vision recognition model based on optimizing the output spike counts. The SCM is achieved by structural risk minimization (SRM) and a specially designed spike counting loss. The spike counting loss counts the output spikes of the SNN by using the L0-norm, and the SRM maximizes the distance between the margin boundaries of the classifier to ensure the generalization of the model. The SCM is non-smooth and non-differentiable, and we design a two-stage algorithm with fast convergence to solve the problem. Experiment results demonstrate that the SCM performs satisfactorily in most cases. Using the output spikes for prediction, the accuracies of SCM are 2.12%~16.50% higher than the popular training losses on the CIFAR10-DVS dataset. The code is available at https://github.com/TJXTT/SCM-SNN.
Jianxiong Tang, Jian-Huang Lai, Xiaohua Xie, Lingxiao Yang
IJCAI3
2023 RuleMatch: Matching Abstract Rules for Semi-supervised Learning of Human Standard Intelligence Tests
abstract
Raven's Progressive Matrices (RPM), one of the standard intelligence tests in human psychology, has recently emerged as a powerful tool for studying abstract visual reasoning (AVR) abilities in machines. Although existing computational models for RPM problems achieve good performance, they require a large number of labeled training examples for supervised learning. In contrast, humans can efficiently solve unlabeled RPM problems after learning from only a few example questions. Here, we develop a semi-supervised learning (SSL) method, called RuleMatch, to train deep models with a small number of labeled RPM questions along with other unlabeled questions. Moreover, instead of using pixel-level augmentation in object perception tasks, we exploit the nature of RPM problems and augment the data at the level of abstract rules. Specifically, we disrupt the possible rules contained among context images in an RPM question and force the two augmented variants of the same unlabeled sample to obey the same abstract rule and predict a common pseudo label for training. Extensive experiments show that the proposed RuleMatch achieves state-of-the-art performance on two popular RAVEN datasets. Our work makes an important stride in aligning abstract analogical visual reasoning abilities in machines and humans. Our Code is at https://github.com/ZjjConan/AVR-RuleMatch.
Yunlong Xu 0001, Lingxiao Yang, Hongzhi You, Zonglei Zhen, Da-Hui Wang, Xiaohong Wan, Xiaohua Xie, Ru-Yuan Zhang
IJCAI7
2023 Few Shot Object Detection with Incompletely Annotated Samples
abstract
Few shot object detection aims to generalize the model to previously unseen classes with only a few training samples, which has been attached great attention due to its practicability in real scenes. Many existing methods hold the missed detection issue, mainly due to the problem of incompletely annotated samples. Specifically, only partial objects in a sample are labeled, resulting in unlabeled objects being used as the background for training. This problem is especially serious for few-shot learning. To solve this noisy label problem, we first propose a label calibration method based on confidence to correct potential incorrect labels, and then introduce the class center library to eliminate the negative impact of unlabeled objects. We conduct extensive experiments on PASCAL VOC and MS-COCO benchmarks, which proves the effectiveness of our approach and achieves the state-of-the-art results.
Bo Qiao 0003, Huajun Zhou, Lingxiao Yang, Xiaohua Xie
IJCNN4
2023 Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion
abstract
Emotional Voice Conversion aims to manipulate a speech according to a given emotion while preserving non-emotion components. Existing approaches cannot well express fine-grained emotional attributes. In this paper, we propose an Attention-based Interactive diseNtangling Network (AINN) that leverages instance-wise emotional knowledge for voice conversion. We introduce a two-stage pipeline to effectively train our network: Stage I utilizes inter-speech contrastive learning to model fine-grained emotion and intra-speech disentanglement learning to better separate emotion and content. In Stage II, we propose to regularize the conversion with a multi-view consistency mechanism. This technique helps us transfer fine-grained emotion and maintain speech content. Extensive experiments show that our AINN outperforms state-of-the-arts in both objective and subjective metrics.
Lingxiao Yang, Qi Chen 0013, Jian-Huang Lai, Xiaohua Xie
INTERSPEECH5
2023 Formulating Discrete Probability Flow Through Optimal Transport
abstract
Continuous diffusion models are commonly acknowledged to display a deterministic probability flow, whereas discrete diffusion models do not. In this paper, we aim to establish the fundamental theory for the probability flow of discrete diffusion models. Specifically, we first prove that the continuous probability flow is the Monge optimal transport map under certain conditions, and also present an equivalent evidence for discrete cases. In view of these findings, we are then able to define the discrete probability flow in line with the principles of optimal transport. Finally, drawing upon our newly established definitions, we propose a novel sampling method that surpasses previous discrete diffusion models in its ability to generate more certain outcomes. Extensive experiments on the synthetic toy dataset and the CIFAR-10 dataset have validated the effectiveness of our proposed discrete probability flow. Code is released at: https://github.com/PangzeCheung/Discrete-Probability-Flow.
Pengze Zhang, Hubery Yin, Xiaohua Xie
NeurIPS4
2023 Modality Balancing Mechanism for RGB-Infrared Object Detection in Aerial Image
Weibo Cai, Junhao Dong 0001, Jian-Huang Lai, Xiaohua Xie
PRCV (12)5
2023 Feature Disentanglement and Adaptive Fusion for Improving Multi-modal Tracking
Weibo Cai, Junhao Dong 0001, Jian-Huang Lai, Xiaohua Xie
PRCV (12)5
2023 Multi-Resolution Edge-aware Lighting Enhancement Network
Wenyong Gong, Wenzhu Chen, Zhongwei Yu, Xiaohua Xie
Comput. Graph.4
2023 AC2AS: Activation Consistency Coupled ANN-SNN framework for fast and memory-efficient SNN training
Jianxiong Tang, Jian-Huang Lai, Xiaohua Xie, Lingxiao Yang, Wei-Shi Zheng 0001
Pattern Recognit.3
2023 Activation to Saliency: Forming High-Quality Labels for Unsupervised Salient Object Detection
abstract
This paper focuses on the Unsupervised Salient Object Detection (USOD) issue. We come up with a two-stage Activation-to-Saliency (A2S) framework that effectively excavates saliency cues to train a robust saliency detector. It is worth noting that our method does not require any manual annotation in the whole process. In the first stage, we transform an unsupervisedly pre-trained network to aggregate multi-level features into a single activation map, where an Adaptive Decision Boundary (ADB) is proposed to assist the training of the transformed network. Moreover, a new loss function is proposed to facilitate the generation of high-quality pseudo labels. In the second stage, a self-rectification learning strategy is developed to train a saliency detector and refine the pseudo labels online. In addition, we construct a lightweight saliency detector using two Residual Attention Modules (RAMs) to learn robust saliency information. Extensive experiments on several SOD benchmarks prove that our framework reports significant performance compared with existing USOD methods. Moreover, training our framework on 3,000 images consumes about 1 hour, which is over 10 times faster than previous state-of-the-art methods. Code has been published athttps://github.com/moothes/A2S-USOD.
Huajun Zhou, Peijia Chen, Lingxiao Yang, Xiaohua Xie, Jian-Huang Lai
IEEE Trans. Circuits Syst. Video Technol.4
2023 Restricted Black-Box Adversarial Attack Against DeepFake Face Swapping
abstract
DeepFake face swapping presents a significant threat to online security and social media, which can replace the source face in an arbitrary photo/video with the target face of an entirely different person. In order to prevent this fraud, some researchers have begun to study the adversarial methods against DeepFake or face manipulation. However, existing works mainly focus on the white-box setting or the black-box setting driven by abundant queries, which severely limits the practical application of these methods. To tackle this problem, we introduce a practical adversarial attack that does not require any queries to the facial image forgery model. Our method is built on a substitute model based on face reconstruction and then transfers adversarial examples from the substitute model directly to inaccessible black-box DeepFake models. Specially, we propose the Transferable Cycle Adversary Generative Adversarial Network (TCA-GAN) to construct the adversarial perturbation for disrupting unknown DeepFake systems. We also present a novel post-regularization module for enhancing the transferability of generated adversarial examples. To comprehensively measure the effectiveness of our approaches, we construct a challenging baseline of DeepFake adversarial attacks for future development. Extensive experiments impressively show that the proposed adversarial attack method makes the visual quality of DeepFake face images plummet so that they are easier to be detected by humans and algorithms. Moreover, we demonstrate that the proposed algorithm can be generalized to offer face image protection against various face translation methods.
Junhao Dong 0001, Yuan Wang 0030, Jian-Huang Lai, Xiaohua Xie
IEEE Trans. Inf. Forensics Secur.4
2023 Toward Intrinsic Adversarial Robustness Through Probabilistic Training
abstract
Modern deep neural networks have made numerous breakthroughs in real-world applications, yet they remain vulnerable to some imperceptible adversarial perturbations. These tailored perturbations can severely disrupt the inference of current deep learning-based methods and may induce potential security hazards to artificial intelligence applications. So far, adversarial training methods have achieved excellent robustness against various adversarial attacks by involving adversarial examples during the training stage. However, existing methods primarily rely on optimizing injective adversarial examples correspondingly generated from natural examples, ignoring potential adversaries in the adversarial domain. This optimization bias can induce the overfitting of the suboptimal decision boundary, which heavily jeopardizes adversarial robustness. To address this issue, we propose Adversarial Probabilistic Training (APT) to bridge the distribution gap between the natural and adversarial examples via modeling the latent adversarial distribution. Instead of tedious and costly adversary sampling to form the probabilistic domain, we estimate the adversarial distribution parameters in the feature level for efficiency. Moreover, we decouple the distribution alignment based on the adversarial probability model and the original adversarial example. We then devise a novel reweighting mechanism for the distribution alignment by considering the adversarial strength and the domain uncertainty. Extensive experiments demonstrate the superiority of our adversarial probabilistic training method against various types of adversarial attacks in different datasets and scenarios.
Junhao Dong 0001, Lingxiao Yang, Yuan Wang 0030, Xiaohua Xie, Jian-Huang Lai
IEEE Trans. Image Process.4
2023 Learning Shadow Removal From Unpaired Samples via Reciprocal Learning
abstract
We focus on addressing the problem of shadow removal for an image, and attempt to make a weakly supervised learning model that does not depend on the pixelwise-paired training samples, but only uses the samples with image-level labels that indicate whether an image contains shadow or not. To this end, we propose a deep reciprocal learning model that interactively optimizes the shadow remover and the shadow detector to improve the overall capability of the model. On the one hand, shadow removal is modeled as an optimization problem with a latent variable of the detected shadow mask. On the other hand, a shadow detector can be trained using the prior from the shadow remover. A self-paced learning strategy is employed to avoid fitting to intermediate noisy annotation during the interactive optimization. Furthermore, a color-maintenance loss and a shadow-attention discriminator are both designed to facilitate model optimization. Extensive experiments on the pairwise ISTD dataset, SRD dataset, and unpaired USR dataset demonstrate the superiority of the proposed deep reciprocal model.
Xiaohua Xie, Kuoyu Deng, Lingxiao Yang, Jian-Huang Lai
IEEE Trans. Image Process.2
2023 Learning Weak Semantics by Feature Graph for Attribute-Based Person Search
abstract
Attribute-based person search aims to find the target person from the gallery images based on the given query text. It often plays an important role in surveillance systems when visual information is not reliable, such as identifying a criminal from a few witnesses. Although recent works have made great progress, most of them neglect the attribute labeling problems that exist in the current datasets. Moreover, these problems also increase the risk of non-alignment between attribute texts and visual images, leading to large semantic gaps. To address these issues, in this paper, we propose Weak Semantic Embeddings (WSEs), which can modify the data distribution of the original attribute texts and thus improve the representability of attribute features. We also introduce feature graphs to learn more collaborative and calibrated information. Furthermore, the relationship modeled by our feature graphs between all semantic embeddings can reduce the semantic gap in text-to-image retrieval. Extensive evaluations on three challenging benchmarks - PETA, Market-1501 Attribute, and PA100K, demonstrate the effectiveness of the proposed WSEs, and our method outperforms existing state-of-the-art methods.
Qiyang Peng, Lingxiao Yang, Xiaohua Xie, Jian-Huang Lai
IEEE Trans. Image Process.3
2023 Wavelet-Guided Promotion-Suppression Transformer for Surface-Defect Detection
abstract
Surface-defect detection aims to accurately locate and classify defect areas in images via pixel-level annotations. Different from the objects in traditional image segmentation, defect areas comprise a small group of pixels with random shapes, characterized by uncommon textures and edges that are inconsistent with the normal surface patterns of industrial products. This task-specific knowledge is hardly considered in the current methods. Therefore, we propose a two-stage "promotion-suppression" transformer (PST) framework, which explicitly adopts the wavelet features to guide the network to focus on the detailed features in the images. Specifically, in the promotion stage, we propose the Haar augmentation module to improve the backbone's sensitivity to high-frequency details. However, the background noise is inevitably amplified as well because it also constitutes high-frequency information. Therefore, a quadratic feature-fusion module (QFFM) is proposed in the suppression stage, which exploits the two properties of noise: independence and attenuation. The QFFM analyzes the similarities and differences between noise and defect features to achieve noise suppression. Compared with the traditional linear-fusion approach, the QFFM is more sensitive to high-frequency details; thus, it can afford highly discriminative features. Extensive experiments are conducted on three datasets, namely DAGM, MT, and CRACK500, which demonstrate the superiority of the proposed PST framework.
Jian-Huang Lai, Jun-Yong Zhu, Xiaohua Xie
IEEE Trans. Image Process.4
2023 Cross-Camera Trajectories Help Person Retrieval in a Camera Network
abstract
We are concerned with retrieving a query person from multiple videos captured by a non-overlapping camera network. Existing methods often rely on purely visual matching or consider temporal constraints but ignore the spatial information of the camera network. To address this issue, we propose a pedestrian retrieval framework based on cross-camera trajectory generation that integrates both temporal and spatial information. To obtain pedestrian trajectories, we propose a novel cross-camera spatio-temporal model that integrates pedestrians' walking habits and the path layout between cameras to form a joint probability distribution. Such a cross-camera spatio-temporal model can be specified using sparsely sampled pedestrian data. Based on the spatio-temporal model, cross-camera trajectories can be extracted by the conditional random field model and further optimised by restricted non-negative matrix factorization. Finally, a trajectory re-ranking technique is proposed to improve the pedestrian retrieval results. To verify the effectiveness of our method, we construct the first cross-camera pedestrian trajectory dataset, the Person Trajectory Dataset, in real surveillance scenarios. Extensive experiments verify the effectiveness and robustness of the proposed method.
Xin Zhang 0113, Xiaohua Xie, Jian-Huang Lai, Wei-Shi Zheng 0001
IEEE Trans. Image Process.2
2022 Uncertainty Modeling with Second-Order Transformer for Group Re-identification
abstract
Group re-identification (G-ReID) focuses on associating the group images containing the same persons under different cameras. The key challenge of G-ReID is that all the cases of the intra-group member and layout variations are hard to exhaust. To this end, we propose a novel uncertainty modeling, which treats each image as a distribution depending on the current member and layout, then digs out potential group features by random samplings. Based on potential and original group features, uncertainty modeling can learn better decision boundaries, which is implemented by two modules, member variation module (MVM) and layout variation module (LVM). Furthermore, we propose a novel second-order transformer framework (SOT), which is inspired by the fact that the position modeling in the transformer is coped with the G-ReID task. SOT is composed of the intra-member module and inter-member module. Specifically, the intra-member module extracts the first-order token for each member, and then the inter-member module learns a second-order token as a group feature by the above first-order tokens, which can be regarded as the token of tokens. A large number of experiments have been conducted on three available datasets, including CSG, DukeGroup and RoadGroup. The experimental results show that the proposed SOT outperforms all previous state-of-the-art methods.
Jian-Huang Lai, Zhan-Xiang Feng, Xiaohua Xie
AAAI4
2022 Self-supervised Image-specific Prototype Exploration for Weakly Supervised Semantic Segmentation
abstract
Weakly Supervised Semantic Segmentation (WSSS) based on image-level labels has attracted much attention due to low annotation costs. Existing methods often rely on Class Activation Mapping (CAM) that measures the correlation between image pixels and classifier weight. However, the classifier focuses only on the discriminative regions while ignoring other useful information in each image, resulting in incomplete localization maps. To address this issue, we propose a Self-supervised Image-specific Prototype Exploration (SIPE) that consists of an Image-specific Prototype Exploration (IPE) and a General-Specific Consistency (GSC) loss. Specifically, IPE tailors prototypes for every image to capture complete regions, formed our Image-Specific CAM (IS-CAM), which is realized by two sequential steps. In addition, GSC is proposed to construct the consistency of general CAM and our specific IS-CAM, which further optimizes the feature representation and empowers a self-correction ability of prototype exploration. Extensive experiments are conducted on PASCAL VOC 2012 and MS COCO 2014 segmentation benchmark and results show our SIPE achieves new state-of-the-art performance using only image-level labels. The code is available at https://github.com/chenqi1126/SIPE.
Qi Chen 0013, Lingxiao Yang, Jian-Huang Lai, Xiaohua Xie
CVPR4
2022 Improving Adversarially Robust Few-shot Image Classification with Generalizable Representations
abstract
Few-Shot Image Classification (FSIC) aims to recognize novel image classes with limited data, which is significant in practice. In this paper, we consider the FSIC problem in the case of adversarial examples. This is an extremely challenging issue because current deep learning methods are still vulnerable when handling adversarial examples, even with massive labeled training samples. For this problem, existing works focus on training a network in the meta-learning fashion that depends on numerous sampled few-shot tasks. In comparison, we propose a simple but effective baseline through directly learning generalizable representations without tedious task sampling, which is robust to unforeseen adversarial FSIC tasks. Specifically, we introduce an adversarial-aware mechanism to establish auxiliary supervision via feature-level differences between legitimate and adversarial examples. Furthermore, we design a novel adversarial-reweighted training manner to alleviate the imbalance among adversarial examples. The feature purifier is also employed as post-processing for adversarial features. Moreover, our method can obtain generalizable representations to remain superior transferability, even facing cross-domain adversarial examples. Extensive experiments show that our method can significantly outperform state-of-the-art adversarially robust FSIC methods on two standard benchmarks.
Junhao Dong 0001, Yuan Wang 0030, Jian-Huang Lai, Xiaohua Xie
CVPR4
2022 Modeling 3D Layout For Group Re-Identification
abstract
Group re-identification (GReID) attempts to correctly associate groups with the same members under different cameras. The main challenge is how to resist the membership and layout variations. Existing works attempt to incorporate layout modeling on the basis of appearance features to achieve robust group representations. However, layout ambiguity is introduced because these methods only consider the 2D layout on the imaging plane. In this paper, we overcome the above limitations by 3D layout modeling. Specifically, we propose a novel 3D transformer (3DT) that reconstructs the relative 3D layout relationship among members, then applies sampling and quantification to preset a series of layout tokens along three dimensions, and selects the corresponding tokens as layout features for each member. Furthermore, we build a synthetic GReID dataset, City1M, including 1.84M images, 45K persons and 11.5K groups with 3D annotations to alleviate data shortages and poor annotations. To the best of our knowledge, 3DT is the first work to address GReID with 3D perspective, and the City1M is the currently largest dataset. Several experiments show the superiority of our 3DT and City1M. Our project has been released on https://github.com/LinlyAC/City1M-dataset.
Kaiheng Dang, Jian-Huang Lai, Zhan-Xiang Feng, Xiaohua Xie
CVPR5
2022 Exploring Dual-task Correlation for Pose Guided Person Image Generation
abstract
Pose Guided Person Image Generation (PGPIG) is the task of transforming a person image from the source pose to a given target pose. Most of the existing methods only focus on the ill-posed source-to-target task and fail to capture reasonable texture mapping. To address this problem, we propose a novel Dual-task Pose Transformer Network (DPTN), which introduces an auxiliary task (i.e., source-to-source task) and exploits the dual-task correlation to promote the performance of PGPIG. The DPTN is of a Siamese structure, containing a source-to-source self-reconstruction branch, and a transformation branch for source-to-target generation. By sharing partial weights between them, the knowledge learned by the source-to-source task can effectively assist the source-to-target learning. Furthermore, we bridge the two branches with a proposed Pose Transformer Module (PTM) to adaptively explore the correlation between features from dual tasks. Such correlation can establish the fine-grained mapping of all the pixels between the sources and the targets, and promote the source texture transmission to enhance the details of the generated target images. Extensive experiments show that our DPTN outperforms state-of-the-arts in terms of both PSNR and LPIPS. In addition, our DPTN only contains 9.79 million parameters, which is significantly smaller than other approaches. Our code is available at: https://github.com/PangzeCheung/Dual-task-Pose-Transformer-Network.
Pengze Zhang, Lingxiao Yang, Jian-Huang Lai, Xiaohua Xie
CVPR4
2022 Decoupled Contrastive Learning for Intra-Camera Supervised Person Re-identification
abstract
Intra-camera supervised (ICS) person re-identification (Re-ID) assumes that people are annotated independently in each camera, which is a newly proposed setting to reduce the cost of manual annotation. Most existing methods are developed on generating global pseudo labels to learn camera-agnostic features by classification loss. However, they do not utilize intra-camera labels well for cross-camera association. In this paper, we propose a Decoupled Contrastive Learning (DCL) strategy to tackle the issue. Concretely, we first conduct an intra-camera pre-train stage to reduce the intra-identity variance. Then an inter-identity association and an inter-camera learning step are alternatively iterated to improve the feature representation gradually. We propose a decoupled identity contrastive loss in the inter-camera learning stage to make the training more effective. In order to improve the compactness of the associated cluster, we further adopt a hard-aware contrastive loss. Extensive experiments on two large-scale Re-ID datasets demonstrate that the proposed method outperforms most ICS methods and performs comparably to fully supervised methods.
Shiteng Hu, Xiaohua Xie
ICPR3
2022 Variance of Local Contribution: an Unsupervised Image Quality Assessment for Face Recognition
abstract
In recent years, Face Image Quality Assessment (FIQA) plays an important role in the face recognition system. However, how to define face image quality is still an open question. In this work, we argue that a high-quality face image should have more identity-related information than a low-quality face image. Thus, we propose a novel unsupervised Face Image Quality Assessment with the variance of local contribution (VLC-FIQA). In our approach, we alternately mask partial pixels of the face image, then quantify the importance of these pixels and compute the variation of the importance of different parts as the quality of the image. Extensive experiments show that our VLC-FIQA outperforms state-of-the-art approaches on LFW. Our approach can be easily used for any recognition system and be extended to other recognition tasks such as person re-identification.
Qiye Lian, Xiaohua Xie, Huicheng Zheng, Yongdong Zhang 0002
ICPR2
2022 Learning Bi-directional Feature Propagation with Latent Layout Modeling for Group Re-identification
abstract
Group re-identification (G-ReID) aims to identify the same group of persons across the disjoint cameras. The key challenge of G-ReID is the robust feature extraction against the potential group layout and membership varitions. However, previous works focus more on the appearance modeling and less on the importance of group layout. In this paper, we propose a bi-directional feature propagation framework, which propagates information between group layout and member appearance. In addition, we propose the spatial generation framework, which analyses the group image and generates new images with different group layouts to simulate various layouts in the real world. Moreover, we propose a network that learns latent layout representations and propagates the layout representations with the member appearance representations. The proposed network achieves SOTA performance on two widely used G-ReID datasets, i.e., 87.9% mAP and 89.2% Rank-1 on CSG, 92.7% mAP and 90.1% Rank-1 on RoadGroup.
Yuan Wang 0030, Jian-Huang Lai, Xiaohua Xie, Junhao Dong 0001
ICPR4
2022 Cross-level Attention and Ratio Consistency Network for Ship Detection
abstract
In ship detection task, target objects with extreme aspect ratios are common in practical applications. However, existing ship detection methods seldom make efforts to tackle this issue. In this paper, we present a novel Cross-level Attention and Ratio Consistency (CARC) Network for ship detection. First, we propose a Cross-Level Attention (CLA) module to generate attention signals by integrating information from both higher and lower level features. Specifically, for each feature, we calculate its similarity with features from adjacent levels. These similarities are utilized as weights to enhance the channels that consist of different-level information. By fusing multi-level information, a channel-wise attention vector is generated to enhance the learned representations in the base feature. Second, we propose a Ratio Consistency loss that promotes the networks to localize the ships with more accurate aspect ratios. Existing ship detection methods have different sensitiveness to width and height predictions, significantly increasing the learning difficulty for localizing target ships. We append an auxiliary supervision signal to the detection head in our method. This supervision signal measures the error between the predicted and the ground truth aspect ratios of target ships. Experiment results show that our model achieves significant performance gains compared to existing methods.
Biaohua Ye, Huajun Zhou, Jian-Huang Lai, Xiaohua Xie
ICPR5
2022 Learning Contextual Embedding Deep Networks for Accurate and Efficient Image Deraining
Guangguang Yang, Jun Chen 0013, Xiaohua Xie, Jian-Huang Lai
PRCV (4)4
2022 Improving Pre-trained Masked Autoencoder via Locality Enhancement for Person Re-identification
Yanzuo Lu, Manlin Zhang, Andy Jinhua Ma, Xiaohua Xie, Jian-Huang Lai
PRCV (2)5
2022 Relaxation LIF: A gradient-based spiking neuron for direct training deep spiking neural networks
Jianxiong Tang, Jian-Huang Lai, Wei-Shi Zheng 0001, Lingxiao Yang, Xiaohua Xie
Neurocomputing5
2022 Successive Consensus Clustering for Unsupervised Video-Based Person Re-Identification
abstract
Person re-identification is to match the same person between non-overlapping cameras. This paper focuses on unsupervised video-based person re-identification. The mainstream approach is to obtain pseudo-labels by clustering samples for training the classification model. In this scheme, a potential threat is that noisy pseudo-labels may damage the optimization of the model. To mitigate this danger, we propose using a Successive Consensus Clustering framework for optimizing the pseudo-labels and the model iteratively. First, we leverage consensus clustering with respect to multiple frames of a video, which can generate high-quality pseudo-labels for pedestrians. Secondly, we develop contrastive learning based on the cluster successive memory mechanism, which can establish the correlation between different epochs of clustering so that makes the training of the model stable. Experiments on three large-scale data sets show that our method outperforms the previous state-of-the-art method, surpassing 10.6% for rank-1 and 18.6% for mAP on Mars, and 9.6% for rank-1 and 13.3% for mAP on DukeMTMC-VideoReID.
Jinhao Qian, Xiaohua Xie
IEEE Signal Process. Lett.2
2022 Lightweight Texture Correlation Network for Pose Guided Person Image Generation
abstract
Pose Guided Person Image Generation (PGPIG) is a popular task in deepfake, which aims at generating a person image with the given pose based on the source image. However, existing methods cannot comprehensively model the correlation between the source and the target domain. Most of them only focus on the correlation of the keypoints but ignore detail textures. In this paper, we propose a novel Texture Correlation Network (TCN) to simultaneously build pose and texture correlations. Specifically, our TCN adopts a two-stage design, including two networks: Pose Guided Person Alignment Network (PGPAN) and Texture Correlation Attention Network (TCAN). The PGPAN generates a coarse person image aligned with the target pose, while the TCAN produces a target generated image with the guidance of multiple correlations. The key component of TCAN is our new module, Texture Correlation Attention Module (TCAM), which explicitly builds geometry and texture correlation between the source image and the coarse target image. Those kinds of correlations facilitate to transfer real textures from the source to the target. Extensive experiments on the DeepFashion and Market1501 benchmarks demonstrate the superior performance of the proposed method. In addition, our model only uses 8.5 million parameters, which is significantly smaller than other methods.
Pengze Zhang, Lingxiao Yang, Xiaohua Xie, Jian-Huang Lai
IEEE Trans. Circuits Syst. Video Technol.3
2022 Selective Intra-Image Similarity for Personalized Fixation-Based Object Segmentation
abstract
Personalized Fixation-based Object Segmentation (PFOS) aims at segmenting the gazed objects in images conditioned on personalized fixations. However, the performances of existing PFOS methods are degraded when facing anomalous fixation maps (some fixations fall in the background) or enormous objects because of their poor localization ability. In this paper, we propose a novel Selective Intra-image Similarity Network (SISNet) that achieves significant performance by precisely localizing the gazed objects. First, we propose a Response Purifying Module (RPM) to eliminate the false response regions caused by anomalous fixations in the background. By suppressing these false responses, we can significantly reduce the negative impacts caused by anomalous fixations. Second, we propose an intra-image similarity module (ISM) to better localize large objects by integrating more long-range information. In addition, we propose a new Discriminative Intersection-over-Union metric that evaluates whether PFOS methods can produce distinctive predictions for varying fixations. Experiments on the PFOS and our proposed OSIE-CFPS-UN datasets prove that our network achieves remarkable improvements and outperforms existing state-of-the-art methods. Code has been published athttps://www.github.com/moothes/SISNet.
Huajun Zhou, Lingxiao Yang, Xiaohua Xie, Jian-Huang Lai
IEEE Trans. Circuits Syst. Video Technol.3
2022 Seeing Like a Human: Asynchronous Learning With Dynamic Progressive Refinement for Person Re-Identification
abstract
Learning discriminative and rich features is an important research task for person re-identification. Previous studies have attempted to capture global and local features at the same time and layer of the model in a non-interactive manner, which are called synchronous learning. However, synchronous learning leads to high similarity, and further defects in model performance. To this end, we propose asynchronous learning based on the human visual perception mechanism. Asynchronous learning emphasizes the time asynchrony and space asynchrony of feature learning and achieves mutual promotion and cyclical interaction for feature learning. Furthermore, we design a dynamic progressive refinement module to improve local features with the guidance of global features. The dynamic property allows this module to adaptively adjust the network parameters according to the input image, in both the training and testing stage. The progressive property narrows the semantic gap between the global and local features, which is due to the guidance of global features. Finally, we have conducted several experiments on four datasets, including Market1501, CUHK03, DukeMTMC-ReID, and MSMT17. The experimental results show that asynchronous learning can effectively improve feature discrimination and achieve strong performance.
Jian-Huang Lai, Zhan-Xiang Feng, Xiaohua Xie
IEEE Trans. Image Process.4
2021 Hole Detection with Texture-Suppression on Wooden Plate Surfaces
Xiaojie An, Xiaohua Xie, Xiang Chen 0007
ICIG (1)2
2021 Visually Maintained Image Disturbance Against Deepfake Face Swapping
abstract
As a deep learning-based application, DeepFake can generate malicious images or videos through replacing the face of a source image with the target face, which poses a significant threat to social media. In this paper, we propose a scheme to prevent such tampering by exploring adversarial examples against DeepFake. Specifically, adversarial examples are produced by adding tailored distortion to source images. The added distortion is imperceptible to human vision but can mislead the generation of face-swapped images effectively. We present three novel adversarial attacks against DeepFake autoencoders from perspectives of adversarial transferability and latent representation. Our first method synthesizes universal perturbation, which is image-agnostic. By contrast, the latter two methods directly perform the preciser perturbation specific to a source image. Extensive experiments demonstrate the effectiveness of our adversarial examples against DeepFake in terms of both reference and non-reference image quality assessment.
Junhao Dong 0001, Xiaohua Xie
ICME2
2021 SimAM: A Simple, Parameter-Free Attention Module for Convolutional Neural Networks
abstract
In this paper, we propose a conceptually simple but very effective attention module for Convolutional Neural Networks (ConvNets). In contrast to existing channel-wise and spatial-wise attention modules, our module instead infers 3-D attention weights for the feature map in a layer without adding parameters to the original networks. Specifically, we base on some well-known neuroscience theories and propose to optimize an energy function to find the importance of each neuron. We further derive a fast closed-form solution for the energy function, and show that the solution can be implemented in less than ten lines of code. Another advantage of the module is that most of the operators are selected based on the solution to the defined energy function, avoiding too many efforts for structure tuning. Quantitative evaluations on various visual tasks demonstrate that the proposed module is flexible and effective to improve the representation ability of many ConvNets. Our code is available at Pytorch-SimAM.
Lingxiao Yang, Ru-Yuan Zhang, Lida Li, Xiaohua Xie
ICML4
2021 Scale-Aware Multi-branch Decoder for Salient Object Detection
Huajun Zhou, Xiaohua Xie, Jian-Huang Lai
PRCV (1)3
2021 Flounder-Net: An efficient CNN for crowd counting by aerial photography
Shengjie Xiu, Xiang Chen 0007, Xiaohua Xie
Neurocomputing5
2021 Optical Flow Estimation Based on the Frequency-Domain Regularization
abstract
Accurate optical flow estimation with the frequency-domain regularization is a challenging problem in computer vision. In this paper, we solve this issue by introducing a novel optical flow method related to the frequency domain that uses TV-wavelet regularization. Specifically, we regard TV-wavelet regularization as a filtering process. After wavelet transform for optical flow field, we firstly remove outliers by performing a threshold operation. Then, we make up for lost motion information (such as flow edges and important motion details) determined by these missing or damaged wavelet coefficients by adding TV-wavelet coefficients that are obtained from transform spectrum of the prior flow geometrical features, which are controlled by the image structures. By combining the advantages of total variation to recover geometric structures with the strengths of wavelet representation to remove outliers, the proposed method significantly outperforms the current frequency-domain optical flow methods in removing outliers, preserving sharp flow edges, and restoring important motion details. It also shows competitive optical flow evaluation results on the challenging MPI-Sintel, Kitti, and Middlebury datasets.
Jun Chen 0013, Jian-Huang Lai, Xiaohua Xie
IEEE Trans. Circuits Syst. Video Technol.4
2021 Weakly Supervised Learning for Raindrop Removal on a Single Image
abstract
In this paper, we address a problem of view-disturbing raindrop removal on a single image. In existing methods to tackle this problem, machine learning based ones seem promising but require elaborate pairwise images, i.e., the raindrop-degraded image and the corresponding clean image of the same scene, for training. To overcome this drawback, we propose a weakly supervised learning based model in the absence of pairwise training examples, which needs only a collection of images with image-level annotations indicating the presence/absence of raindrops for training. Specifically, we train a raindrop detector for highlighting regions of raindrops in a multi-task learning manner. Then, we propose an attention-based generative network for raindrop removal and introduce a weighted preservation loss to retain the non-raindrop details. Specially, our model can be mixedly trained with pairwise and unpaired samples, which enables us to conveniently adapt the model to a new domain. Experiments verify the effect of the proposed method. Especially, using only weakly-supervised learning, our method can achieve comparable results with state-of-the-art strongly-supervised learning methods.
Jian-Huang Lai, Xiaohua Xie
IEEE Trans. Circuits Syst. Video Technol.3
2021 Contour-Aware Loss: Boundary-Aware Learning for Salient Object Segmentation
abstract
We present a learning model that makes full use of boundary information for salient object segmentation. Specifically, we come up with a novel loss function, i.e., Contour Loss, which leverages object contours to guide models to perceive salient object boundaries. Such a boundary-aware network can learn boundary-wise distinctions between salient objects and background, hence effectively facilitating the salient object segmentation. Yet the Contour Loss emphasizes the boundaries to capture the contextual details in the local range. We further propose the hierarchical global attention module (HGAM), which forces the model hierarchically to attend to global contexts, thus captures the global visual saliency. Comprehensive experiments on six benchmark datasets show that our method achieves superior performance over state-of-the-art ones. Moreover, our model has a real-time speed of 26 fps on a TITAN X GPU.
Huajun Zhou, Jian-Huang Lai, Lingxiao Yang, Xiaohua Xie
IEEE Trans. Image Process.5
2021 Resolution-Aware Knowledge Distillation for Efficient Inference
abstract
Minimizing the computation complexity is essential for the popularization of deep networks in practical applications. Nowadays, most researches attempt to accelerate deep networks by designing new network structure or compressing the network parameters. Meanwhile, transfer learning techniques such as knowledge distillation are utilized to keep the performance of deep models. In this paper, we focus on accelerating deep models and relieving the computation burden by using low-resolution (LR) images as inputs while maintaining competitive performance, which is rarely researched in the current literature. Deep networks may encounter serious performance degradation when using LR inputs because many details are unavailable from LR images. Besides, the existing approaches may fail to learn discriminative features for LR images because of the dramatic appearance variations between LR and high-resolution (HR) images. To tackle with the above problems, we propose a resolution-aware knowledge distillation (RKD) framework to narrow the cross-resolution variations by transferring knowledge from HR domain to LR domain. The proposed framework consists of a HR teacher network and a LR student network. First, we introduce a discriminator and propose an adversarial learning strategy to shrink the variations between inputs with changing resolution. Then we design a cross-resolution knowledge distillation (CRKD) loss to train discriminative student network by exploiting the knowledge of the teacher network. The CRKD loss is consisted of a resolution-aware distillation loss, a pair-wise constraint, and a maximum mean discrepancy loss. Experimental results on person re-identification, image classification, face recognition, and defect segmentation tasks demonstrate that RKD outperforms traditional knowledge distillation method by achieving better performance with lower computation complexities. Furthermore, CRKD surpasses the state-of-the-art knowledge distillation methods in transferring knowledge across different resolutions under RKD framework, especially when coping with large resolution differences.
Zhan-Xiang Feng, Jian-Huang Lai, Xiaohua Xie
IEEE Trans. Image Process.3
2021 Homogeneous-to-Heterogeneous: Unsupervised Learning for RGB-Infrared Person Re-Identification
abstract
RGB-Infrared (RGB-IR) cross-modality person re-identification (re-ID) is attracting more and more attention due to requirements for 24-h scene surveillance. However, the high cost of labeling person identities of an RGB-IR dataset largely limits the scalability of supervised models in real-world scenarios. In this paper, we study the unsupervised RGB-IR person re-ID problem (or briefly uRGB-IR re-ID) in which no identity annotations are available in RGB-IR cross-modality datasets. Considering that intra-modality (i.e., RGB-RGB or IR-IR) re-ID is much easier than cross-modality re-ID and can provide shared knowledge for RGB-IR re-ID, we propose a two-stage method to solve the uRGB-IR re-ID, namely homogeneous-to-heterogeneous learning. In the first stage, the unsupervised self-learning method is conducted to learn the intra-modality feature representation and to generate the pseudo-labeled identities of person images separately for each modality. In the second stage, heterogeneous learning is used to learn a shared discriminative feature representation by distilling the knowledge from intra-modality pseudo-labels, to align two modalities via a modality-based consistent learning module, and finally to target modality-invariant learning via a pseudo-labeled positive instance selection module. With the use of homogeneous-to-heterogeneous learning, the proposed unsupervised framework greatly reduces the modality gap and thus learns a robust feature representation against RGB and infrared modalities, leading to promising accuracy. We also propose a novel cross-modality re-ranking approach that includes a self-modality search and a cycle-modality search to tailor the uRGB-IR re-ID. Unlike conventional re-ranking, the proposed re-ranking method takes a modality-based constraint into re-ranking and thus can select more reliable nearest neighbors, which greatly improves uRGB-IR re-ID. The experimental results demonstrate the superiority of our approach on the SYSU-MM01 and RegDB datasets.
Wenqi Liang, Guangcong Wang, Jian-Huang Lai, Xiaohua Xie
IEEE Trans. Image Process.4
2021 RelightGAN: Instance-level Generative Adversarial Network for Face Illumination Transfer
abstract
Face illumination perception and processing is a significantly difficult issue especially due to asymmetric shadings, local highlights, and local shadows. This study focuses on the face illumination transfer problem, which is to transfer the illumination style from a reference face image to a target face image while preserving other attributes. Such an instance-level transfer task is more challenging than the domain-level one that only considers the pre-defined lighting categories. To tackle this problem, we develop an instance-level conditional Generative Adversarial Networks (GAN). Specifically, face identifier is integrated into GAN learning, which enables an individual-specific low-level visual generation. Moreover, the illumination-inspired attention mechanism is conducted to allow GAN to well handle the local lighting effect. Our method requires neither lighting categorization, 3D information, nor strict face alignment, which are often employed by traditional methods. Experiments demonstrate that our method achieves significantly better results than previous methods.
Xiaohua Xie, Jian-Huang Lai
IEEE Trans. Image Process.2
2021 Learning Modal-Invariant Angular Metric by Cyclic Projection Network for VIS-NIR Person Re-Identification
abstract
Person re-identification across visible and near-infrared cameras (VIS-NIR Re-ID) has widespread applications. The challenge of this task lies in heterogeneous image matching. Existing methods attempt to learn discriminative features via complex feature extraction strategies. Nevertheless, the distributions of visible and near-infrared features are disparate caused by modal gap, which significantly affects feature metric and makes the performance of the existing models poor. To address this problem, we propose a novel approach from the perspective of metric learning. We conduct metric learning on a well-designed angular space. Geometrically, features are mapped from the original space to the hypersphere manifold, which eliminates the variations of feature norm and concentrates on the angle between the feature and the target category. Specifically, we propose a cyclic projection network (CPN) that transforms features into an angle-related space while identity information is preserved. Furthermore, we proposed three kinds of loss functions, AICAL, LAL and DAL, in angular space for angular metric learning. Multiple experiments on two existing public datasets, SYSU-MM01 and RegDB, show that performance of our method greatly exceeds the SOTA performance.
Jian-Huang Lai, Xiaohua Xie
IEEE Trans. Image Process.3
2020 Interactive Two-Stream Decoder for Accurate and Fast Saliency Detection
abstract
Recently, contour information largely improves the performance of saliency detection. However, the discussion on the correlation between saliency and contour remains scarce. In this paper, we first analyze such correlation and then propose an interactive two-stream decoder to explore multiple cues, including saliency, contour and their correlation. Specifically, our decoder consists of two branches, a saliency branch and a contour branch. Each branch is assigned to learn distinctive features for predicting the corresponding map. Meanwhile, the intermediate connections are forced to learn the correlation by interactively transmitting the features from each branch to the other one. In addition, we develop an adaptive contour loss to automatically discriminate hard examples during learning process. Extensive experiments on six benchmarks well demonstrate that our network achieves competitive performance with a fast speed around 50 FPS. Moreover, our VGG-based model only contains 17.08 million parameters, which is significantly smaller than other VGG-based approaches. Code has been made available at: https://github.com/moothes/ITSD-pytorch.
Huajun Zhou, Xiaohua Xie, Jian-Huang Lai, Lingxiao Yang
CVPR2
2020 Open-World Group Retrieval with Ambiguity Removal: A Benchmark
abstract
Group retrieval has attracted plenty of attention in artificial intelligence, traditional group retrieval researches assume that members in a group are unique and do not change under different cameras. However, the assumption may not be met for practical situations such as open-world and group-ambiguity scenarios. This paper tackles an important yet non-studied problem: re-identifying changing groups of people under the open-world and group-ambiguity scenarios in different camera fields. The open-world scenario considers that there are probably non-target people for the probe set appear in the searching gallery, while the group-ambiguity scenario means the group members may change. The open-world and group-ambiguity issue is very challenging for the existing methods because the changing of group members results in dramatic visual variations. Nevertheless, as far as we know, the existing literature lacks benchmarks which target on coping with this issue. In this paper, we propose a new group retrieval dataset named OWGA-Campus to consider these challenges. Moreover, we propose a person-to-group similarity matching based ambiguity removal (P2GSM-AR) method to solve these problems and realize the intention of group retrieval. Experimental results on OWGA-Campus dataset demonstrate the effectiveness and robustness of the proposed P2GSM-AR approach in improving the performance of the state-of-the-art feature extraction methods of person re-id towards the open-world and ambiguous group retrieval task.
Ling Mei 0001, Jian-Huang Lai, Zhan-Xiang Feng, Xiaohua Xie
ICPR4
2020 LG-VTON: Fashion Landmark Meets Image-Based Virtual Try-On
Zhenyu Xie, Jian-Huang Lai, Xiaohua Xie
PRCV (3)3
2020 From pedestrian to group retrieval via siamese network and correlation
Ling Mei 0001, Jian-Huang Lai, Zhan-Xiang Feng, Xiaohua Xie
Neurocomputing4
2020 Motion-Appearance Interactive Encoding for Object Segmentation in Unconstrained Videos
abstract
We present a two-stage framework of integrating motion and appearance cues for foreground object segmentation in unconstrained videos. Unlike conventional methods which encode motion and appearance patterns individually, our method puts particular emphasis on their mutual assistance. We propose an interactively constrained encoding (ICE) scheme to incorporate motion and appearance patterns into a graph that leads to a spatiotemporal energy optimization. Specifically, we construct a saliency network to infer the initial foreground maps and use optical flow to capture the initial motion information. After that, we perform ICE in the refinement stage for object segmentation. This scheme allows our method to consistently capture structural patterns about object perceptions throughout the whole framework. Our method can be operated on superpixels instead of raw pixels to reduce the number of graph nodes by two orders of magnitude. Moreover, we propose to tackle the object localization problem with inter-occlusion by weighted bipartite graph matching. The comprehensive experiments on two benchmark datasets (i.e., SegTrack-v2 and DAVIS2016) demonstrate the effectiveness of our approach compared with the state-of-the-art methods.
Chun-Chao Guo, Jian-Huang Lai, Xiaohua Xie
IEEE Trans. Circuits Syst. Video Technol.4
2020 Illumination-Invariance Optical Flow Estimation Using Weighted Regularization Transform
abstract
Many recent variational optical flow methods are not robust for illumination variance, and they only consider local image relation in terms of illumination. In this paper, we propose a new efficient illumination-invariance total variation optical flow method called the weighted regularization transform, which uses and optimizes the Weber's Law. Our method exploits unequal probability as the weight that has non-local information to estimate stable optical flow despite illumination changes. The proposed method uses a coarse-to-fine pyramid model to reduce the influence on the data term from illumination. Then, an energy optimization procedure is introduced to constrain the minimization of the data term with the non-local regularization. Experimentation with the proposed method has been performed on three optical flow datasets and a face liveness detection database, which have challenging illumination variations, and the results demonstrate that the proposed method is quite robust with respect to variations in illumination.
Ling Mei 0001, Jian-Huang Lai, Xiaohua Xie, Jun-Yong Zhu, Jun Chen 0013
IEEE Trans. Circuits Syst. Video Technol.3
2020 Learning Modality-Specific Representations for Visible-Infrared Person Re-Identification
abstract
Traditional person re-identification (re-id) methods perform poorly under changing illuminations. This situation can be addressed by using dual-cameras that capture visible images in a bright environment and infrared images in a dark environment. Yet, this scheme needs to solve the visible-infrared matching issue, which is largely under-studied. Matching pedestrians across heterogeneous modalities is extremely challenging because of different visual characteristics. In this paper, we propose a novel framework that employ modality-specific networks to tackle with the heterogeneous matching problem. The proposed framework utilizes the modality-related information and extracts modality-specific representations (MSR) by constructing an individual network for each modality. In addition, a cross-modality Euclidean constraint is introduced to narrow the gap between different networks. We also integrate the modality-shared layers into modality-specific networks to extract shareable information and use a modality-shared identity loss to facilitate the extraction of modality-invariant features. Then a modality-specific discriminant metric is learned for each domain to strengthen the discriminative power of MSR. Eventually, we use a view classifier to learn view information. The experiments demonstrate that the MSR effectively improves the performance of deep networks on VI-REID and remarkably outperforms the state-of-the-art methods.
Zhan-Xiang Feng, Jian-Huang Lai, Xiaohua Xie
IEEE Trans. Image Process.3
2019 Spatial-Temporal Person Re-Identification
abstract
Most of current person re-identification (ReID) methods neglect a spatial-temporal constraint. Given a query image, conventional methods compute the feature distances between the query image and all the gallery images and return a similarity ranked table. When the gallery database is very large in practice, these approaches fail to obtain a good performance due to appearance ambiguity across different camera views. In this paper, we propose a novel two-stream spatial-temporal person ReID (st-ReID) framework that mines both visual semantic information and spatial-temporal information. To this end, a joint similarity metric with Logistic Smoothing (LS) is introduced to integrate two kinds of heterogeneous information into a unified framework. To approximate a complex spatial-temporal probability distribution, we develop a fast Histogram-Parzen (HP) method. With the help of the spatial-temporal constraint, the st-ReID model eliminates lots of irrelevant images and thus narrows the gallery database. Without bells and whistles, our st-ReID method achieves rank-1 accuracy of 98.1% on Market-1501 and 94.4% on DukeMTMC-reID, improving from the baselines 91.2% and 83.8%, respectively, outperforming all previous state-of-theart methods by a large margin.
Guangcong Wang, Jian-Huang Lai, Peigen Huang, Xiaohua Xie
AAAI4
2019 Low Resolution Person Re-identification by an Adaptive Dual-Branch Network
Zhan-Xiang Feng, Jian-Huang Lai, Xiaohua Xie
ICIG (1)4
2019 Siamese Network for Pedestrian Group Retrieval: A Benchmark
Ling Mei 0001, Jian-Huang Lai, Xiaohua Xie
ICIG (1)3
2019 A Filtering-Based Framework for Optical Flow Estimation
abstract
We present a novel optical flow estimation framework that reinterprets most popular algorithms (e.g., the Horn-Schunck model, PDE-based models, TV-based models, and nonlocal-based models) from an iterative filtering perspective. Different regularizers in the related optical flow estimation algorithms can be achieved by adopting different filtering operations. The key merit of the proposed framework is that a variety of filtering models and corresponding optimization strategies, which are used in image processing, can be directly utilized in the regularization term, leading to a convenient way of designing appropriate optical flow estimation algorithms. Under the proposed filtering framework, we demonstrate easily designing optical flow estimation algorithms with respect to translation consistency, rotation consistency, and divergence consistency constraints by adopting different spatial filters, respectively. We also derive a novel optical flow estimation algorithm with 3D filtering-based model for the regularization term, which makes use of nonlocal self-similarity and sparse characteristic of optical flow field. Benefited from the advantages of patch-based nonlocal sparse regularizer, the derived algorithm can remove outliers while preserving sharp flow edges and important motion details. At present, the proposed algorithm achieves a high ranking on the challenging MPI-Sintel dataset and shows good performance on the Kitti flow 2015 and Middlebury datasets.
Jun Chen 0013, Jian-Huang Lai, Xiaohua Xie
IEEE Trans. Circuits Syst. Video Technol.4
2019 Efficient Segmentation-Based PatchMatch for Large Displacement Optical Flow Estimation
abstract
Efficient optical flow estimation with high accuracy is a challenging problem in computer vision. In this paper, we present a simple but efficient segmentation-based PatchMatch framework to address this issue. Specifically, it firstly generates sparse seeds without losing important motion information by over-segmentation, and then yields sparse matches by adopting a coarse-to-fine PatchMatch with sparse seeds. Such a scheme enhances the robustness of global regularization and yields better matching results compared with the existing NNF techniques while leading to a significant speed-up due to the sparsity of these seeds. Simultaneously, we introduce an extended nonlocal propagation and adaptive random search to address the basic limitation of the traditional coarse-to-fine framework in handing motion details that often vanish at coarser levels. Finally, we obtain dense matches at the finest level through an efficient sparse-to-dense matching according to the cues of over-segmentation. While performing an efficient approximation for over-segmentation, the proposed algorithm runs significantly fast and is robust to large displacements while preserving important motion details. It also achieves good performance on the challenging MPI-Sintel and Kitti flow 2015 datasets.
Jun Chen 0013, Jian-Huang Lai, Xiaohua Xie
IEEE Trans. Circuits Syst. Video Technol.4
2018 Learning Intrinsic Image Decomposition by Deep Neural Network with Perceptual Loss
abstract
Intrinsic Image Decomposition (IID) refers to recovering the albedo and shading from images, and it plays important roles in addressing computer vision tasks such as illumination-invariant object recognition and image recoloring. IID is an ill-posed problem and lacks of actual labelled samples for learning. This paper presents a deep neural network (DNN) based method to address this problem. To facilitate the training of DNN, we synthesize an intrinsic image dataset through rendering \pmb 3D models. To make the learnt model well generalize to realworld images with better visual results, we employ the perceptual loss in model learning. The perceptual loss is constructed upon the activations of a neural network pre-trained on real images, such as the VGG network. Such a loss function implicitly introduces knowledge from real-world images and has the ability of multi-level semantic understanding on the decomposed results. Experimental results show that our model trained on synthetic single-object dataset can produce well decomposition results not only on synthetic images but also on real-world scene-level images (containing multiple objects).
Guangyun Han, Xiaohua Xie, Wei-Shi Zheng 0001, Jian-Huang Lai
ICPR2
2018 Face Image Illumination Processing Based on Generative Adversarial Nets
abstract
It is a well-known fact that the variations in illumination could seriously affect the performance of 2D face analysis algorithms, such as face landmarking and face recognition. Unfortunately, the illumination condition is usually uncontrolled and unpredictable in most practical applications. Numerous methods have been developed to tackle this problem but the results is poor, especially for images with extreme lighting condition. Furthermore, most traditional illumination processing methods only demonstrate on grayscale images and require strict alignment of face images, resulting in limited applications in real world. In this paper, we proposed to reformulate the face image illumination processing problem as a style translation task with a Generative Adversarial Network (GAN). The key insight is to use the powerful mapping ability of GAN between two domains without knowing their true distributions. In this new sight, we developed a new multi-scale dual discriminate nets and employed multi-scale adversarial learning for visually realistic illumination processing. Advocating the use of the insights from traditional method, we also use reconstruction learning and add two new loss items of image quality assessment to enforce the preservation of all other illumination excluding details on the generated image. Experiments on CMU Multi-PIE and FRGC datasets show that our method can obtain promising illumination normalization results and preserve a superior visual quality.
Xiaohua Xie, Chong Yin, Jian-Huang Lai
ICPR2
2018 Towards Automatic Detection of Monkey Faces
abstract
An automated monkey face detection system confers distinct advantages in the protection of wild monkeys, sociological studies, monkey feeding and management and so on. The monkey face and human face have similar structures, but still hold some very important differences in appearance. Therefore, whether the mainstream human face detection algorithms can be adapted to the detection of monkey face is still unknown. To investigate this problem, we collected a database of monkey face (with more than 20,000 macaque faces) and conducted several experiments in our database. Experimental results reveal some interesting results. Firstly, the classical Viola-Jones Adaboost algorithm on monkey faces does not work as well as that on human faces. A in-depth study for this result will be given by taking insight into the selected features by Adaboost. In particular, lips and eyebrows are very important to human face recognition. However, the lack of these prominent features in the monkey's face causes the Viola-Jones algorithm to choose more local Haar-like features, resulting in a higher false positive rate. Secondly, the Faster R-CNN works effectively for monkey face detection but requires a large number of training samples. A pre-training with human faces helps to tackle the problem of shortage of monkey faces for training. Above conclusion indicate that an automatic monkey face detector can be learnt from a human face detector, yet a model with complex features should be employed.
Manning Zhang, Susu Guo, Xiaohua Xie
ICPR3
2018 Conditional Face Synthesis for Data Augmentation
Xiaohua Xie, Jian-Huang Lai, Zhan-Xiang Feng
PRCV (3)2
2018 Feature Visualization Based Stacked Convolutional Neural Network for Human Body Detection in a Depth Image
Xiao Liu 0023, Ling Mei 0001, Dakun Yang, Jian-Huang Lai, Xiaohua Xie
PRCV (2)5
2018 Asymmetric Two-Stream Networks for RGB-Disparity Based Object Detection
Ruizhi Lu, Jian-Huang Lai, Xiaohua Xie
PRCV (4)3
2018 Face Image Illumination Processing Based on GAN with Dual Triplet Loss
Xiaohua Xie, Jian-Huang Lai, Jun-Yong Zhu
PRCV (3)2
2018 Convolutional LSTM Based Video Object Detection
Xiaohua Xie, Jian-Huang Lai
PRCV (2)2
2018 Exploring Multi-scale Deep Feature Fusion for Object Detection
Jian-Huang Lai, Xiaohua Xie, Jun-Yong Zhu
PRCV (4)3
2018 Image super-resolution via a densely connected recursive network
Zhan-Xiang Feng, Jian-Huang Lai, Xiaohua Xie, Jun-Yong Zhu
Neurocomputing3
2018 Learning discriminative visual elements using part-based convolutional neural network
Lingxiao Yang, Xiaohua Xie, Jian-Huang Lai
Neurocomputing2
2018 Learning an Intrinsic Image Decomposer Using Synthesized RGB-D Dataset
abstract
Intrinsic image decomposition refers to recover the albedo and shading from images, which is an ill-posed problem in signal processing. As realistic labeled data are severely lacking, it is difficult to apply learning methods in this issue. In this letter, we propose using a synthesized dataset to facilitate the solving of this problem. A physically based renderer is used to generate color images and their underlying ground-truth albedo and shading from three-dimensional models. Additionally, we render a Kinect-like noisy depth map for each instance. We utilize this synthetic dataset to train a deep neural network for intrinsic image decomposition and further fine-tune it for real-world images. Our model supports both RGB and RGB-D as input, and it employs both high-level and low-level features to avoid blurry outputs. Experimental results verify the effectiveness of our model on realistic images.
Guangyun Han, Xiaohua Xie, Jian-Huang Lai, Wei-Shi Zheng 0001
IEEE Signal Process. Lett.2
2018 Fast Optical Flow Estimation Based on the Split Bregman Method
abstract
Fast and accurate optical flow estimation is a challenging problem in computer vision. In this paper, we present a novel model to solve the optical flow problem by combining the strengths of the Split Bregman method with the advantages of an efficient variational framework. It allows us to employ different regularization tensors for the Split Bregman regularizer to preserve motion discontinuities. Simultaneously, a novel nonlocal Split Bregman method with adaptive support weights is also developed for the smoothness regularization to preserve motion details, repel outliers, enhance contrast, and reduce motion blurring and staircase effect. The proposed algorithm shows significant improvements in preserving sharp flow edges and important motion details. Most of all, it requires only a few iterations to achieve fast convergence and runs faster than the classic TV-L1 flow method. Our method significantly outperforms the current state-of-the-art methods on the challenging MPI-Sintel dataset and shows good performance on the Kitti flow 2015 and Middlebury datasets.
Jun Chen 0013, Jian-Huang Lai, Xiaohua Xie
IEEE Trans. Circuits Syst. Video Technol.4
2018 P2SNet: Can an Image Match a Video for Person Re-Identification in an End-to-End Way?
abstract
We address a new person re-identification problem, in which our goal is to directly match image-based scenarios with video-based ones. This differs significantly from the conventional person re-identification problem, which aims to match two image-based scenarios (and it is assumed that the available video frames have been manually selected to form the image-based scenarios). To solve this more challenging and realistic problem without the implicit assumption of manual selection, we propose an end-to-end matching framework called a point-to-set network (P2SNet), which consists of: 1) a k-nearest neighbor triplet module, which functions as a “denoiser” by letting the network sequentially focus on the available frames, while ignoring the other useless frames in a video and 2) a novel deep neural network that uses videos and images as input to jointly learn the feature representations and a point-to-set distance metric in a unified way. Our P2SNet is evaluated on three new image-to-video person re-identification data sets, i-LIDS-VID-P2S, PRID2011-P2S, and MARS-P2S, which are modified from i-LIDS-VID, PRID 2011, and MARS, respectively. The experimental results demonstrate the superior performance of our model over the other state-of-the-art methods.
Guangcong Wang, Jian-Huang Lai, Xiaohua Xie
IEEE Trans. Circuits Syst. Video Technol.3
2018 Learning View-Specific Deep Networks for Person Re-Identification
abstract
In recent years, a growing body of research has focused on the problem of person re-identification (re-id). The re-id techniques attempt to match the images of pedestrians from disjoint non-overlapping camera views. A major challenge of the re-id is the serious intra-class variations caused by changing viewpoints. To overcome this challenge, we propose a deep neural network-based framework which utilizes the view information in the feature extraction stage. The proposed framework learns a view-specific network for each camera view with a cross-view Euclidean constraint (CV-EC) and a cross-view center loss. We utilize the CV-EC to decrease the margin of the features between diverse views and extend the center loss metric to a view-specific version to better adapt the re-id problem. Moreover, we propose an iterative algorithm to optimize the parameters of the view-specific networks from coarse to fine. The experiments demonstrate that our approach significantly improves the performance of the existing deep networks and outperforms the state-of-the-art methods on the VIPeR, CUHK01, CUHK03, SYSU-mReId, and Market-1501 benchmarks.
Zhan-Xiang Feng, Jian-Huang Lai, Xiaohua Xie
IEEE Trans. Image Process.3
2017 Deep Growing Learning
abstract
Semi-supervised learning (SSL) is an import paradigm to make full use of a large amount of unlabeled data in machine learning. A bottleneck of SSL is the overfitting problem when training over the limited labeled data, especially on a complex model like a deep neural network. To get around this bottleneck, we propose a bio-inspired SSL framework on deep neural network, namely Deep Growing Learning (DGL). Specifically, we formulate the SSL as an EM-like process, where the deep network alternately iterates between automatically growing convolutional layers and selecting reliable pseudo-labeled data for training. The DGL guarantees that a shallow neural network is trained with labeled data, while a deeper neural network is trained with growing amount of reliable pseudo-labeled data, so as to alleviate the overfitting problem. Experiments on different visual recognition tasks have verified the effectiveness of DGL.
Guangcong Wang, Xiaohua Xie, Jian-Huang Lai, Jiaxuan Zhuo
ICCV2
2017 Face recognition by landmark pooling-based CNN with concentrate loss
abstract
Face recognition has been a hot research topic in recent years, convolutional neural network (CNN) based methods have achieved state of the art results and significantly improve the performance. Along with the CNN framework, we propose a novel loss function called concentrate loss which focuses on the class centers in the mini-batch. The concentrate loss aims to push the samples towards corresponding class centers and simultaneously enlarge the gap between different class centers. Additionally, we ultilize facial landmark pooling technique to take full advantage of facial structure information. Experiment results on Labeled Faces in the Wild (LFW), YouTube Faces (YTF), and the BluFR benchmark demonstrate the efficiency of our proposal.
Xiaohua Xie, Zhan-Xiang Feng, Jian-Huang Lai
ICIP2
2017 Part-based convolutional neural network for visual recognition
abstract
Mid-level element based representations have been proven to be very effective for visual recognition. We present a method to discover discriminative elements based on deep Convolutional Neural Networks (CNNs), namely Part-based CNN (P-CNN), which acts as the role of encoding module in part-based representation. The P-CNN can be attached at arbitrary layer of a pre-trained CNN and be trained using image-level labels. The training of P-CNN essentially corresponds to the optimization and selection of discriminative mid-level visual elements. For an input image, the output of P-CNN is naturally the part-based coding and can be directly used for image recognition. By applying P-CNN to multiple layers of a pretrained CNN, more diverse visual elements can be obtained for visual recognitions. Experiments are conducted on two recognition tasks and their results demonstrate the effectiveness of the proposed method.
Lingxiao Yang, Xiaohua Xie, Peihua Li, David Zhang 0001, Lei Zhang 0006
ICIP2
2017 Sparse transfer for facial shape-from-shading
Jianfang Hu, Wei-Shi Zheng 0001, Xiaohua Xie, Jian-Huang Lai
Pattern Recognit.3
2016 Face hallucination by deep traversal network
abstract
In this paper, we propose a novel patch-based face hallucination method that consists of two patch-based sparse autoencoder (SAE) networks and a deep fully connected network (namely traversal network). The SAE networks are used to capture the intrinsic features of low-resolution (LR) images and high-resolution (HR) images in the hidden layers, while the traversal network is used to map features from the LR hidden layer to the HR hidden layer. In the training stage, these three networks are jointly optimized. Compared with previous network-based methods that learn an end-to-end mapping from LR images to HR images, our method learns the mapping between hidden layers, which can better alleviate the over-fitting problem. Experimental results demonstrate that our method is efficient and robust for hallucinating face images from both lab environment and the wild. The proposal achieves state-of-the-art performance when conducting face hallucination in CAS-PEAL-R1 database, CMU-PIE database and Casia database.
Zhan-Xiang Feng, Jian-Huang Lai, Xiaohua Xie, Dakun Yang, Ling Mei 0001
ICPR3
2016 HEp-2 specimen classification via deep CNNs and pattern histogram
abstract
Automatic classification of Human Epithelial Type-2 (HEp-2) specimen patterns is an important yet challenging problem in medical image analysis. Most prior works have primarily focused on cells images classification problem which is one of the early essential steps in the system pipeline, while less attention has been paid to the classification of whole-specimen ones. In this work, a specimen pattern recognition system combining convolutional neural networks (CNNs) and pattern histogram was proposed. The pattern histograms were obtained based on the prediction of each single cell inside the specimens. Two strategies were designed to predicted the pattern of a whole specimen: 1) the most dominant cell pattern in pattern histogram was represented as the specimen pattern, 2) the pattern histograms were employed as bags of patterns and then were trained and predicted separately by a SVM classifier. Experimental results show that the proposed system is effective and achieves high classification accuracy on public benchmark datasets. We further evaluate the robustness of the proposed framework by testing trained CNNs on another different dataset, demonstrating that the system is robust to inter-lab data.
Hongwei Li 0004, Wei-Shi Zheng 0001, Xiaohua Xie, Jianguo Zhang 0001
ICPR4
2016 Facial skin beautification via sparse representation over learned layer dictionary
abstract
In this paper, we propose a facial skin beautification framework to remove facial spots based on layer dictionary learning and sparse representation. More precisely, we first decompose the face image into three layers: lighting layer, detail layer and color layer. The corresponding detail layer dictionary are learned by using 60 thousands beauty images collected from the Internet. Thereafter, the detail layer of the image is reconstructed by using sparse representation. Moreover, a binary mask obtained from the learned layer is used to transform detail information from original detail layer to the learned one. The experiment results demonstrate that the proposed method is more effective in eliminating moles, flaws and wrinkles in face image compared with representative commercial systems like PicTreat, Portrait+, Portraitrue and MeituPic.
Xiaobin Chang, Xiaohua Xie, Jianfang Hu, Wei-Shi Zheng 0001
IJCNN3
2016 Learning object-specific DAGs for multi-label material recognition
Xiaohua Xie, Lingxiao Yang, Wei-Shi Zheng 0001
Comput. Vis. Image Underst.1
2016 Exploiting object semantic cues for Multi-label Material Recognition
Lingxiao Yang, Xiaohua Xie
Neurocomputing2
2015 Recovering intrinsic images from image sequences using total variation models
abstract
Recovering intrinsic images from natural photos is one of the foundational problems in computer vision. This mission always falls into an ill-posed problem. In order to attain reasonable estimations, one strategy is to use multiple images of the scene under various lightings so as to narrow the solution space, whereas another is to utilize priori knowledge as constraints. In this paper, we present an approach to deriving intrinsic images (including illumination images and reflectance images) that employs both strategies. Specifically, the Total Variation (TV) constraint is imposed because of its excellent edge preservation ability and simple parameter settings. To solve this objective function efficiently, we propose using the Alternating Direction Method of Multipliers (AD-MM) to build an iterative numerical scheme. Experimental results illustrate the effectiveness of the proposed model and the numerical scheme.
Xiaohua Xie, Wenyong Gong, Minglun Gong, Tieru Wu
ICIP1
2015 Max-margin analysis based patch sampling for discovery of mid-level parts
abstract
Discovering representative, discriminative mid-level parts is crucial for visual recognition models such as Bag-Of-Parts. We present a weakly-supervised approach to learn class-specific mid-level parts from a database. In our approach, only the image-level labels but no additional human annotations are used. As a start, we employ a SVM-like model to sample discriminant patches from each image. The employed SVM-like model corresponds to a max-margin analysis between a specific image patch and other patches from the whole training set, which can be easily solved in a closed form. For each class, the sampled patches are then clustered in an agglomerative manner to generate the final semantic parts, in the meantime the less-representative patches are discarded. The proposed approach is effective since it sequentially discards the non-discriminative and non-representative patches. The approach is also efficient since the clustering operation only needs to handle a small number of discriminant patches. The state-of-the-art results are observed in scene classification benchmarks when using the learned parts as a visual codebook.
Lingxiao Yang, Xiaohua Xie
ICIP2
2014 Illumination preprocessing for face images based on empirical mode decomposition
Xiaohua Xie
Signal Process.1
2013 Sketch-to-Design: Context-Based Part Assembly
abstract
Abstract Designing 3D objects from scratch is difficult, especially when the user intent is fuzzy and lacks a clear target form. We facilitate design by providing reference and inspiration from existing model contexts. We rethink model design as navigating through different possible combinations of part assemblies based on a large collection of pre‐segmented 3D models. We propose an interactive sketch‐to‐design system, where the user sketches prominent features of parts to combine. The sketched strokes are analysed individually, and more importantly, in context with the other parts to generate relevant shape suggestions via adesign galleryinterface. As a modelling session progresses and more parts get selected, contextual cues become increasingly dominant, and the model quickly converges to a final form. As a key enabler, we use pre‐learned part‐based contextual information to allow the user to quickly explore different combinations of parts. Our experiments demonstrate the effectiveness of our approach for efficiently designing new variations from existing shape collections.
Xiaohua Xie, Kai Xu 0004, Niloy J. Mitra, Daniel Cohen-Or, Wenyong Gong, Baoquan Chen
Comput. Graph. Forum1
2013 Face hallucination based on morphological component analysis
Xiaohua Xie, Jian-Huang Lai
Signal Process.2
2012 Radical Extraction Using Affine Sparse Matrix Factorization for Printed Chinese Characters Recognition
abstract
Each Chinese character is comprised of radicals, where a single character (compound character) contains one (or more than one) radicals. For human cognitive perspective, a Chinese character can be recognized by identifying its radicals and their spatial relationship. This human cognitive law may be followed in computer recognition. However, extracting Chinese character radicals automatically by computer is still an unsolved problem. In this paper, we propose using an improved sparse matrix factorization which integrates affine transformation, namely affine sparse matrix factorization (ASMF), for automatically extracting radicals from Chinese characters. Here the affine transformation is vitally important because it can address the poor-alignment problem of characters that may be caused by internal diversity of radicals and image segmentation. Consequently we develop a radical-based Chinese character recognition model. Because the number of radicals is much less than the number of Chinese characters, the radical-based recognition performs a far smaller category classification than the whole character-based recognition, resulting in a more robust recognition system. The experiments on standard Chinese character datasets show that the proposed method gets higher recognition rates than related Chinese character recognition methods.
Jun Tan 0001, Xiaohua Xie, Wei-Shi Zheng 0001, Jian-Huang Lai
Int. J. Pattern Recognit. Artif. Intell.2
2011 Non-ideal class non-point light source quotient image for face relighting
Xiaohua Xie, Jian-Huang Lai, Ching Y. Suen, Wei-Shi Zheng 0001
Signal Process.1
2011 Normalization of Face Illumination Based on Large-and Small-Scale Features
abstract
A face image can be represented by a combination of large-and small-scale features. It is well-known that the variations of illumination mainly affect the large-scale features (low-frequency components), and not so much the small-scale features. Therefore, in relevant existing methods only the small-scale features are extracted as illumination-invariant features for face recognition, while the large-scale intrinsic features are always ignored. In this paper, we argue that both large-and small-scale features of a face image are important for face restoration and recognition. Moreover, we suggest that illumination normalization should be performed mainly on the large-scale features of a face image rather than on the original face image. A novel method of normalizing both the Small-and Large-scale (S&L) features of a face image is proposed. In this method, a single face image is first decomposed into large-and small-scale features. After that, illumination normalization is mainly performed on the large-scale features, and only a minor correction is made on the small-scale features. Finally, a normalized face image is generated by combining the processed large-and small-scale features. In addition, an optional visual compensation step is suggested for improving the visual quality of the normalized image. Experiments on CMU-PIE, Extended Yale B, and FRGC 2.0 face databases show that by using the proposed method significantly better recognition performance and visual results can be obtained as compared to related state-of-the-art methods.
Xiaohua Xie, Wei-Shi Zheng 0001, Jian-Huang Lai, Pong C. Yuen, Ching Y. Suen
IEEE Trans. Image Process.1
2010 Face Hallucination under an Image Decomposition Perspective
abstract
In this paper we propose to convert the task of face hallucination into an image decomposition problem, and then use the morphological component analysis (MCA) for hallucinating a single face image, based on a novel three-step framework. Firstly, a low-resolution input image is up-sampled by interpolation. Then, the MCA is employed to decompose the interpolated image into a high-resolution image and an unsharp masking, as MCA can properly decompose a signal into special parts according to typical dictionaries. Finally, a residue compensation, which is based on the neighbor reconstruction of patches, is performed to enhance the facial details. The proposed method can effectively exploit the facial properties for face hallucination under the image decomposition perspective. Experimental results demonstrate the effectiveness of our method, in terms of the visual quality of the hallucinated face images.
Jian-Huang Lai, Xiaohua Xie, Wanquan Liu
ICPR3
2010 Restoration of a Frontal Illuminated Face Image Based on KPCA
abstract
In this paper, we propose a novel illumination-normalization method. By using the combination of the Kernel Principal Component Analysis (KPCA) and Pre-image technology, this method can restore the frontal-illuminated face image from a single non-frontal-illuminated face image. In this method, a frontal-illumination subspace is first learned by KPCA. For each input face image, we project its large-scale features, which are affected by illumination variations, onto this subspace to normalize the illumination. Then the frontal-illuminated face image is reconstructed by combining the small- and the normalized large- scale features. Unlike most existing techniques, the proposed method does not require any shape modeling or lighting estimation. As a holistic reconstruction, KPCA+Pre-image technology incurs less local distortion. Compared to directly applying KPCA+Pre-image technology on the original image, our proposed method can be better at processing an image of a face that is outside the training set. Experiments on CMU-PIE and Extended Yale B face databases show that the proposed method outperforms state-of-the-art algorithms.
Xiaohua Xie, Wei-Shi Zheng 0001, Jian-Huang Lai, Ching Y. Suen
ICPR1
2010 Extraction of illumination invariant facial features from a single image using nonsubsampled contourlet transform
Xiaohua Xie, Jian-Huang Lai, Wei-Shi Zheng 0001
Pattern Recognit.1
2008 Face illumination normalization on large and small scale features
abstract
It is well known that the effect of illumination is mainly on the large-scale features (low-frequency components) of a face image. In solving the illumination problem for face recognition, most (if not all) existing methods either only use extracted small-scale features while discard large-scale features, or perform normalization on the whole image. In the latter case, small-scale features may be distorted when the large-scale features are modified. In this paper, we argue that large-scale features of face image are important and contain useful information for face recognition as well as visual quality of normalized image. Moreover, this paper suggests that illumination normalization should mainly perform on large-scale features of face image rather than the whole face image. Along this line, a novel framework for face illumination normalization is proposed. In this framework, a single face image is first decomposed into large- and small- scale feature images using logarithmic total variation (LTV) model. After that, illumination normalization is performed on large-scale feature image while small-scale feature image is smoothed. Finally, a normalized face image is generated by combination of the normalized large-scale feature image and smoothed small-scale feature image. CMU PIE and (Extended) YaleB face databases with different illumination variations are used for evaluation and the experimental results show that the proposed method outperforms existing methods.
Xiaohua Xie, Wei-Shi Zheng 0001, Jian-Huang Lai, Pong C. Yuen
CVPR1