Shijuan Huang

dblp:335/0950 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2026
0009-0000-2177-5110ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dual-stream Relation-modeling Disentanglement for Cloth-Changing Person Re-Identification
abstract
Cloth-changing person re-identification (CC-ReID) aims to identify individuals across non-overlapping cameras despite clothing variations. Existing methods are often constrained by two primary limitations: approaches using auxiliary modalities typically rely on a single specific cue, limiting their robustness, while feature disentanglement methods struggle with discrete labels that create inconsistencies between ground truth labels and modality semantic similarity. To overcome these limitations, we propose DRDnet, a unified framework that synergistically integrates dual auxiliary cues and advanced relation modeling. Specifically, our Dual-Stream Disentanglement (DSD) module leverages textual descriptions and parsing images to decouple clothing factors through high-level semantic supervision and pixel-level operations, yielding robust clothing-agnostic features. Simultaneously, our Modal Relation Modeling (MRM) module constructs feature memory banks and employs adaptive soft label smoothing, effectively enhancing image-text semantic alignment and reinforcing identity consistency across clothing changes. We evaluate DRDnet on several CC-ReID benchmarks to demonstrate its effectiveness and provide state-of-the-art performance across all benchmarks.
Shijuan Huang, Zongyi Li, Zhao Lv
AAAI1
2025 Exploring the Potential of Large Vision-Language Models for Unsupervised Text-Based Person Retrieval
abstract
The aim of text-based person retrieval is to identify pedestrians using natural language descriptions within a large-scale image gallery. Traditional methods rely heavily on manually annotated image-text pairs, which are resource-intensive to obtain. With the emergence of Large Vision-Language Models (LVLMs), the advanced capabilities of contemporary models in image understanding have led to the generation of highly accurate captions. Therefore, this paper explores the potential of employing Large Vision-Language Models for unsupervised text-based pedestrian image retrieval and proposes a Multi-grained Uncertainty Modeling and Alignment framework (MUMA). Initially, multiple Large Vision-Language Models are employed to generate diverse and hierarchically structured pedestrian descriptions across different styles and granularities. However, the generated captions inevitably introduce noise. To address this issue, an uncertainty-guided sample filtration module is proposed to estimate and filter out unreliable image-text pairs. Additionally, to simulate the diversity of styles and granularities in captions, a multi-grained uncertainty modeling approach is applied to model the distributions of captions, with each caption represented as a multivariate Gaussian distribution. Finally, a multi-level consistency distillation loss is employed to integrate and align the multi-grained captions, aiming to transfer knowledge across different granularities. Experimental evaluations conducted on three widely-used datasets demonstrate the significant advancements achieved by our approach.
Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Shijuan Huang, Linnan Tu, Fei Shen 0004
AAAI5
2025 Detecting Adversarial Data Using Perturbation Forgery
abstract
As a defense strategy against adversarial attacks, adversarial detection aims to identify and filter out adversarial data from the data flow based on discrepancies in distribution and noise patterns between natural and adversarial data. Although previous detection methods achieve high performance in detecting gradient-based adversarial attacks, new attacks based on generative models with imbalanced and anisotropic noise patterns evade detection. Even worse, the significant inference time overhead and limited performance against unseen attacks make existing techniques impractical for real-world use. In this paper, we explore the proximity relationship among adversarial noise distributions and demonstrate the existence of an open covering for these distributions. By training on the open covering of adversarial noise distributions, a detector with strong generalization performance against various types of unseen attacks can be developed. Based on this insight, we heuristically propose Perturbation Forgery, which includes noise distribution perturbation, sparse mask generation, and pseudo-adversarial data production, to train an adversarial detector capable of detecting any unseen gradient-based, generative-based, and physical adversarial attacks. Comprehensive experiments conducted on multiple general and facial datasets, with a wide spectrum of attacks, validate the strong generalization of our method.1
Qian Wang 0001, Shijuan Huang, Ruoxi Jia 0001, Ning Yu 0006
CVPR5
2025 Multi-Branch Clothes-Agnostic Feature Learning for Cloth-Changing Person Re-Identification
abstract
Person Re-Identification (Re-ID) is crucial for video surveillance and multi-camera tracking, yet traditional methods struggle with clothing changes that undermine their reliability. This paper introduces a novel multi-branch clothes-agnostic feature learning framework to address cloth-changing person re-identification (CC-ReID), which comprises two key modules: Multi-grained Clothes Caption Generation (MCG), and Multi-Branch Clothes-Agnostic Feature Extraction (MAE). MCG leverages Large Vision-Language Models to generate diverse coarse-to-fine clothing descriptions, reducing the impact of clothing on feature extraction. MAE employs a dual-branch architecture combining Semantic-Guided Feature Extraction (SGE) and Parsing Image Feature Extraction (PIE) to focus on identity-related features while minimizing dependence on clothing characteristics. Extensive experiments on benchmark datasets demonstrate that our method achieves state-of-the-art performance for CC-ReID tasks, showcasing our method’s effectiveness in real-world applications.
Shijuan Huang
ECAI1
2025 Seeing Through Changes: Clothing Information Suppression and Momentum Enhancement for Cloth-Changing Person Re-Identification
abstract
Cloth-Changing Person Re-Identification (CC-ReID) aims to identify individuals across different camera views despite changes in their clothing. Recent approaches focus on leveraging clothing-invariant features such as body shape and gait. However, these methods often fail to suppress clothing-related information while maintaining robust features effectively. This issue arises because clothing variations can overshadow more stable identity-related features during feature extraction. To alleviate these problems, we propose the Clothing Information Suppression and Momentum Enhancement (CSME) framework. This innovative solution incorporates two key components: the Low-level Clothing Information Suppression (LCS) module, which effectively suppresses the impact of clothing changes by filtering out clothing-related features in the early stages of feature extraction; and the Momentum-Enhanced Feature Representation (MER) module that enhances the model’s ability to capture more stable features through momentum-updated mechanisms. Experimental results on several widely used datasets demonstrate that the CSME method significantly improves retrieval performance, achieving state-of-the-art results in CC-ReID tasks.
Shijuan Huang
IJCNN1
2025 Autoregressive Motion Generation with Gaussian Mixture-Guided Latent Sampling
abstract
Existing efforts in motion synthesis typically utilize either generative transformers with discrete representations or diffusion models with continuous representations. However, the discretization process in generative transformers can introduce motion errors, while the sampling process in diffusion models tends to be slow. In this paper, we propose a novel text-to-motion synthesis method GMMotion that combines a continuous motion representation with an autoregressive model, using the Gaussian mixture model (GMM) to represent the conditional probability distribution. Unlike autoregressive approaches relying on residual vector quantization, our model employs continuous motion representations derived from the VAE's latent space. This choice streamlines both the training and the inference processes. Specifically, we utilize a causal transformer to learn the distributions of continuous motion representations, which are modeled with a learnable Gaussian mixture model. Extensive experiments demonstrate that our model surpasses existing state-of-the-art models in the motion synthesis task.
Linnan Tu, Lingwei Meng, Zongyi Li, Shijuan Huang
NeurIPS5
2025 Cross-Modality Relation and Uncertainty Exploration for Text-Based Person Search
abstract
Text-based person search aims to retrieve specific individuals from an extensive image gallery using textual queries. Recent approaches have delved into aligning global and part features in both text and image modalities, yielding substantial improvements. However, these methods often overlook intra-modality instance relations and uncertainties inherent in text-based person search. In response to these challenges, we propose the Cross-Modality Relation and Uncertainty Exploration (CRUE) method to model the relations and uncertainties in the matching procedure. To alleviate the strict alignment issues arising from hard labels in the original contrastive loss, an Intra-Modality Relation Exploration (IRE) module is introduced. This module is specifically designed to smooth hard-matching relations by modeling intra-modality similarity. Additionally, to address uncertain matching problems stemming from many-to-many relations, we propose a novel Uncertainty-Guided Modeling (UGM) module. This module is specifically designed to handle weak and noise-matched image–text pairs by modeling features as distributions, thereby alleviating instability and noise. Both the IRE and UGM modules effectively consider genuine intra-modality similarities and reduce the negative impact of uncertainties. Experimental results demonstrate significant improvements across three widely used person search datasets, thereby validating the efficacy of the CRUE method in enhancing text-based person search. Our code will be available on GitHub at https://github.com/ShijuanHuang/CRUE .
Shijuan Huang, Zongyi Li
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Cross-modal Generation and Alignment via Attribute-guided Prompt for Unsupervised Text-based Person Retrieval
Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Runsheng Wang, Shijuan Huang
IJCAI7
2024 Knowledge Consistency Distillation for Weakly Supervised One Step Person Search
abstract
Weakly supervised person search targets to detect and identify a person with only bounding box annotations. Recent approaches have focused on learning person relations in a single model, ignoring the conflicts between the detection and Re-ID heads, along with the influence of background elements, which may lead to noisy pseudo labels and inaccurate Re-ID features. To address this challenge, we introduce a novel framework named Knowledge Consistency Distillation (KCD) for weakly supervised person search, which explores the capabilities of an advanced unsupervised person re-identification (Re-ID) model to mitigate the conflicts and background influences. We propose hierarchical consistency alignments, including feature-level, cluster-level, and instance-level consistency alignment, to synchronize the knowledge from the state-of-the-art unsupervised Re-ID model. Specifically, the feature-level consistency aligns the feature through both context and relation alignment. The cluster-level consistency aligns the teacher cluster information by reusing its OIM module. To tackle the inconsistency problem between student instances and teacher cluster centroids, we incorporate pseudo-label refinement to assist the student model in comprehending the teacher’s knowledge at cluster-level while mitigating the negative effects of noisy labels. Finally, an instance-level consistency loss weighted by the similarity between the instance and its corresponding cluster is proposed to align the positive instance correlations. Our approach aims to train a one-step weakly supervised model for person search by exploiting the characteristics of unsupervised person Re-ID. Extensive experiments illustrate that our method achieves state-of-the-art performance on two widely-used person search datasets, CUHK-SYSU and PRW. Our code will be available on GitHub athttps://github.com/zongyi1999/KCD.
Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Runsheng Wang, Chengxin Zhao, Qian Wang 0001, Shijuan Huang
IEEE Trans. Circuits Syst. Video Technol.8
2022 EXPERT: transfer learning-enabled context-aware microbial community classification
abstract
Microbial community classification enables identification of putative type and source of the microbial community, thus facilitating a better understanding of how the taxonomic and functional structure were developed and maintained. However, previous classification models required a trade-off between speed and accuracy, and faced difficulties to be customized for a variety of contexts, especially less studied contexts. Here, we introduced EXPERT based on transfer learning that enabled the classification model to be adaptable in multiple contexts, with both high efficiency and accuracy. More importantly, we demonstrated that transfer learning can facilitate microbial community classification in diverse contexts, such as classification of microbial communities for multiple diseases with limited number of samples, as well as prediction of the changes in gut microbiome across successive stages of colorectal cancer. Broadly, EXPERT enables accurate and context-aware customized microbial community classification, and potentiates novel microbial knowledge discovery.
Hui Chong, Yuguo Zha, Qingyang Yu, Mingyue Cheng 0003, Guangzhou Xiong, Xinhe Huang, Shijuan Huang, Chuqing Sun, Sicheng Wu, Wei-Hua Chen, Luís Pedro Coelho, Kang Ning 0001
Briefings Bioinform.8