VLDB 2026 Research / reviewers in the wild / expert
Si Chen 0002
dblp:93/5439-2
· DBLP profile ↗
55ranked-venue papers
16as first author
36since 2021 · last 2026
0000-0002-5631-7942ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 9 first-author · 23 since 2021Artificial intelligence and machine learning · 27 · 7 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fair Facial Attribute Recognition via Group-Decoupled Vision Transformer with Mask-Guided Correlation SuppressionabstractFacial Attribute Recognition (FAR) holds significant potential for wide-ranging applications. However, traditionally trained FAR models exhibit unfairness, largely due to data bias—where certain sensitive attributes correlate statistically with target attributes. To address this, we propose a group-attention mechanism: first, each image is categorized into subgroups (e.g., Male/Female&short hair, Male/Female&long hair). Within the attention mechanism, distinct Query parameters are used for each group, with shared Key and Value parameters. As group-specific Query parameters are trained on subgrouped data, the noted bias is effectively mitigated. Consequently, integrating this Group-Attention into Vision Transformer (ViT) yields our novel Group-Decoupled ViT (GD-ViT) model. Moreover, to further attenuate the statistical correlation between sensitive and target attributes, we propose a Mask-Guided Correlation Suppression learning strategy. Specifically, in Stage 1, it first leverages a min-max dual-loss optimization strategy to train GD-ViT in capturing key regions related to sensitive attributes yet irrelevant to target attributes. Then, in Stage 2, it trains another GD-ViT by masking sensitive regions identified in Stage 1, fusing the masked output (as intermediate input) with the model’s intermediate outputs. This weakens regions associated with sensitive attributes while enhancing others, suppressing the learning of key features related to sensitive attributes. Consequently, it encourages the model to focus more on intrinsic target attribute regions and balances the learning process between the sensitive attribute and the target attribute. Extensive experiments demonstrate that our method achieves superior performance across three benchmark datasets for fair facial attribute recognition. Huichang Huang, Kunchi Li, Si Chen 0002, Dahan Wang |
AAAI | 3 |
| 2026 | HDFNet:Hybrid-domain fusion network for medical image restoration
Liqun Lin, Shunzhou Wang, Si Chen 0002, Chao Zeng 0005, Nanfeng Jiang, Dahan Wang |
Expert Syst. Appl. | 4 |
| 2026 | MoKA-HP: Motion-aware KAdaptation with historical prompts for efficient and robust RGB-T tracking
Zhixi Wu, Si Chen 0002, Dahan Wang, Shunzhi Zhu |
Neurocomputing | 2 |
| 2026 | Learning relationship-guided vision-language transformer for facial attribute recognition
Si Chen 0002, Mingxuan Lei, Dahan Wang, Xu-Yao Zhang, Yan Yan 0001, Shunzhi Zhu |
Pattern Recognit. | 1 |
| 2026 | ProMoT: Progressive Prompting of Modality and Temporal Dynamics for RGB-T TrackingabstractRGB-T tracking benefits from the complementary nature of RGB and TIR modalities, yet their relative reliability for target localization often shifts over time. Most existing trackers fail to adapt to such modality and temporal dynamics in a unified and effective manner, resulting in target representations that are neither discriminative nor temporally consistent. In this paper, we propose ProMoT, a novel tracking framework that jointly integrates cross-modal and temporal cues into a progressive prompting process, enabling continuous retrieval of target-aware representations. Specifically, we design an adaptive target query generator (QueryGen), which selectively aggregates informative spatio-temporal cues from diverse ghost representations through the dynamic sparse ghost fusion mechanism, thereby enabling the generation of target-aware queries. To further preserve fine-grained, temporally consistent target cues, we introduce a high-order contextual prompt updater (PromptUpdater), which encodes high-order cross-modal representations from current and previous frames. These prompts establish the compact and discriminative inter-frame context to not only refine the current frame’s features but also guide target localization in future frames. All components are built upon a parameter-shared backbone for RGB and TIR inputs, forming our complete ProMoT framework. Extensive experiments on both complete and missing modality RGB-T tracking benchmarks show that ProMoT consistently achieves state-of-the-art performance while balancing efficiency. Rui Xu 0028, Si Chen 0002, Yuzhen Niu, Yan Yan 0001, Dahan Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | HOH-Net: High-Order Hierarchical Middle-Feature Learning Network for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is a cross-modality retrieval task that aims to match images of the same person across visible (VIS) and infrared (IR) modalities. Existing VI-ReID methods ignore high-order structure information of features and struggle to learn a reliable common feature space due to the modality discrepancy between VIS and IR images. To alleviate the above issues, we propose a novel high-order hierarchical middle-feature learning network (HOH-Net) for VI-ReID. We introduce a high-order structure learning (HSL) module to explore the high-order relationships of short- and long-range feature nodes, for significantly mitigating model collapse and effectively obtaining discriminative features. We further develop a fine-coarse graph attention alignment (FCGA) module, which efficiently aligns multi-modality feature nodes from node-level and region-level perspectives, ensuring reliable middle-feature representations. Moreover, we exploit a hierarchical middle-feature agent learning (HMAL) loss to hierarchically reduce the modality discrepancy at each stage of the network by using the agents of middle features. The proposed HMAL loss also exchanges detailed and semantic information between low- and high-stage networks. Finally, we introduce a modality-range identity-center contrastive (MRIC) loss to minimize the distances between VIS, IR, and middle features. Extensive experiments demonstrate that the proposed HOH-Net yields state-of-the-art performance on the image-based and video-based VI-ReID datasets. The code is available at: https://github.com/Jaulaucoeng/HOS-Net. Liuxiang Qiu, Si Chen 0002, Jing-Hao Xue, Dahan Wang, Shunzhi Zhu, Yan Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | CAMD: Context-Aware Masked Distillation for General Self-Supervised Facial Representation Pre-TrainingabstractSelf-supervised pre-training has been shown to effectively learn transferable representations from unlabeled images in many visual tasks. However, existing self-supervised pre-training methods lack sufficient context-awareness and are difficult to obtain fine-grained facial representations, thus resulting in the weak generalization ability of the model to deal with various facial analysis tasks. To address this issue, we propose a Context-Aware Masked Distillation method, termed CAMD, to effectively learn general facial representations for fine-grained facial analysis tasks. The CAMD method first designs an innovative local-to-global masked image modeling framework to learn the contextual spatial structures and semantic relationships between local and global features, enabling effective self-supervised pre-training. In this framework, our pre-training task predicts the dense global feature representations based on the visible local feature representations after masking, so as to achieve semantic alignment across local and global views and significantly enhance spatial sensitivity. Moreover, the CAMD method leverages an attention-driven cross-view hierarchical distillation module to fully distill the features of related regions between different encoder layers of the online and target encoders. This module can learn contextual dependencies and capture discriminative fine-grained facial feature representations. Our method is evaluated on multiple downstream facial analysis tasks, including face alignment, face parsing, facial attribute recognition, facial expression recognition, and head pose estimation, all achieving state-of-the-art results and exhibiting the strong generality and effectiveness. The code is available at: https://github.com/mumumu-wss/CAMD. Sensen Wang 0001, Si Chen 0002, Dahan Wang, Yang Hua 0001, Yan Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Video-Level Cross-Modal Temporal-Navigation for RGBT TrackingabstractRGBT tracking has recently garnered significant attention due to its all-weather tracking capability. Traditional RGBT tracking methods primarily concentrate on the fusion of cross-modal spatial information. However, these methods ignore contextual relationships between consecutive video frames and lack effective interactions between modalities, easily resulting in tracking drift gradually due to the accumulation of errors. To avoid this limitation, we propose a novel Video-Level Cross-Modal Temporal-Navigation method termed VCT for robust RGBT tracking, which fully leverages the complementary spatio-temporal information across modalities to improve cross-modal tracking accuracy. The VCT employs a simple, flexible, and effective video-level dual-stream architecture that accommodates video sequences of arbitrary length, enabling RGB and TIR streams to capture and synergize spatio-temporal features across frames. To achieve temporal consistency and adaptability, we design a Cross-Modal Temporal Prompt Navigator (CM-TPN) that dynamically aggregates and compresses the historical frame context to navigate predictions of subsequent frames through temporal prompts. In addition, we introduce a Modality-Specific Mixture of Adapters (MS-MoA) to promote the spatio-temporal interaction both within and between modalities, thereby dramatically adapting to appearance changes. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the four popular RGBT tracking benchmarks. Si Chen 0002, Dahan Wang, Xu-Yao Zhang, Yan Yan 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | AMST: Object tracking based on collaborative framework with adaptive multi-strategy
Rui Xu 0028, Si Chen 0002, Yan Yan 0001, Dahan Wang, Shunzhi Zhu |
Inf. Sci. | 2 |
| 2025 | Sampling Consensus by Neighborhood Interaction Information for Remote Sensing Image MatchingabstractIn remote sensing and photogrammetry, establishing reliable feature correspondences between two sets of feature points is a critical preprocessing step. In this letter, we propose a novel outlier removal method, termed sampling consensus by neighborhood interaction information (SACNI), to accurately distinguish true matches (i.e., inliers) from false matches (i.e., outliers) for remote sensing image matching. Inspired by social networks, where two individuals with close ties share numerous common connections, we propose a neighborhood interactive representation (NIR) strategy to effectively evaluate correlations among feature matches, thereby guiding the efficient sampling of outlier-free data subsets. This representation is integrated into both the initial data subset selection and optimization phases, enhancing the overall matching performance. Extensive experiments on challenging remote sensing datasets show the superiority of the proposed SACNI over several other state-of-the-art methods. Hanlin Guo, Zhao Deng, Si Chen 0002, Dahan Wang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | DHLA: Dynamic Hybrid Label Assignment for End-to-End Object DetectionabstractThe recent one-to-one label assignment plays a crucial role in removing the last non-differentiable component, i.e., Non-Maximum Suppression (NMS), used in the post-processing step of the one-to-many label assignment, thus building an efficient end-to-end detection system. However, due to the limited number of foreground samples, the one-to-one label assignment often suffers from insufficient representation learning, and its performance is inferior to that of traditional detectors trained using the one-to-many label assignment. To solve these problems, we introduce a novel Dynamic Hybrid Label Assignment (DHLA) method, including a Hybrid Sample Selection (HSS) strategy and a Stage-aware Soft-label Adjustment (SSA) mechanism. In order to enhance the ability of representation learning of the one-to-one label assignment, the HSS strategy subtly integrates the one-to-many and the one-to-one label assignment rules to form a simple and effective hybrid assignment rule, where high-quality samples are selected for training according to an effective task consistency metric. Moreover, the SSA mechanism dynamically adjusts the contributions of different foreground samples at different training stages, thus effectively achieving the transition from one-to-many to one-to-one label assignment. In addition, we leverage a ranking loss function to widen the score gaps between the highest scoring position and surrounding areas for effectively removing duplicate bounding boxes. As a result, our method not only learns robust feature representations during training but also performs efficient end-to-end detection during inference. Extensive experiments demonstrate our method achieves competitive performance compared to state-of-the-art detectors on the challenging COCO and CrowdHuman datasets. Zhi-Liang Hu, Si Chen 0002, Yang Hua 0001, Dahan Wang, Shunzhi Zhu, Yan Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Hierarchical Attention-Enhanced Correlation Refinement for Robust Visual TrackingabstractIn recent years, visual tracking has witnessed remarkable advancements with the exploration of feature extraction and correlation modeling techniques. However, inadequate robustness of either the backbone network or the correlation operation continues to plague existing trackers, leading to frustrating drift when confronted with similar distractors or cluttered backgrounds. To address this problem, we propose a hierarchical attention-enhanced correlation refinement network (HarNet) for achieving robust visual tracking. Specifically, a gated dual-view attention (GDA) module is first designed to aggregate the intra-layer attention and the inter-layer self-attention based on a fusion gate, so as to enhance hierarchical feature representations of the template. Meanwhile, a target-aware attention (TA) module introduces the template information to the inter-layer self-attention, which can highlight the target information in the search region. Moreover, a graph guided correlation (GGC) module leverages the pixel-to-local and pixel-to-global correlations to fully exploit both local-and global-spatial information between the template and the search region, and then uses the graph convolutional network (GCN) to further learn the node relationships of the correlation map for more finegrained correlations. Thus, with the above three elaborately designed modules, the HarNet is beneficial for the enhancement of feature representation and the precise localization of the target. Extensive experiments on popular visual tracking datasets (including OTB100, VOT2016, VOT2018, VOT2019, UAV123, UAV20L, GOT-10k, and LaSOT) demonstrate the superiority of our proposed method against several state-of-the-art tracking methods. Si Chen 0002, Rui Xu 0028, Yan Yan 0001, Yang Hua 0001, Dahan Wang, Shunzhi Zhu |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | Hierarchical Token-Aware Cross-Modality Reconstruction for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) aims to query the same pedestrian's visible (infrared) images in the gallery set from the infrared (visible) images. VI-ReID not only needs to deal with the challenging factors like pose variation and occlusion, but also requires handling the large modality discrepancy. Previous methods mainly focus on learning single-scale modality-shared features and do not effectively explore the multi-scale features of two modalities from both short-range and long-range perspectives. In order to solve these problems, this paper proposes a novel Hierarchical Token-Aware Cross-Modality Reconstruction (HTCR) network to significantly mitigate the modality discrepancy for effective VI-ReID. The HTCR network consists of two main components, i.e., Hierarchical Token-aware Fusion (HTF) and Cross-modality Feature Reconstruction (CFR). The HTF module first bidirectionally exchanges the short-range and long-range multi-scale modality-shared features with a few learnable tokens to achieve discriminative pedestrian features by making full use of the advantages of both Convolutional Neural Network (CNN) and Transformer. Moreover, the CFR module reconstructs global and local pedestrian features of one modality by using the token sequence of the other modality with multi-scale cues to further explore the relationship between the two distinct modalities and alleviate the modality discrepancy. In addition, the Modality-shared feature Reconstruction (MR) loss is leveraged to reduce the noises between the reconstructed and the target features. Experimental results indicate that the proposed HTCR can significantly improve the VI-ReID performance and outperform the state-of-the-art methods on the cross-modality SYSU-MM01, RegDB, and LLCM datasets. Si Chen 0002, Liuxiang Qiu, Dahan Wang, Wentao Zhu 0002, Yang Hua 0001, Yan Yan 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Knowledge Distillation Meets Label Noise Learning: Ambiguity-Guided Mutual Label RefineryabstractKnowledge distillation (KD), which aims at transferring the knowledge from a complex network (a teacher) to a simpler and smaller network (a student), has received considerable attention in recent years. Typically, most existing KD methods work on well-labeled data. Unfortunately, real-world data often inevitably involve noisy labels, thus leading to performance deterioration of these methods. In this article, we study a little-explored but important issue, i.e., KD with noisy labels. To this end, we propose a novel KD method, called ambiguity-guided mutual label refinery KD (AML-KD), to train the student model in the presence of noisy labels. Specifically, based on the pretrained teacher model, a two-stage label refinery framework is innovatively introduced to refine labels gradually. In the first stage, we perform label propagation (LP) with small-loss selection guided by the teacher model, improving the learning capability of the student model. In the second stage, we perform mutual LP between the teacher and student models in a mutual-benefit way. During the label refinery, an ambiguity-aware weight estimation (AWE) module is developed to address the problem of ambiguous samples, avoiding overfitting these samples. One distinct advantage of AML-KD is that it is capable of learning a high-accuracy and low-cost student model with label noise. The experimental results on synthetic and real-world noisy datasets show the effectiveness of our AML-KD against state-of-the-art KD methods and label noise learning (LNL) methods. Code is available at https://github.com/Runqing-forMost/ AML-KD. Runqing Jiang, Yan Yan 0001, Jing-Hao Xue, Si Chen 0002, Nannan Wang 0001, Hanzi Wang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | High-Order Structure Based Middle-Feature Learning for Visible-Infrared Person Re-identificationabstractVisible-infrared person re-identification (VI-ReID) aims to retrieve images of the same persons captured by visible (VIS) and infrared (IR) cameras. Existing VI-ReID methods ignore high-order structure information of features while being relatively difficult to learn a reasonable common feature space due to the large modality discrepancy between VIS and IR images. To address the above problems, we propose a novel high-order structure based middle-feature learning network (HOS-Net) for effective VI-ReID. Specifically, we first leverage a short- and long-range feature extraction (SLE) module to effectively exploit both short-range and long-range features. Then, we propose a high-order structure learning (HSL) module to successfully model the high-order relationship across different local features of each person image based on a whitened hypergraph network. This greatly alleviates model collapse and enhances feature representations. Finally, we develop a common feature space learning (CFL) module to learn a discriminative and reasonable common feature space based on middle features generated by aligning features from different modalities and ranges. In particular, a modality-range identity-center contrastive (MRIC) loss is proposed to reduce the distances between the VIS, IR, and middle features, smoothing the training process. Extensive experiments on the SYSU-MM01, RegDB, and LLCM datasets show that our HOS-Net achieves superior state-of-the-art performance. Our code is available at https://github.com/Jaulaucoeng/HOS-Net. Liuxiang Qiu, Si Chen 0002, Yan Yan 0001, Jing-Hao Xue, Dahan Wang, Shunzhi Zhu |
AAAI | 2 |
| 2024 | Adjustable Gating Prompt Transformer for Facial Attribute Recognition with Limited Labeled Data
Qinxian Ye, Si Chen 0002, Dahan Wang, Nanfeng Jiang, Yanfei Su, Yan Yan 0001 |
ICPR (28) | 2 |
| 2024 | GCAT: graph calibration attention transformer for robust object tracking
Si Chen 0002, Xinxin Hu, Dahan Wang, Yan Yan 0001, Shunzhi Zhu |
Neural Comput. Appl. | 1 |
| 2024 | HASI: Hierarchical Attention-Aware Spatio-Temporal Interaction for Video-Based Person Re-IdentificationabstractVideo-based person re-identification (re-ID) aims to match the same pedestrian of video sequences across non-overlapping cameras. Video re-ID methods generally adopt frame-level feature extraction for different video frames, but they still lack effective spatio-temporal interaction, easily leading to the multi-frame misalignment problem. In this paper, we propose a Hierarchical Attention-aware Spatio-temporal Interaction (HASI) network, including an Attention-aware Temporal Interaction (ATI) module and a Hierarchical Local-spatial Enhancement (HLE) module for video-based person re-ID. In order to avoid the spatial misalignment between video frames, the ATI module employs multiple Frame-to-Frame Temporal Interaction (2FTI) blocks with the Multi-head Inter-frame Alignment Attention (MIAA) to make the current frame iteratively interact with each rest frame of a video in a positive single-cycle manner, rather than only interacting with the adjacent frame or directly building the relationship of all frames at once. This module can not only obtain the long-range non-adjacent temporal information, but also learn the pairwise frame-to-frame relationships. Moreover, the HLE module is designed to enhance the local fine-grained features from multiple Transformer layers, whilst delivering low-level information to further enrich middle-level and high-level semantic knowledge. Thus, our method can learn multi-perspective pedestrian information, including inter-frame long-range interaction information and intra-frame multi-layer global and local information. Extensive experiments demonstrate the superiority of the proposed HASI method compared with the state-of-the-art methods on the three challenging video-based re-ID datasets, i.e., MARS, iLIDS-VID, and PRID-2011. Si Chen 0002, Hui Da, Dahan Wang, Xu-Yao Zhang, Yan Yan 0001, Shunzhi Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Relationship-Guided Knowledge Transfer for Class-Incremental Facial Expression RecognitionabstractHuman emotions contain both basic and compound facial expressions. In many practical scenarios, it is difficult to access all the compound expression categories at one time. In this paper, we investigate comprehensive facial expression recognition (FER) in the class-incremental learning paradigm, where we define well-studied and easily-accessible basic expressions as initial classes and learn new compound expressions incrementally. To alleviate the stability-plasticity dilemma in our incremental task, we propose a novel Relationship-Guided Knowledge Transfer (RGKT) method for class-incremental FER. Specifically, we develop a multi-region feature learning (MFL) module to extract fine-grained features for capturing subtle differences in expressions. Based on the MFL module, we further design a basic expression-oriented knowledge transfer (BET) module and a compound expression-oriented knowledge transfer (CET) module, by effectively exploiting the relationship across expressions. The BET module initializes the new compound expression classifiers based on expression relevance between basic and compound expressions, improving the plasticity of our model to learn new classes. The CET module transfers expression-generic knowledge learned from new compound expressions to enrich the feature set of old expressions, facilitating the stability of our model against forgetting old classes. Extensive experiments on three facial expression databases show that our method achieves superior performance in comparison with several state-of-the-art methods. Yuanling Lv, Yan Yan 0001, Jing-Hao Xue, Si Chen 0002, Hanzi Wang |
IEEE Trans. Image Process. | 4 |
| 2024 | Visual-Textual Attribute Learning for Class-Incremental Facial Expression RecognitionabstractIn this paper, we study facial expression recognition (FER) in the class-incremental learning (CIL) setting, which defines the classification of well-studied and easily-accessible basic expressions as an initial task while learning new compound expressions gradually. Motivated by the fact that compound expressions are meaningful combinations of basic expressions, we treat basic expressions as attributes (i.e., semantic descriptors), and thus compound expressions are represented in terms of attributes. To this end, we propose a novel visual-textual attribute learning network (VTA-Net), mainly consisting of a textual-guided visual module (TVM) and a textual compositional module (TCM), for class-incremental FER. Specifically, TVM extracts textual-aware visual features and classifies expressions by incorporating the textual information into visual attribute learning. Meanwhile, TCM generates visual-aware textual features and predicts expressions by exploiting the dependency between textual attributes and category names of old and new expressions based on a textual compositional graph. In particular, a visual-textual distillation loss is introduced to calibrate TVM and TCM during incremental learning. Finally, the outputs from TVM and TCM are fused to make a final prediction. On the one hand, at each incremental task, the representations of visual attributes are enhanced since visual attributes are shared across old and new expressions. This increases the stability of our method. On the other hand, the textual modality, which involves rich prior knowledge of the relevance between expressions, facilitates our model to identify subtle visual distinctions between compound expressions, improving the plasticity of our method. Experimental results on both in-the-lab and in-the-wild facial expression databases show the superiority of our method against several state-of-the-art methods for class-incremental FER. Yuanling Lv, Guangyu Huang, Yan Yan 0001, Jing-Hao Xue, Si Chen 0002, Hanzi Wang |
IEEE Trans. Multim. | 5 |
| 2023 | Multi-Zone Transformer Based on Self-Distillation for Facial Attribute RecognitionabstractRecently, transformers have shown great promising performance in various computer vision tasks. However, the current transformer based methods ignore the information exchanges between transformer blocks, and they have not been applied in the facial attribute recognition task. In this paper, we propose a multi-zone transformer based on self-distillation for FAR, termed MZTS, to predict the facial attributes. A multi-zone transformer encoder is firstly presented to achieve the interactions of the different transformer encoder blocks, thus avoiding forgetting the effective information between the transformer encoder block groups during the iteration process. Furthermore, we introduce a new self-distillation mechanism based on class tokens, which distills the class tokens obtained from the last transformer encoder block group to the other shallow groups by interacting with the significant information between the different transformer blocks through attention. Extensive experiments on the challenging CelebA and LFWA datasets have demonstrated the excellent performance of the proposed method for FAR. Si Chen 0002, Xueyan Zhu, Dahan Wang, Shunzhi Zhu, Yun Wu 0001 |
FG | 1 |
| 2023 | SiamCCF: Siamese visual tracking via cross-layer calibration fusionabstractAbstract Siamese networks have attracted wide attention in visual tracking due to their competitive accuracy and speed. However, the existing Siamese trackers usually leverage a fixed linear aggregation of feature maps, which does not effectively fuse the different layers of features with attention. Besides, most of Siamese trackers calculate the similarity between the template and the search region through a cross‐correlation operation between the features of the last blocks from the two branches, which might introduce the redundant noise information. In order to solve these problems, this study proposes a novel Siamese visual tracking method via cross‐layer calibration fusion, termed SiamCCF. An attention‐based feature fusion module is employed using local attention and non‐local attention to fuse the features from the deep and shallow layers, so as to capture both local details and high‐level semantic information. Moreover, a cross‐layer calibration module can use the fused features to calibrate the features of the last network blocks and build the cross‐layer long‐range spatial and inter‐channel dependencies around each spatial location. Extensive experiments demonstrate that the proposed method has achieved competitive tracking performance compared with state‐of‐the‐art trackers on challenging benchmarks, including OTB100, OTB2013, UAV123, UAV20L, and LaSOT. Si Chen 0002, Shunzhi Zhu, Huarong Xu, Dahan Wang |
IET Comput. Vis. | 1 |
| 2023 | SPL-Net: Spatial-Semantic Patch Learning Network for Facial Attribute Recognition with Limited Labeled Data
Yan Yan 0001, Ying Shu, Si Chen 0002, Jing-Hao Xue, Chunhua Shen, Hanzi Wang |
Int. J. Comput. Vis. | 3 |
| 2023 | Learning an attention-aware parallel sharing network for facial attribute recognition
Si Chen 0002, Xinyu Lai, Yan Yan 0001, Dahan Wang, Shunzhi Zhu |
J. Vis. Commun. Image Represent. | 1 |
| 2023 | MTNet: Mutual tri-training network for unsupervised domain adaptation on person re-identification
Si Chen 0002, Liuxiang Qiu, Zimin Tian, Yan Yan 0001, Dahan Wang, Shunzhi Zhu |
J. Vis. Commun. Image Represent. | 1 |
| 2023 | HC-GCN: hierarchical contrastive graph convolutional network for unsupervised domain adaptation on person re-identification
Si Chen 0002, Bolun Xu, Yan Yan 0001, Xia Du, Weiwei Zhuang, Yun Wu 0001 |
Multim. Syst. | 1 |
| 2023 | Identity-Aware Contrastive Knowledge Distillation for Facial Attribute RecognitionabstractFacial attribute recognition (FAR) is an important and yet challenging multi-label learning task in computer vision. Existing FAR methods have achieved promising performance with the development of deep learning. However, they usually suffer from prohibitive computational and memory costs. In this paper, we propose an identity-aware contrastive knowledge distillation method, termed ICKD, to compress the FAR model. A nonlinear weight-sharing mapping (NWSM) mechanism is firstly designed to avoid the difficulty of directly matching features of the teacher and student networks due to the lower representation ability of the student network. Furthermore, an identity-aware contrastive distillation (ICD) loss is employed to guide the student network to effectively learn the mutual relations between samples with multiple attributes. In addition, an adjustable ladder distillation (ALD) loss is developed to automatically adjust the importance of different distillation points with the progress of training. Extensive experiments demonstrate that our method can significantly improve the performance of student networks and outperforms the existing FAR methods on the public challenging datasets. Si Chen 0002, Xueyan Zhu, Yan Yan 0001, Shunzhi Zhu, Shaozi Li, Dahan Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | When Facial Expression Recognition Meets Few-Shot Learning: A Joint and Alternate Learning FrameworkabstractHuman emotions involve basic and compound facial expressions. However, current research on facial expression recognition (FER) mainly focuses on basic expressions, and thus fails to address the diversity of human emotions in practical scenarios. Meanwhile, existing work on compound FER relies heavily on abundant labeled compound expression training data, which are often laboriously collected under the professional instruction of psychology. In this paper, we study compound FER in the cross-domain few-shot learning setting, where only a few images of novel classes from the target domain are required as a reference. In particular, we aim to identify unseen compound expressions with the model trained on easily accessible basic expression datasets. To alleviate the problem of limited base classes in our FER task, we propose a novel Emotion Guided Similarity Network (EGS-Net), consisting of an emotion branch and a similarity branch, based on a two-stage learning framework. Specifically, in the first stage, the similarity branch is jointly trained with the emotion branch in a multi-task fashion. With the regularization of the emotion branch, we prevent the similarity branch from overfitting to sampled base classes that are highly overlapped across different episodes. In the second stage, the emotion branch and the similarity branch play a “two-student game” to alternately learn from each other, thereby further improving the inference ability of the similarity branch on unseen compound expressions. Experimental results on both in-the-lab and in-the-wild compound expression datasets demonstrate the superiority of our proposed method against several state-of-the-art methods. Xinyi Zou, Yan Yan 0001, Jing-Hao Xue, Si Chen 0002, Hanzi Wang |
AAAI | 4 |
| 2022 | Learn-to-Decompose: Cascaded Decomposition Network for Cross-Domain Few-Shot Facial Expression Recognition
Xinyi Zou, Yan Yan 0001, Jing-Hao Xue, Si Chen 0002, Hanzi Wang |
ECCV (19) | 4 |
| 2022 | MSFL-Net: Multi-Semantic Feature Learning Network for Occluded Person Re-IdentificationabstractRecently, occluded person re-identification (Re-ID) has received significant interest due to its widespread real-world applications. However, most existing occluded person Re-ID methods ignore semantic granularities that indicate different levels of occluded information of the human body, leading to sub-optimal performance. To address this, we propose a Multi-Semantic Feature Learning Network (MSFL-Net) for occluded person Re-ID. Specifically, MSFL-Net involves a backbone network and a Multi-branch Feature Learning sub-network (MFL). MFL consists of two local-global branches and a global branch to learn multisemantic features in a multi-branch deep network architecture. In each local-global branch, we design a local subbranch and a semantic-guided global sub-branch to extract discriminative features at a certain level of feature granularity and semantic granularity. In the global branch, we learn global features at the largest level of semantic granularity. In particular, a patch contrastive loss is developed to explicitly encourage the semantic feature maps to capture the information from specific body parts. By extracting multi-semantic features, our method is effective in dealing with person Re-ID at different occlusion levels. Experimental results on an occluded person Re-ID dataset (Occluded-REID) and two partial person Re-ID datasets (Partial-iLIDS and Partial-REID) show the superiority of our method against state-of-the-art person Re-ID methods. Guangyu Huang, Yan Yan 0001, Si Chen 0002, Wentao Zhu 0002, Hanzi Wang |
IJCB | 4 |
| 2022 | Graph Attention Transformer Network for Robust Visual Tracking
Si Chen 0002, Dahan Wang, Shunzhi Zhu |
ICONIP (4) | 2 |
| 2022 | Adaptive Deep Disturbance-Disentangled Learning for Facial Expression Recognition
Delian Ruan, Rongyun Mo, Yan Yan 0001, Si Chen 0002, Jing-Hao Xue, Hanzi Wang |
Int. J. Comput. Vis. | 4 |
| 2022 | Learning meta-adversarial features via multi-stage adaptation network for robust visual object tracking
Si Chen 0002, Yan Yan 0001, Dahan Wang, Shunzhi Zhu |
Neurocomputing | 1 |
| 2022 | Stage-Aware Feature Alignment Network for Real-Time Semantic Segmentation of Street ScenesabstractOver the past few years, deep convolutional neural network-based methods have made great progress in semantic segmentation of street scenes. Some recent methods align feature maps to alleviate the semantic gap between them and achieve high segmentation accuracy. However, they usually adopt the feature alignment modules with the same network configuration in the decoder and thus ignore the different roles of stages of the decoder during feature aggregation, leading to a complex decoder structure. Such a manner greatly affects the inference speed. In this paper, we present a novel Stage-aware Feature Alignment Network (SFANet) based on the encoder-decoder structure for real-time semantic segmentation of street scenes. Specifically, a Stage-aware Feature Alignment module (SFA) is proposed to align and aggregate two adjacent levels of feature maps effectively. In the SFA, by taking into account the unique role of each stage in the decoder, a novel stage-aware Feature Enhancement Block (FEB) is designed to enhance spatial details and contextual information of feature maps from the encoder. In this way, we are able to address the misalignment problem with a very simple and efficient multi-branch decoder structure. Moreover, an auxiliary training strategy is developed to explicitly alleviate the multi-scale object problem without bringing additional computational costs during the inference phase. Experimental results show that the proposed SFANet exhibits a good balance between accuracy and speed for real-time semantic segmentation of street scenes. In particular, based on ResNet-18, SFANet respectively obtains 78.1% and 74.7% mean of class-wise Intersection-over-Union (mIoU) at inference speeds of 37 FPS and 96 FPS on the challenging Cityscapes and CamVid test datasets by using only a single GTX 1080Ti GPU. Xi Weng, Yan Yan 0001, Si Chen 0002, Jing-Hao Xue, Hanzi Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Learning Spatial-Semantic Relationship for Facial Attribute Recognition With Limited Labeled DataabstractRecent advances in deep learning have demonstrated excellent results for Facial Attribute Recognition (FAR), typically trained with large-scale labeled data. However, in many real-world FAR applications, only limited labeled data are available, leading to remarkable deterioration in performance for most existing deep learning-based FAR methods. To address this problem, here we propose a method termed Spatial-Semantic Patch Learning (SSPL). The training of SSPL involves two stages. First, three auxiliary tasks, consisting of a Patch Rotation Task (PRT), a Patch Segmentation Task (PST), and a Patch Classification Task (PCT), are jointly developed to learn the spatial-semantic relationship from large-scale unlabeled facial data. We thus obtain a powerful pre-trained model. In particular, PRT exploits the spatial information of facial images in a self-supervised learning manner. PST and PCT respectively capture the pixel-level and image-level semantic information of facial images based on a facial parsing model. Second, the spatial-semantic knowledge learned from auxiliary tasks is transferred to the FAR task. By doing so, it enables that only a limited number of labeled data are required to fine-tune the pre-trained model. We achieve superior performance compared with state-of-the-art methods, as substantiated by extensive experiments and studies. Ying Shu, Yan Yan 0001, Si Chen 0002, Jing-Hao Xue, Chunhua Shen, Hanzi Wang |
CVPR | 3 |
| 2021 | D³Net: Dual-Branch Disturbance Disentangling Network for Facial Expression RecognitionabstractOne of the main challenges in facial expression recognition (FER) is to address the disturbance caused by various disturbing factors, including common ones (such as identity, pose, and illumination) and potential ones (such as hairstyle, accessory, and occlusion). Recently, a number of FER methods have been developed to explicitly or implicitly alleviate the disturbance involved in facial images. However, these methods either consider only a few common disturbing factors or neglect the prior information of these disturbing factors, thus resulting in inferior recognition performance. In this paper, we propose a novel Dual-branch Disturbance Disentangling Network (D3Net), mainly consisting of an expression branch and a disturbance branch, to perform effective FER. In the disturbance branch, a label-aware sub-branch (LAS) and a label-free sub-branch (LFS) are elaborately designed to cope with different types of disturbing factors. On the one hand, LAS explicitly captures the disturbance due to some common disturbing factors by transfer learning on a pretrained model. On the other hand, LFS implicitly encodes the information of potential disturbing factors in an unsupervised manner. In particular, we introduce an Indian buffet process (IBP) prior to model the distribution of potential disturbing factors in LFS. Moreover, we leverage adversarial training to increase the differences between disturbance features and expression features, thereby enhancing the disentanglement of disturbing factors. By disentangling the disturbance from facial images, we are able to extract discriminative expression features. Extensive experiments demonstrate that our proposed method performs favorably against several state-of-the-art FER methods on both in-the-lab and in-the-wild databases. Rongyun Mo, Yan Yan 0001, Jing-Hao Xue, Si Chen 0002, Hanzi Wang |
ACM Multimedia | 4 |
| 2020 | Deep Disturbance-Disentangled Learning for Facial Expression RecognitionabstractTo achieve effective facial expression recognition (FER), it is of great importance to address various disturbing factors, including pose, illumination, identity, and so on. However, a number of FER databases merely provide the labels of facial expression, identity, and pose, but lack the label information for other disturbing factors. As a result, many methods are only able to cope with one or two disturbing factors, ignoring the heavy entanglement between facial expression and multiple disturbing factors. In this paper, we propose a novel Deep Disturbance-disentangled Learning (DDL) method for FER. DDL is capable of simultaneously and explicitly disentangling multiple disturbing factors by taking advantage of multi-task learning and adversarial transfer learning. The training of DDL involves two stages. First, a Disturbance Feature Extraction Model (DFEM) is pre-trained to perform multi-task learning for classifying multiple disturbing factors on the large-scale face database (which has the label information for various disturbing factors). Second, a Disturbance-Disentangled Model (DDM), which contains a global shared sub-network and two task-specific (i.e., expression and disturbance) sub-networks, is learned to encode the disturbance-disentangled information for expression recognition. The expression sub-network adopts a multi-level attention mechanism to extract expression-specific features, while the disturbance sub-network leverages adversarial transfer learning to extract disturbance-specific features based on the pre-trained DFEM. Experimental results on both the in-the-lab FER databases (including CK+, MMI, and Oulu-CASIA) and the in-the-wild FER databases (including RAF-DB and SFEW) demonstrate the superiority of our proposed method compared with several state-of-the-art methods. Delian Ruan, Yan Yan 0001, Si Chen 0002, Jing-Hao Xue, Hanzi Wang |
ACM Multimedia | 3 |
| 2020 | Object-adaptive LSTM network for real-time visual tracking with adversarial data augmentation
Yihan Du, Yan Yan 0001, Si Chen 0002, Yang Hua 0001 |
Neurocomputing | 3 |
| 2020 | Low-resolution facial expression recognition: A filter learning perspective
Yan Yan 0001, Zizhao Zhang 0002, Si Chen 0002, Hanzi Wang |
Signal Process. | 3 |
| 2020 | Joint Deep Learning of Facial Expression Synthesis and RecognitionabstractRecently, deep learning based facial expression recognition (FER) methods have attracted considerable attention and they usually require large-scale labelled training data. Nonetheless, the publicly available facial expression databases typically contain a small amount of labelled data. In this paper, to overcome the above issue, we propose a novel joint deep learning of facial expression synthesis and recognition method for effective FER. More specifically, the proposed method involves a two-stage learning procedure. Firstly, a facial expression synthesis generative adversarial network (FESGAN) is pre-trained to generate facial images with different facial expressions. To increase the diversity of the training images, FESGAN is elaborately designed to generate images with new identities from a prior distribution. Secondly, an expression recognition network is jointly learned with the pre-trained FESGAN in a unified framework. In particular, the classification loss computed from the recognition network is used to simultaneously optimize the performance of both the recognition network and the generator of FESGAN. Moreover, in order to alleviate the problem of data bias between the real images and the synthetic images, we propose an intra-class loss with a novel real data-guided back-propagation (RDBP) algorithm to reduce the intra-class variations of images from the same class, which can significantly improve the final performance. Extensive experimental results on public facial expression databases demonstrate the superiority of the proposed method compared with several state-of-the-art FER methods. Yan Yan 0001, Si Chen 0002, Chunhua Shen, Hanzi Wang |
IEEE Trans. Multim. | 3 |
| 2019 | A Cascaded Noise-Robust Deep CNN for Face RecognitionabstractState-of-the-art face recognition methods have achieved excellent performance on the clean datasets. However, in real-world applications, the captured face images are usually contaminated with noise, which significantly decreases the performance of these face recognition methods. In this paper, we propose a cascaded noise-robust deep convolutional neural network (CNR-CNN) method, consisting of two sub-networks, i.e., a denoising sub-network and a face recognition sub-network, for face recognition under noise. Instead of separately training the two sub-networks, we jointly train them in a cascaded manner. As a result, the images generated from the denoising sub-network are beneficial to the training of the face recognition sub-network. Furthermore, the dense connectivity is used to concatenate the feature maps layer-by-layer in the denoising sub-network, which can effectively exploit the shallow and deep features of CNN. Experimental results on public face datasets demonstrate the superior performance of the proposed method over several state-of-the-art methods. Xiangbang Meng, Yan Yan 0001, Si Chen 0002, Hanzi Wang |
ICIP | 3 |
| 2019 | Adaptive deep metric embeddings for person re-identification under occlusions
Wanxiang Yang, Yan Yan 0001, Si Chen 0002 |
Neurocomputing | 3 |
| 2018 | Object-Adaptive LSTM Network for Visual TrackingabstractConvolutional Neural Networks (CNNs) have shown outstanding performance in visual object tracking. However, most of classification-based tracking methods using CNNs are time-consuming due to expensive computation of complex online fine-tuning and massive feature extractions. Besides, these methods suffer from the problem of over-fitting since the training and testing stages of CNN models are based on the videos from the same domain. Recently, matching-based tracking methods (such as Siamese networks) have shown remarkable speed superiority, while they cannot well address target appearance variations and complex scenes for inherent lack of online adaptability and background information. In this paper, we propose a novel object-adaptive LSTM network, which can effectively exploit sequence dependencies and dynamically adapt to the temporal object variations via constructing an intrinsic model for object appearance and motion. In addition, we develop an efficient strategy for proposal selection, where the densely sampled proposals are firstly pre-evaluated using the fast matching-based method and then the well-selected high-quality proposals are fed to the sequence-specific learning LSTM network. This strategy enables our method to adaptively track an arbitrary object and operate faster than conventional CNN-based classification tracking methods. To the best of our knowledge, this is the first work to apply an LSTM network for classification in visual object tracking. Experimental results on OTB and TC-128 benchmarks show that the proposed method achieves state-of-the-art performance, which exhibits great potentials of recurrent structures for visual object tracking. Yihan Du, Yan Yan 0001, Si Chen 0002, Yang Hua 0001, Hanzi Wang |
ICPR | 3 |
| 2018 | Multi-task Learning of Cascaded CNN for Facial Attribute ClassificationabstractRecently, facial attribute classification (FAC) has attracted significant attention in the computer vision community. Great progress has been made along with the availability of challenging FAC datasets. However, conventional FAC methods usually firstly pre-process the input images (i.e., perform face detection and alignment) and then predict facial attributes. These methods ignore the inherent dependencies among these tasks (i.e., face detection, facial landmark localization and FAC). Moreover, some methods using convolutional neural network are trained based on the fixed loss weights without considering the differences between facial attributes. In order to address the above problems, we propose a novel multi-task learning of cascaded convolutional neural network method, termed MCFA, for predicting multiple facial attributes simultaneously. Specifically, the proposed method takes advantage of three cascaded sub-networks (i.e., S_Net, M_Net and L_Net corresponding to the neural networks under different scales) to jointly train multiple tasks in a coarse-to-fine manner, which can achieve end-to-end optimization. Furthermore, the proposed method automatically assigns the loss weight to each facial attribute based on a novel dynamic weighting scheme, thus making the proposed method concentrate on predicting the more difficult facial attributes. Experimental results show that the proposed method outperforms several state-of-the-art FAC methods on the challenging CelebA and LFWA datasets. Ni Zhuang, Yan Yan 0001, Si Chen 0002, Hanzi Wang |
ICPR | 3 |
| 2018 | Expression-targeted feature learning for effective facial expression recognition
Yan Yan 0001, Si Chen 0002, Hanzi Wang |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Multi-label learning based deep transfer neural network for facial attribute classification
Ni Zhuang, Yan Yan 0001, Si Chen 0002, Hanzi Wang, Chunhua Shen |
Pattern Recognit. | 3 |
| 2017 | An efficient deep neural networks training framework for robust face recognitionabstractIn recent years, the triplet loss-based deep neural networks (DNN) are widely used in the task of face recognition and achieve the state-of-the-art performance. However, the complexity of training the triplet loss-based DNN is significantly high due to the difficulty in generating high-quality training samples. In this paper, we propose a novel DNN training framework to accelerate the training process of the triplet loss-based DNN and meanwhile to improve the performance of face recognition. More specifically, the proposed framework contains two stages: 1) The DNN initialization. A deep architecture based on the softmax loss function is designed to initialize the DNN. 2) The adaptive fine-tuning. Based on the trained model, a set of high-quality triplet samples is generated and used to fine-tune the network, where an adaptive triplet loss function is introduced to improve the discriminative ability of DNN. Experimental results show that, the model obtained by the proposed DNN training framework achieves 97.3% accuracy on the LFW benchmark with low training complexity, which verifies the efficiency and effectiveness of the proposed framework. Canping Su, Yan Yan 0001, Si Chen 0002, Hanzi Wang |
ICIP | 3 |
| 2017 | A novel robust model fitting approach towards multiple-structure data segmentation
Yan Yan 0001, Min Liu 0006, Si Chen 0002 |
Neurocomputing | 3 |
| 2017 | Image feature detection algorithm based on the spread of Hessian source
Shunzhi Zhu, Lizhao Liu, Si Chen 0002 |
Multim. Syst. | 3 |
| 2016 | Sparse similarity metric learning for kinship verificationabstractMetric learning technique learns a linear transformation of the given training data which can significantly promote the performance of a prediction task, such as kinship verification. However, many of the existing metric learning methods do not explicitly regularize for sparsity or low-rank, which in practice usually results in high-rank solutions that are not only time-consuming but also tend to overfitting. In addition, some methods simply neglect the positive semidefinite (PSD) constraint causing the learned metric to be potentially noisy. In this paper, we propose an effective sparse similarity metric learning (SSML) method which enforces both the group sparsity and the PSD constraints on the learned similarity matrix for kinship verification. In order to solve the proposed optimization problem efficiently, we successfully apply the alternating direction method of multipliers (ADMM) to obtain the optimal solution. Experimental results demonstrate that the proposed method achieves competitive results compared with other state-of-the-art metric learning methods on widely used kinship datasets. Yan Yan 0001, Si Chen 0002, Hanzi Wang |
VCIP | 3 |
| 2016 | Discriminative local collaborative representation for online object tracking
Si Chen 0002, Shaozi Li, Rongrong Ji, Yan Yan 0001, Shunzhi Zhu |
Knowl. Based Syst. | 1 |
| 2016 | Robust visual tracking via online semi-supervised co-boosting
Si Chen 0002, Shunzhi Zhu, Yan Yan 0001 |
Multim. Syst. | 1 |
| 2016 | Quadratic projection based feature extraction with its application to biometric recognition
Yan Yan 0001, Hanzi Wang, Si Chen 0002, Xiaochun Cao, David Zhang 0001 |
Pattern Recognit. | 3 |
| 2014 | Online MIL tracking with instance-level semi-supervised learning
Si Chen 0002, Shaozi Li, Songzhi Su, Qi Tian 0001, Rongrong Ji |
Neurocomputing | 1 |
| 2014 | Online semi-supervised compressive coding for robust visual tracking
Si Chen 0002, Shaozi Li, Songzhi Su, Donglin Cao, Rongrong Ji |
J. Vis. Commun. Image Represent. | 1 |