VLDB 2026 Research / reviewers in the wild / expert
Yan Yan 0001
dblp:13/3953-1
· DBLP profile ↗
153ranked-venue papers
9as first author
92since 2021 · last 2026
0000-0002-3674-7160ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 93 · 3 first-author · 59 since 2021Artificial intelligence and machine learning · 67 · 6 first-author · 33 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 7 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Security and privacy · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Joint Implicit and Explicit Language Learning for Pedestrian Attribute RecognitionabstractPedestrian attribute recognition (PAR) has received increasing attention due to its wide application in video surveillance and pedestrian analysis. Some text-enhanced methods tackle this task by converting attributes into language descriptions to facilitate interactive learning between attributes and visual images. However, these generic languages fail to uniquely describe different pedestrian images, missing individual characteristics. In this paper, we propose a Joint Implicit and Explicit Language Guidance Enhancement Learning (JGEL) method, which converts each pedestrian image into a language description with dual language learning to effectively learn enhanced attribute information. Specifically, we first propose an Implicit Language Guidance Learning (ILGL) stream. It projects visual image features into the text embedding space to generate pseudo-word tokens, implicitly modeling image attributes and providing personalized descriptions. Moreover, we propose an Explicit Attribute Enhancement Learning (EAEL) stream to guide the generated pseudo-word tokens obtained by ILGL explicitly aligned with pedestrian attributes, which can effectively align the pseudo-word tokens with the attribute concepts in the text embedding space. Extensive experiments show that JGEL has significant advantages in improving the performance of PAR and the challenging zero-shot PAR task. Yang Lu 0009, Yan Yan 0001, Hanzi Wang |
AAAI | 4 |
| 2026 | A survey of robotic manipulation: From bottom-up approaches to end-to-end paradigms with LLMs
Qing Li 0001, Zhijian He, Bowen Zhang 0005, Xianghua Fu, Zhi-Qi Cheng, Yan Yan 0001, Xiaojiang Peng |
Neurocomputing | 9 |
| 2026 | Learning relationship-guided vision-language transformer for facial attribute recognition
Si Chen 0002, Mingxuan Lei, Dahan Wang, Xu-Yao Zhang, Yan Yan 0001, Shunzhi Zhu |
Pattern Recognit. | 5 |
| 2026 | Global-Local Disturbance Decoupling for Federated Facial Expression RecognitionabstractMost existing facial expression recognition (FER) methods are designed for centralized model training on largescale data. Unfortunately, accessing massive facial expression data can be difficult due to privacy concerns in practice. In this paper, we study an important but little-explored task, federated FER, which allows us to train an FER model with decentralized expression data. To this end, we propose a novel global-local disturbance decoupling (GLDD) method for federated FER. Specifically, for local disturbance decoupling on each client, we develop a dual-branch feature decoupling network consisting of a backbone network, an expression branch, and a disturbance branch, to perform local FER. In the disturbance branch, we design an entropy-guided feature encoding module to extract priorbased disturbance features. This greatly facilitates the extraction of client-specific disturbance features. For global disturbance decoupling on the server, we introduce orthogonal decoupling, which is a global-level feature disentanglement technique that enforces mutual orthogonality between the global expression feature class centers and the global disturbance feature centers, thereby eliminating cross-client disturbances across clients and yields decoupled global expression feature class centers. These centers are then used to retrain the global model, substantially enhancing disturbance invariance and classification performance. By jointly performing disturbance decoupling at local and global levels, our method effectively addresses the unique challenges of heterogeneous expression data and heterogeneous disturbances in federated FER. Experimental results on two real-world facial expression databases show that, on the federated FER task, our method significantly outperforms several state-of-the-art federated learning methods and FER methods. The code will be released soon. Hu Ding 0005, Yan Yan 0001, Yang Lu 0009, Jing-Hao Xue, Hanzi Wang |
IEEE Trans. Affect. Comput. | 2 |
| 2026 | ProMoT: Progressive Prompting of Modality and Temporal Dynamics for RGB-T TrackingabstractRGB-T tracking benefits from the complementary nature of RGB and TIR modalities, yet their relative reliability for target localization often shifts over time. Most existing trackers fail to adapt to such modality and temporal dynamics in a unified and effective manner, resulting in target representations that are neither discriminative nor temporally consistent. In this paper, we propose ProMoT, a novel tracking framework that jointly integrates cross-modal and temporal cues into a progressive prompting process, enabling continuous retrieval of target-aware representations. Specifically, we design an adaptive target query generator (QueryGen), which selectively aggregates informative spatio-temporal cues from diverse ghost representations through the dynamic sparse ghost fusion mechanism, thereby enabling the generation of target-aware queries. To further preserve fine-grained, temporally consistent target cues, we introduce a high-order contextual prompt updater (PromptUpdater), which encodes high-order cross-modal representations from current and previous frames. These prompts establish the compact and discriminative inter-frame context to not only refine the current frame’s features but also guide target localization in future frames. All components are built upon a parameter-shared backbone for RGB and TIR inputs, forming our complete ProMoT framework. Extensive experiments on both complete and missing modality RGB-T tracking benchmarks show that ProMoT consistently achieves state-of-the-art performance while balancing efficiency. Rui Xu 0028, Si Chen 0002, Yuzhen Niu, Yan Yan 0001, Dahan Wang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | TCFF-Adapter: Text-Driven Adaption of CLIP for Few-Shot Image ClassificationabstractIn recent years, few-shot image classification has achieved substantial progress. Although existing methods have achieved promising performance, the limited availability of training data often leads to the problem of model overfitting. Model overfitting affects generalization and restricts the effective transfer of knowledge to unseen classes. Moreover, existing methods maintain independence between the image and text modalities during the encoding process, lacking mutual collaboration. This limitation restricts their ability to fully exploit task-specific semantic relationships between visual concepts and textual descriptions. To address this challenge, we propose a text-driven cross-modal feature fusion adapter (TCFF-Adapter) for few-shot image classification. TCFF-Adapter introduces two core components: a cross-modal feature fusion module that constructs joint representations by aligning image and text semantics, and a text-driven adapter that optimizes fused features and dynamically adjusts feature weights in a meta-learning paradigm. By integrating multimodal knowledge with parameter-efficient tuning, our method achieves robust generalization to unseen data without requiring additional fine-tuning. Extensive experiments on eight benchmark datasets demonstrate that the proposed TCFF-Adapter significantly outperforms various state-of-the-art few-shot image classification methods. Guanlin Du, Hanzi Wang, Xintao Xu, Yan Yan 0001, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | HOH-Net: High-Order Hierarchical Middle-Feature Learning Network for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is a cross-modality retrieval task that aims to match images of the same person across visible (VIS) and infrared (IR) modalities. Existing VI-ReID methods ignore high-order structure information of features and struggle to learn a reliable common feature space due to the modality discrepancy between VIS and IR images. To alleviate the above issues, we propose a novel high-order hierarchical middle-feature learning network (HOH-Net) for VI-ReID. We introduce a high-order structure learning (HSL) module to explore the high-order relationships of short- and long-range feature nodes, for significantly mitigating model collapse and effectively obtaining discriminative features. We further develop a fine-coarse graph attention alignment (FCGA) module, which efficiently aligns multi-modality feature nodes from node-level and region-level perspectives, ensuring reliable middle-feature representations. Moreover, we exploit a hierarchical middle-feature agent learning (HMAL) loss to hierarchically reduce the modality discrepancy at each stage of the network by using the agents of middle features. The proposed HMAL loss also exchanges detailed and semantic information between low- and high-stage networks. Finally, we introduce a modality-range identity-center contrastive (MRIC) loss to minimize the distances between VIS, IR, and middle features. Extensive experiments demonstrate that the proposed HOH-Net yields state-of-the-art performance on the image-based and video-based VI-ReID datasets. The code is available at: https://github.com/Jaulaucoeng/HOS-Net. Liuxiang Qiu, Si Chen 0002, Jing-Hao Xue, Dahan Wang, Shunzhi Zhu, Yan Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Vision-Language Enhancement Network Based on Decoupling-Joint Adaptation for Few-Shot Action RecognitionabstractLearning robust and generalizable feature extractors to generate discriminative prototypes is crucial for few-shot action recognition. However, most existing methods rely on fine-tuning large pre-trained image models, easily leading to transferability and overfitting issues. In this paper, we propose a novel vision-language enhancement network based on decoupling-joint adaptation (VEDA) for few-shot action recognition, which decouples visual features into temporal and spatial branches, followed by a joint operation that integrates these two branches using an adapter-tuning paradigm. VEDA can gradually equip the model with spatio-temporal reasoning capabilities. Since relying exclusively on local frame feature matching results in inaccurate performance, we design a video-level relation module (VLR) to enhance video context awareness through global feature matching. In addition, we design a vision-language fusion module (VLF) that introduces multimodal information to alleviate the data scarcity issue. Simultaneously, we apply adapter-tuning to both visual and textual branches to enhance the generalization ability. Based on the proposed components above, our network can extract both informative and discriminative prototypes, resulting in excellent recognition performance. Experimental results on five challenging benchmarks demonstrate the effectiveness of the proposed VEDA. The code will be released soon at https://github.com/ReverseSuzhou/VEDA. Suzhou Que, Hanyu Guo, Kaiwen Du, Yan Yan 0001, Yanwei Pang, Hanzi Wang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | CAMD: Context-Aware Masked Distillation for General Self-Supervised Facial Representation Pre-TrainingabstractSelf-supervised pre-training has been shown to effectively learn transferable representations from unlabeled images in many visual tasks. However, existing self-supervised pre-training methods lack sufficient context-awareness and are difficult to obtain fine-grained facial representations, thus resulting in the weak generalization ability of the model to deal with various facial analysis tasks. To address this issue, we propose a Context-Aware Masked Distillation method, termed CAMD, to effectively learn general facial representations for fine-grained facial analysis tasks. The CAMD method first designs an innovative local-to-global masked image modeling framework to learn the contextual spatial structures and semantic relationships between local and global features, enabling effective self-supervised pre-training. In this framework, our pre-training task predicts the dense global feature representations based on the visible local feature representations after masking, so as to achieve semantic alignment across local and global views and significantly enhance spatial sensitivity. Moreover, the CAMD method leverages an attention-driven cross-view hierarchical distillation module to fully distill the features of related regions between different encoder layers of the online and target encoders. This module can learn contextual dependencies and capture discriminative fine-grained facial feature representations. Our method is evaluated on multiple downstream facial analysis tasks, including face alignment, face parsing, facial attribute recognition, facial expression recognition, and head pose estimation, all achieving state-of-the-art results and exhibiting the strong generality and effectiveness. The code is available at: https://github.com/mumumu-wss/CAMD. Sensen Wang 0001, Si Chen 0002, Dahan Wang, Yang Hua 0001, Yan Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Spatial-Temporal Scene Graph Generation for Open-Vocabulary Multiple Object TrackingabstractOpen-vocabulary multiple object tracking (MOT) aims to track arbitrary objects in the real world. Although significant progress has been achieved in object classification by leveraging the knowledge from large vision-language models, advances in data association for open-vocabulary MOT remain limited. Existing methods primarily rely on appearance cues to establish associations. However, these cues are often unreliable in the face of occlusions and ambiguous object appearances, resulting in suboptimal tracking performance in complex scenarios. In this paper, we propose a novel open-vocabulary MOT method, Spatial-temporal Scene Graph Tracker (SSGTrack), which introduces a fundamentally different approach to data association by building a Spatial-temporal Scene Graph (SSG) that captures rich semantic and spatial relationships between objects across adjacent frames. Specifically, SSGTrack constructs proposal-level relationships by extracting diverse contextual information from the multi-head self-attention layers of the Transformer decoder. These relationships, derived from the keyframe and reference frame, are compressed into the compact SSG, where nodes represent detected objects, and edge weights denote frame-level connectivity. Furthermore, to address the challenge of differentiating visually similar objects and background distractors, we propose a Context-aware Contrastive Learning (CCL) strategy. By identifying background features that significantly differ from positive samples and incorporating them as negative samples, CCL enhances the ability of the model to learn discriminative representations, thus improving tracking robustness. Extensive experiments conducted on several challenging MOT benchmarks demonstrate the effectiveness of our method, which achieves superior tracking performance. Siping Zhuang, Yajun Jian, Yan Yan 0001, Hanzi Wang |
IEEE Trans. Image Process. | 4 |
| 2026 | Fine-Grained Self-Paced Relational Preserving Network for Cross-Domain Few-Shot Facial Expression RecognitionabstractCross-domain few-shot facial expression recognition (CF-FER) aims to adapt models trained on basic expressions to recognize novel compound expressions using only a few annotated examples. Although vision-language models (VLMs) have shown promise in few-shot learning, their application to CF-FER remains challenging due to two key issues: coarse-grained textual prompts that fail to capture subtle variations among compound expressions, and episodic training that tends to overfit on highly overlapping few-shot tasks. To address these issues, we propose a fine-grained self-paced relational preserving network (FSR-Net), which introduces fine-grained action unit (AU)-aware textual descriptions generated by large language models (LLMs) to enrich semantic representations and provide more discriminative prototypes. Based on this, we introduce a self-paced relational preserving regularization (SPR) strategy that leverages structural discrepancies between teacher-student visual features and textual-enhanced prototypes as reliability indicators. By progressively weighting reliable samples while filtering out harder ones, the regularization strategy explicitly preserves relational consistency across samples and mitigates overfitting in CF-FER. Comprehensive experiments on multiple CF-FER benchmarks confirm the effectiveness of FSR-Net, yielding average improvements of 5.78% (1-shot) and 3.80% (5-shot) over prior state-of-the-art methods. These results demonstrate its superior capacity for capturing subtle expression cues and enhancing cross-domain transferability. Kaiyun Wang, Hanzi Wang, Yan Yan 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | Video-Level Cross-Modal Temporal-Navigation for RGBT TrackingabstractRGBT tracking has recently garnered significant attention due to its all-weather tracking capability. Traditional RGBT tracking methods primarily concentrate on the fusion of cross-modal spatial information. However, these methods ignore contextual relationships between consecutive video frames and lack effective interactions between modalities, easily resulting in tracking drift gradually due to the accumulation of errors. To avoid this limitation, we propose a novel Video-Level Cross-Modal Temporal-Navigation method termed VCT for robust RGBT tracking, which fully leverages the complementary spatio-temporal information across modalities to improve cross-modal tracking accuracy. The VCT employs a simple, flexible, and effective video-level dual-stream architecture that accommodates video sequences of arbitrary length, enabling RGB and TIR streams to capture and synergize spatio-temporal features across frames. To achieve temporal consistency and adaptability, we design a Cross-Modal Temporal Prompt Navigator (CM-TPN) that dynamically aggregates and compresses the historical frame context to navigate predictions of subsequent frames through temporal prompts. In addition, we introduce a Modality-Specific Mixture of Adapters (MS-MoA) to promote the spatio-temporal interaction both within and between modalities, thereby dramatically adapting to appearance changes. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the four popular RGBT tracking benchmarks. Si Chen 0002, Dahan Wang, Xu-Yao Zhang, Yan Yan 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Language Decoupling with Fine-Grained Knowledge Guidance for Referring Multi-Object Tracking
Siping Zhuang, Yajun Jian, Yan Yan 0001, Hanzi Wang |
ICCV | 4 |
| 2025 | You are Your Own Best Teacher: Achieving Centralized-level Performance in Federated Learning under Heterogeneous and Long-Tailed DataabstractData heterogeneity, stemming from local non-IID data and global long-tailed distributions, is a major challenge in federated learning (FL), leading to significant performance gaps compared to centralized learning. Previous research found that poor representations and biased classifiers are the main problems and proposed neural-collapse-inspired synthetic simplex ETF to help representations be closer to neural collapse optima. However, we find that the neural-collapse-inspired methods are not strong enough to reach neural collapse and still have huge gaps to centralized training. In this paper, we rethink this issue from a self-bootstrap perspective and propose FedYoYo (You Are Your Own Best Teacher), introducing Augmented Self-bootstrap Distillation (ASD) to improve representation learning by distilling knowledge between weakly and strongly augmented local samples, without needing extra datasets or models. We further introduce Distribution-aware Logit Adjustment (DLA) to balance the self-bootstrap process and correct biased feature representations. FedYoYo nearly eliminates the performance gap, achieving centralized-level performance even under mixed heterogeneity. It enhances local representation learning, reducing model drift and improving convergence, with feature prototypes closer to neural collapse optimality. Extensive experiments show FedYoYo achieves state-of-the-art results, even surpassing centralized logit adjustment methods by 5.4\% under global long-tailed settings. Shanshan Yan, Zexi Li 0001, Chao Wu 0001, Yang Lu 0009, Yan Yan 0001, Hanzi Wang |
ICCV | 6 |
| 2025 | WarpGAN: Warping-Guided 3D GAN Inversion with Style-Based Novel View Inpaintingabstract3D GAN inversion projects a single image into the latent space of a pre-trained 3D GAN to achieve single-shot novel view synthesis, which requires visible regions with high fidelity and occluded regions with realism and multi-view consistency. However, existing methods focus on the reconstruction of visible regions, while the generation of occluded regions relies only on the generative prior of 3D GAN. As a result, the generated occluded regions often exhibit poor quality due to the information loss caused by the low bit-rate latent code. To address this, we introduce the warping-and-inpainting strategy to incorporate image inpainting into 3D GAN inversion and propose a novel 3D GAN inversion method, WarpGAN. Specifically, we first employ a 3D GAN inversion encoder to project the single-view image into a latent code that serves as the input to 3D GAN. Then, we perform warping to a novel view using the depth map generated by 3D GAN. Finally, we develop a novel SVINet, which leverages the symmetry prior and multi-view image correspondence w.r.t. the same latent code to perform inpainting of occluded regions in the warped image. Quantitative and qualitative experiments demonstrate that our method consistently outperforms several state-of-the-art methods. Kaitao Huang, Yan Yan 0001, Jing-Hao Xue, Hanzi Wang |
NeurIPS | 2 |
| 2025 | IMC-Det: Intra-Inter Modality Contrastive Learning for Video Object Detection
Qiang Qi, Zhenyu Qiu, Yan Yan 0001, Yang Lu 0009, Hanzi Wang |
Int. J. Comput. Vis. | 3 |
| 2025 | Adaptive Middle Modality Alignment Learning for Visible-Infrared Person Re-identification
Yan Yan 0001, Yang Lu 0009, Hanzi Wang |
Int. J. Comput. Vis. | 2 |
| 2025 | AMST: Object tracking based on collaborative framework with adaptive multi-strategy
Rui Xu 0028, Si Chen 0002, Yan Yan 0001, Dahan Wang, Shunzhi Zhu |
Inf. Sci. | 3 |
| 2025 | Consistency-driven feature scoring and regularization network for visible-infrared person re-identification
Xueting Chen, Yan Yan 0001, Jing-Hao Xue, Nannan Wang 0001, Hanzi Wang |
Pattern Recognit. | 2 |
| 2025 | Uncertainty-Aware Label Refinement on Hypergraphs for Personalized Federated Facial Expression RecognitionabstractMost facial expression recognition (FER) models are trained on large-scale expression data with centralized learning. Unfortunately, collecting a large amount of centralized expression data is difficult in practice due to privacy concerns of facial images. In this paper, we investigate FER under the framework of personalized federated learning, which is a valuable and practical decentralized setting for real-world applications. To this end, we develop a novel uncertainty-Aware label refineMent on hYpergraphs (AMY) method. For local training, each local model consists of a backbone, an uncertainty estimation (UE) block, and an expression classification (EC) block. In the UE block, we leverage a hypergraph to model complex high-order relationships between expression samples and incorporate these relationships into uncertainty features. A personalized uncertainty estimator is then introduced to estimate reliable uncertainty weights of samples in the local client. In the EC block, we perform label propagation on the hypergraph, obtaining high-quality refined labels for retraining an expression classifier. Based on the above, we effectively alleviate heterogeneous sample uncertainty across clients and learn a robust personalized FER model in each client. Experimental results on two challenging real-world facial expression databases show that our proposed method consistently outperforms several state-of-the-art methods. This indicates the superiority of hypergraph modeling for uncertainty estimation and label refinement on the personalized federated FER task. The source code will be released athttps://github.com/mobei1006/AMY. Hu Ding 0005, Yan Yan 0001, Yang Lu 0009, Jing-Hao Xue, Hanzi Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Edge Guided Network With Motion Enhancement for Few-Shot Action RecognitionabstractExisting state-of-the-art methods for few-shot action recognition (FSAR) achieve promising performance by spatial and temporal modeling. However, most current methods ignore the importance of edge information and motion cues, leading to inferior performance. For the few-shot task, it is important to effectively explore limited data. Additionally, effectively utilizing edge information is beneficial for exploring motion cues, and vice versa. In this paper, we propose a novel edge guided network with motion enhancement (EGME) for FSAR. To the best of our knowledge, this is the first work to utilize the edge information as guidance in the FSAR task. Our EGME contains two crucial components, including an edge information extractor (EIE) and a motion enhancement module (ME). Specifically, EIE is used to obtain edge information on video frames. Afterward, the edge information is used as guidance to fuse with the frame features. In addition, ME can adaptively capture motion-sensitive features of videos. It adopts a self-gating mechanism to highlight motion-sensitive regions in videos from a large temporal receptive field. Based on the above designed components, EGME can capture edge information and motion cues, resulting in superior recognition performance. Experimental results on four challenging benchmarks show that EGME performs favorably against recent advanced methods. Kaiwen Du, Weirong Ye, Hanyu Guo, Yan Yan 0001, Hanzi Wang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | DHLA: Dynamic Hybrid Label Assignment for End-to-End Object DetectionabstractThe recent one-to-one label assignment plays a crucial role in removing the last non-differentiable component, i.e., Non-Maximum Suppression (NMS), used in the post-processing step of the one-to-many label assignment, thus building an efficient end-to-end detection system. However, due to the limited number of foreground samples, the one-to-one label assignment often suffers from insufficient representation learning, and its performance is inferior to that of traditional detectors trained using the one-to-many label assignment. To solve these problems, we introduce a novel Dynamic Hybrid Label Assignment (DHLA) method, including a Hybrid Sample Selection (HSS) strategy and a Stage-aware Soft-label Adjustment (SSA) mechanism. In order to enhance the ability of representation learning of the one-to-one label assignment, the HSS strategy subtly integrates the one-to-many and the one-to-one label assignment rules to form a simple and effective hybrid assignment rule, where high-quality samples are selected for training according to an effective task consistency metric. Moreover, the SSA mechanism dynamically adjusts the contributions of different foreground samples at different training stages, thus effectively achieving the transition from one-to-many to one-to-one label assignment. In addition, we leverage a ranking loss function to widen the score gaps between the highest scoring position and surrounding areas for effectively removing duplicate bounding boxes. As a result, our method not only learns robust feature representations during training but also performs efficient end-to-end detection during inference. Extensive experiments demonstrate our method achieves competitive performance compared to state-of-the-art detectors on the challenging COCO and CrowdHuman datasets. Zhi-Liang Hu, Si Chen 0002, Yang Hua 0001, Dahan Wang, Shunzhi Zhu, Yan Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Augmentation Matters: A Mix-Paste Method for X-Ray Prohibited Item Detection Under Noisy AnnotationsabstractAutomatic X-ray prohibited item detection is vital for public safety. Existing deep learning-based methods all assume that the annotations of training X-ray images are correct. However, obtaining correct annotations is extremely hard if not impossible for large-scale X-ray images, where item overlapping is ubiquitous. As a result, X-ray images are easily contaminated with noisy annotations, leading to performance deterioration of existing methods. In this paper, we address the challenging problem of training a robust prohibited item detector under noisy annotations (including both category noise and bounding box noise) from a novel perspective of data augmentation, and propose an effective label-aware mixed patch paste augmentation method (Mix-Paste). Specifically, for each item patch, we mix several item patches with the same category label from different images and replace the original patch in the image with the mixed patch. In this way, the probability of containing the correct prohibited item within the generated image is increased. Meanwhile, the mixing process mimics item overlapping, enabling the model to learn the characteristics of X-ray images. Moreover, we design an item-based large-loss suppression (LLS) strategy to suppress the large losses corresponding to potentially positive predictions of additional items due to the mixing operation. We show the superiority of our method on X-ray datasets under noisy annotations. In addition, we evaluate our method on the noisy MS-COCO dataset to showcase its generalization ability. These results clearly indicate the great potential of data augmentation to handle noise annotations. The source code is released athttps://github.com/wscds/Mix-Paste. Ruikang Chen, Yan Yan 0001, Jing-Hao Xue, Yang Lu 0009, Hanzi Wang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | I2OL-Net: Intra-Inter Objectness Learning Network for Point-Supervised X-Ray Prohibited Item DetectionabstractAutomatic detection of prohibited items in X-ray images plays a crucial role in public security. However, existing methods rely heavily on labor-intensive box annotations. To address this, we investigate X-ray prohibited item detection under labor-efficient point supervision and develop an intra-inter objectness learning network (I2OL-Net). I2OL-Net consists of two key modules: an intra-modality objectness learning (intra-OL) module and an inter-modality objectness learning (inter-OL) module. The intra-OL module designs a local focus Gaussian masking block and a global random Gaussian masking block to collaboratively learn the objectness in X-ray images. Meanwhile, the inter-OL module introduces the wavelet decomposition-based adversarial learning block and the objectness block, effectively reducing the modality discrepancy between natural images and X-ray images and transferring the objectness knowledge learned from natural images with box annotations to X-ray images. Based on the above, I2OL-Net greatly alleviates the severe problem of part domination caused by large intra-class variations in X-ray images. Experimental results on four X-ray datasets show that I2OL-Net can achieve superior performance with a significant reduction of annotation cost, thus enhancing its accessibility and practicality. The source code is released athttps://github.com/houjoeng/I2OL-Net. Yan Yan 0001, Jing-Hao Xue, Hanzi Wang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Frequency Domain Nuances Mining for Visible-Infrared Person Re-IdentificationabstractThis paper focuses on the visible-infrared person re-identification (VIReID) task, which is essential for information forensics and security as it enables accurate person re-identification across low-light or nighttime conditions. The primary challenge in the VIReID task is to reduce the modality discrepancy between visible and infrared images. Current methods mainly utilize the spatial information, often neglecting the discriminative potential of frequency information. To address this issue, this paper aims to mitigate the modality discrepancy from a frequency domain perspective. Specifically, we propose a novel Frequency Domain Nuances Mining (FDNM) method, which mainly includes a Salience-guided Phase Enhancement (SPE) module and an Amplitude Nuances Mining (ANM) module, to effectively explore the cross-modality frequency domain information. These two modules are mutually beneficial to jointly explore frequency-domain visible-infrared nuances, thereby significantly reducing the modality discrepancy in the frequency domain. Additionally, we propose a Center-guided Nuances Mining (CNM) loss to ensure that the ANM module retains discriminative identity information while discovering diverse cross-modality nuances. Extensive experiments show that the proposed FDNM has significant advantages in improving the performance of VIReID. For instance, our method respectively outperforms the second-best method by 5.2% in Rank-1 accuracy and 5.8% in mAP on the SYSU-MM01 dataset under the indoor search mode. Furthermore, we also demonstrate the effectiveness and generalization of the proposed FDNM method in the challenging visible-infrared face recognition task. Hanzi Wang, Yang Lu 0009, Yan Yan 0001, Xuelong Li 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Hyperbolic Self-Paced Multi-Expert Network for Cross-Domain Few-Shot Facial Expression RecognitionabstractRecently, cross-domain few-shot facial expression recognition (CF-FER), which identifies novel compound expressions with a few images in the target domain by using the model trained only on basic expressions in the source domain, has attracted increasing attention. Generally, existing CF-FER methods leverage the multi-dataset to increase the diversity of the source domain and alleviate the discrepancy between the source and target domains. However, these methods learn feature embeddings in the Euclidean space without considering imbalanced expression categories and imbalanced sample difficulty in the multi-dataset. Such a way makes the model difficult to capture hierarchical relationships of facial expressions, resulting in inferior transferable representations. To address these issues, we propose a hyperbolic self-paced multi-expert network (HSM-Net), which contains multiple mixture-of-experts (MoE) layers located in the hyperbolic space, for CF-FER. Specifically, HSM-Net collaboratively trains multiple experts in a self-distillation manner, where each expert focuses on learning a subset of expression categories from the multi-dataset. Based on this, we introduce a hyperbolic self-paced learning (HSL) strategy that exploits sample difficulty to adaptively train the model from easy-to-hard samples, greatly reducing the influence of imbalanced expression categories and imbalanced sample difficulty. Our HSM-Net can effectively model rich hierarchical relationships of facial expressions and obtain a highly transferable feature space. Extensive experiments on both in-the-lab and in-the-wild compound expression datasets demonstrate the superiority of our proposed method over several state-of-the-art methods. Code will be released at https://github.com/cxtjl/HSM-Net. Xueting Chen, Yan Yan 0001, Jing-Hao Xue, Hanzi Wang |
IEEE Trans. Image Process. | 2 |
| 2025 | DGC-Net: Dynamic Graph Contrastive Network for Video Object DetectionabstractVideo object detection is a challenging task in computer vision since it needs to handle the object appearance degradation problem that seldom occurs in the image domain. Off-the-shelf video object detection methods typically aggregate multi-frame features at one stroke to alleviate appearance degradation. However, these existing methods do not take supervision knowledge into consideration and thus still suffer from insufficient feature aggregation, resulting in the false detection problem. In this paper, we take a different perspective on feature aggregation, and propose a dynamic graph contrastive network (DGC-Net) for video object detection, including three improvements against existing methods. First, we design a frame-level graph contrastive module to aggregate frame features, enabling our DGC-Net to fully exploit discriminative contextual feature representations to facilitate video object detection. Second, we develop a proposal-level graph contrastive module to aggregate proposal features, making our DGC-Net sufficiently learn discriminative semantic feature representations. Third, we present a graph transformer to dynamically adjust the graph structure by pruning the useless nodes and edges, which contributes to improving accuracy and efficiency as it can eliminate the geometric-semantic ambiguity and reduce the graph scale. Furthermore, inherited from the framework of DGC-Net, we develop DGC-Net Lite to perform real-time video object detection with a much faster inference speed. Extensive experiments conducted on the ImageNet VID dataset demonstrate that our DGC-Net outperforms the performance of current state-of-the-art methods. Notably, our DGC-Net obtains 86.3%/87.3% mAP when using ResNet-101/ResNeXt-101. Qiang Qi, Hanzi Wang, Yan Yan 0001, Xuelong Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Hierarchical Attention-Enhanced Correlation Refinement for Robust Visual TrackingabstractIn recent years, visual tracking has witnessed remarkable advancements with the exploration of feature extraction and correlation modeling techniques. However, inadequate robustness of either the backbone network or the correlation operation continues to plague existing trackers, leading to frustrating drift when confronted with similar distractors or cluttered backgrounds. To address this problem, we propose a hierarchical attention-enhanced correlation refinement network (HarNet) for achieving robust visual tracking. Specifically, a gated dual-view attention (GDA) module is first designed to aggregate the intra-layer attention and the inter-layer self-attention based on a fusion gate, so as to enhance hierarchical feature representations of the template. Meanwhile, a target-aware attention (TA) module introduces the template information to the inter-layer self-attention, which can highlight the target information in the search region. Moreover, a graph guided correlation (GGC) module leverages the pixel-to-local and pixel-to-global correlations to fully exploit both local-and global-spatial information between the template and the search region, and then uses the graph convolutional network (GCN) to further learn the node relationships of the correlation map for more finegrained correlations. Thus, with the above three elaborately designed modules, the HarNet is beneficial for the enhancement of feature representation and the precise localization of the target. Extensive experiments on popular visual tracking datasets (including OTB100, VOT2016, VOT2018, VOT2019, UAV123, UAV20L, GOT-10k, and LaSOT) demonstrate the superiority of our proposed method against several state-of-the-art tracking methods. Si Chen 0002, Rui Xu 0028, Yan Yan 0001, Yang Hua 0001, Dahan Wang, Shunzhi Zhu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Hierarchical Token-Aware Cross-Modality Reconstruction for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) aims to query the same pedestrian's visible (infrared) images in the gallery set from the infrared (visible) images. VI-ReID not only needs to deal with the challenging factors like pose variation and occlusion, but also requires handling the large modality discrepancy. Previous methods mainly focus on learning single-scale modality-shared features and do not effectively explore the multi-scale features of two modalities from both short-range and long-range perspectives. In order to solve these problems, this paper proposes a novel Hierarchical Token-Aware Cross-Modality Reconstruction (HTCR) network to significantly mitigate the modality discrepancy for effective VI-ReID. The HTCR network consists of two main components, i.e., Hierarchical Token-aware Fusion (HTF) and Cross-modality Feature Reconstruction (CFR). The HTF module first bidirectionally exchanges the short-range and long-range multi-scale modality-shared features with a few learnable tokens to achieve discriminative pedestrian features by making full use of the advantages of both Convolutional Neural Network (CNN) and Transformer. Moreover, the CFR module reconstructs global and local pedestrian features of one modality by using the token sequence of the other modality with multi-scale cues to further explore the relationship between the two distinct modalities and alleviate the modality discrepancy. In addition, the Modality-shared feature Reconstruction (MR) loss is leveraged to reduce the noises between the reconstructed and the target features. Experimental results indicate that the proposed HTCR can significantly improve the VI-ReID performance and outperform the state-of-the-art methods on the cross-modality SYSU-MM01, RegDB, and LLCM datasets. Si Chen 0002, Liuxiang Qiu, Dahan Wang, Wentao Zhu 0002, Yang Hua 0001, Yan Yan 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | Knowledge Distillation Meets Label Noise Learning: Ambiguity-Guided Mutual Label RefineryabstractKnowledge distillation (KD), which aims at transferring the knowledge from a complex network (a teacher) to a simpler and smaller network (a student), has received considerable attention in recent years. Typically, most existing KD methods work on well-labeled data. Unfortunately, real-world data often inevitably involve noisy labels, thus leading to performance deterioration of these methods. In this article, we study a little-explored but important issue, i.e., KD with noisy labels. To this end, we propose a novel KD method, called ambiguity-guided mutual label refinery KD (AML-KD), to train the student model in the presence of noisy labels. Specifically, based on the pretrained teacher model, a two-stage label refinery framework is innovatively introduced to refine labels gradually. In the first stage, we perform label propagation (LP) with small-loss selection guided by the teacher model, improving the learning capability of the student model. In the second stage, we perform mutual LP between the teacher and student models in a mutual-benefit way. During the label refinery, an ambiguity-aware weight estimation (AWE) module is developed to address the problem of ambiguous samples, avoiding overfitting these samples. One distinct advantage of AML-KD is that it is capable of learning a high-accuracy and low-cost student model with label noise. The experimental results on synthetic and real-world noisy datasets show the effectiveness of our AML-KD against state-of-the-art KD methods and label noise learning (LNL) methods. Code is available at https://github.com/Runqing-forMost/ AML-KD. Runqing Jiang, Yan Yan 0001, Jing-Hao Xue, Si Chen 0002, Nannan Wang 0001, Hanzi Wang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | MOOD: Leveraging Out-of-Distribution Data to Enhance Imbalanced Semi-Supervised LearningabstractThe imbalanced semi-supervised learning (SSL) has emerged as a critical research area due to the prevalence of class imbalanced and partially labeled data in real-world scenarios. As the requirement for data volume increases, naturally collected datasets inevitably contain out-of-distribution (OOD) samples. However, the performance of existing imbalanced SSL methods experiences a marked deterioration with OOD data. In this article, we propose an imbalanced SSL method called mixup-OOD (MOOD) to address this issue. The core idea is to "turn waste into treasure," exploring the potential of leveraging seemingly detrimental OOD data to expand the feature space, particularly for tail classes. Specifically, we first filter OOD data from unlabeled data, and then fuse it with labeled data to boost feature diversity for the tail classes. To avoid feature overlapping with OOD data, we develop a push-and-pull (PaP) loss to attract in-distribution (ID) instances toward respective class centroids while repelling OOD samples from them. Extensive experiments show that MOOD achieves superior performance compared with other state-of-the-art methods and exhibits robustness across data with different imbalanced ratios and OOD proportions. The source code is available at: https://github.com/xlhuang132/MOODv2. Yang Lu 0009, Xiaolin Huang, Mengke Li 0001, Yan Yan 0001, Chen Gong 0002, Hanzi Wang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | High-Order Structure Based Middle-Feature Learning for Visible-Infrared Person Re-identificationabstractVisible-infrared person re-identification (VI-ReID) aims to retrieve images of the same persons captured by visible (VIS) and infrared (IR) cameras. Existing VI-ReID methods ignore high-order structure information of features while being relatively difficult to learn a reasonable common feature space due to the large modality discrepancy between VIS and IR images. To address the above problems, we propose a novel high-order structure based middle-feature learning network (HOS-Net) for effective VI-ReID. Specifically, we first leverage a short- and long-range feature extraction (SLE) module to effectively exploit both short-range and long-range features. Then, we propose a high-order structure learning (HSL) module to successfully model the high-order relationship across different local features of each person image based on a whitened hypergraph network. This greatly alleviates model collapse and enhances feature representations. Finally, we develop a common feature space learning (CFL) module to learn a discriminative and reasonable common feature space based on middle features generated by aligning features from different modalities and ranges. In particular, a modality-range identity-center contrastive (MRIC) loss is proposed to reduce the distances between the VIS, IR, and middle features, smoothing the training process. Extensive experiments on the SYSU-MM01, RegDB, and LLCM datasets show that our HOS-Net achieves superior state-of-the-art performance. Our code is available at https://github.com/Jaulaucoeng/HOS-Net. Liuxiang Qiu, Si Chen 0002, Yan Yan 0001, Jing-Hao Xue, Dahan Wang, Shunzhi Zhu |
AAAI | 3 |
| 2024 | BlockGCN: Redefine Topology Awareness for Skeleton-Based Action RecognitionabstractGraph Convolutional Networks (GCNs) have long set the state-of-the-art in skeleton-based action recognition, leveraging their ability to unravel the complex dynamics of human joint topology through the graph's adjacency matrix. However, an inherent flaw has come to light in these cutting-edge models: they tend to optimize the adjacency matrix jointly with the model weights. This process, while seemingly efficient, causes a gradual decay of bone connectiv-ity data, resulting in a model indifferent to the very topology it sought to represent. To remedy this, we propose a two-fold strategy: (1) We introduce an innovative approach that encodes bone connectivity by harnessing the power of graph distances to describe the physical topology; we further incorporate action-specific topological representation via persistent homology analysis to depict systemic dynamics. This preserves the vital topological nuances often lost in conventional GCNs. (2) Our investigation also reveals the redundancy in existing GCNs for multi-relational modeling, which we address by proposing an efficient refinement to Graph Convolutions (GC) - the BlockGC. This signif-icantly reduces parameters while improving performance beyond original GCNs. Our full model, BlockGCN, es-tablishes new benchmarks in skeleton-based action recognition across all model categories. Its high accuracy and lightweight design, most notably on the large-scale NTU RGB+D 120 dataset, stand as strong validation of the efficacy of BlockGCN. Yuxuan Zhou 0004, Zhi-Qi Cheng, Yan Yan 0001, Qi Dai 0001, Xian-Sheng Hua 0001 |
CVPR | 4 |
| 2024 | Bi-Directional Motion Attention with Contrastive Learning for few-shot Action RecognitionabstractIn recent years, many few-shot action recognition methods have achieved competitive performance by adopting metric-based techniques. However, they suffer from two limitations: (1) Spatio-temporal relationship is modeled independently, overlooking the spatio-temporal correspondence between target objects across video frames. (2) Inter-class similarities are not well exploited in the task. As a result, their performance is significantly constrained by the presence of similar segments among different classes. In this paper, a novel BiMACL method for few-shot action recognition is presented, consisting of a Temporal Difference Spatial Attention Module (TDSAM) that uses motion attention to effectively capture the spatio-temporal correspondence between video frames, and a Contrastive Temporal-Relational CrossTransformers (CTRX) module to alleviate the adverse effects of similar subsequences of frames among distinct classes. Extensive experimental results demonstrate the superiority of our method over most methods for few-shot action recognition. Code is available at https://github.com/YWCandGHY/BiMACL. Hanyu Guo, Wanchuan Yu, Yan Yan 0001, Hanzi Wang |
ICASSP | 3 |
| 2024 | Proposal Distillation of Multi-Modal Feature Aggregation Network for Video Object DetectionabstractVideo object detection is a challenging task due to deteriorated object appearances. In order to bolster per-frame feature representations, one way is to aggregate features from relevant frames. However, relying exclusively on RGB modal for feature aggregation may limit the detection performance for lacking of motion robustness. We propose a novel proposal distillation of multi-modal feature aggregation network (PDMAN). Specially, it initially aligns the feature domain and flow domain via a lightweight flow module (LFM) and then facilities frame-level feature aggregation. Subsequently, a global-based semantic embedding module (GSEM) is designed to incorporate global semantic features into instance features and introduce a global multi-label classification loss to guide encoding with high class-wise responsiveness. Finally, to alleviate the presence of insufficient and redundant information in multi-modal instance-level feature aggregation, a proposal distilled aggregation module (PDAM) is employed. By distilling the instance set, this approach realizes a fine-grained feature aggregation, ultimately boosting the detection performance. Experimental results demonstrate that the proposed PDMAN achieves a favorable result on the most representative large-scale ImageNet VID dataset. Zhenyu Qiu, Qiang Qi, Yang Lu 0009, Yan Yan 0001, Hanzi Wang |
ICASSP | 4 |
| 2024 | Adjustable Gating Prompt Transformer for Facial Attribute Recognition with Limited Labeled Data
Qinxian Ye, Si Chen 0002, Dahan Wang, Nanfeng Jiang, Yanfei Su, Yan Yan 0001 |
ICPR (28) | 6 |
| 2024 | GLATrack: Global and Local Awareness for Open-Vocabulary Multiple Object TrackingabstractOpen-vocabulary multi-object tracking (MOT) aims to track arbitrary objects encountered in the real world beyond the training set. However, recent methods rely solely on instance-level detection and association of novel objects, which may not consider the valuable fine-grained semantic representations of the targets within key and reference frames. In this paper, we propose a Global and Local Awareness open-vocabulary MOT method (GLATrack), which learns to tackle the task of real-world MOT from both global and instance-level perspectives. Specifically, we introduce a region-aware feature enhancement module to refine global knowledge for complementing local target information, which enhances semantic representation and bridges the distribution gap between the image feature map and the pooled regional features. We propose a bidirectional semantic complementarity strategy to mitigate semantic misalignment arising from missing target information in key frames, which dynamically selects valuable information within reference frames to enrich object representation during the knowledge distillation process. Furthermore, we introduce an appearance richness measurement module to provide appropriate representations for targets with different appearances. The proposed method gains an improvement of 6.9% in TETA and 5.6% in mAP on the large-scale TAO benchmark. Yajun Jian, Yan Yan 0001, Hanzi Wang |
ACM Multimedia | 3 |
| 2024 | Wavelet-domain feature decoupling for weakly supervised multi-object tracking
Yu-Lei Li, Yan Yan 0001, Yang Lu 0009, Hanzi Wang |
Sci. China Inf. Sci. | 2 |
| 2024 | GCAT: graph calibration attention transformer for robust object tracking
Si Chen 0002, Xinxin Hu, Dahan Wang, Yan Yan 0001, Shunzhi Zhu |
Neural Comput. Appl. | 4 |
| 2024 | HASI: Hierarchical Attention-Aware Spatio-Temporal Interaction for Video-Based Person Re-IdentificationabstractVideo-based person re-identification (re-ID) aims to match the same pedestrian of video sequences across non-overlapping cameras. Video re-ID methods generally adopt frame-level feature extraction for different video frames, but they still lack effective spatio-temporal interaction, easily leading to the multi-frame misalignment problem. In this paper, we propose a Hierarchical Attention-aware Spatio-temporal Interaction (HASI) network, including an Attention-aware Temporal Interaction (ATI) module and a Hierarchical Local-spatial Enhancement (HLE) module for video-based person re-ID. In order to avoid the spatial misalignment between video frames, the ATI module employs multiple Frame-to-Frame Temporal Interaction (2FTI) blocks with the Multi-head Inter-frame Alignment Attention (MIAA) to make the current frame iteratively interact with each rest frame of a video in a positive single-cycle manner, rather than only interacting with the adjacent frame or directly building the relationship of all frames at once. This module can not only obtain the long-range non-adjacent temporal information, but also learn the pairwise frame-to-frame relationships. Moreover, the HLE module is designed to enhance the local fine-grained features from multiple Transformer layers, whilst delivering low-level information to further enrich middle-level and high-level semantic knowledge. Thus, our method can learn multi-perspective pedestrian information, including inter-frame long-range interaction information and intra-frame multi-layer global and local information. Extensive experiments demonstrate the superiority of the proposed HASI method compared with the state-of-the-art methods on the three challenging video-based re-ID datasets, i.e., MARS, iLIDS-VID, and PRID-2011. Si Chen 0002, Hui Da, Dahan Wang, Xu-Yao Zhang, Yan Yan 0001, Shunzhi Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Interpretable Heterogeneous Teacher-Student Learning Framework for Hybrid-Supervised Pulmonary Nodule DetectionabstractExisting pulmonary nodule detection methods often train models in a fully-supervised setting that requires strong labels (i.e., bounding box labels) as label information. However, manual annotation of bounding boxes in CT images is very time-consuming and labor-intensive. To alleviate the annotation burden, in this paper, we investigate pulmonary nodule detection by leveraging both strong labels and weak labels (i.e., center point labels) for training, and propose a novel hybrid-supervised pulmonary nodule detection (HND) method. The training of HND involves a heterogeneous teacher-student learning framework in two stages. In the first stage, we design a point-based consistency calibration network (PCC-Net) as a teacher, which is pre-trained to generate high-quality pseudo bounding box labels given point-augmented CT images as inputs. In the second stage, we develop an information bottleneck-guided pulmonary nodule detection network (IBD-Net) as a student to perform pulmonary nodule detection. In particular, we introduce information bottleneck to learn reliable pulmonary nodule-specific heatmaps under the guidance of PCC-Net, largely enhancing the model’s interpretability and improving the final detection performance. Based on the above designs, our method can effectively detect pulmonary nodule regions with only a limited number of bounding box labels. Experimental results on the public pulmonary nodule detection dataset LUNA16 show that our HND method achieves an excellent balance between the annotation cost and the detection performance. Guangyu Huang, Yan Yan 0001, Jing-Hao Xue, Wentao Zhu 0002, Xióngbiao Luó |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Unpaired Caricature-Visual Face Recognition via Feature Decomposition-Restoration-DecompositionabstractExisting caricature-visual face recognition methods train the models based on caricature-visual image pairs from the same identities. Unfortunately, in many real-world applications, facial caricatures and visual facial images are usually unpaired in the training set due to the difficulty of collecting facial caricatures drawn by artists. In this paper, we study caricature-visual face recognition under the practical setting that only unpaired facial caricature and visual facial images are available as training samples, and define this setting as unpaired caricature-visual face recognition. To this end, we develop a novel feature decomposition-restoration-decomposition method (FDRD), which mainly consists of a backbone network, an identity-oriented feature decomposition module, and a modality-oriented feature restoration module, to extract modality-irrelevant identity features. To effectively train FDRD in the case of limited facial caricature training samples, we develop a two-stage learning framework. In the first stage, we perform single-modality restoration, enabling the model to have the basic ability of feature decomposition and restoration for each modality. In the second stage, we perform cross-modality recognition by exchanging new modality features between the two modalities, facilitating the model to focus on the decoupling of identity features and modality features. Experimental results demonstrate that our method performs favorably against several state-of-the-art face recognition methods and cross-modality methods. Our code is available at https://github.com/Capricorn-Karma/FDRD. Yan Yan 0001, Jing-Hao Xue, Yang Hua 0001, Hanzi Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Dual-Mode Learning for Multi-Dataset X-Ray Security Image DetectionabstractWith the recent advance of deep learning, a large number of methods have been developed for prohibited item detection in X-ray security images. Generally, these methods train models on a single X-ray image dataset that may contain only limited categories of prohibited items. To detect more prohibited items, it is desirable to train a model on the multi-dataset that is constructed by combining multiple datasets. However, directly applying existing methods to the multi-dataset cannot guarantee good performance because of the large domain discrepancy between datasets and the occlusion in images. To address the above problems, we propose a novel Dual-Mode Learning Network (DML-Net) to effectively detect all the prohibited items in the multi-dataset. In particular, we develop an enhanced RetinaNet as the architecture of DML-Net, where we introduce a lattice appearance enhanced sub-net to enhance appearance representations. Such a way benefits the detection of occluded prohibited items. Based on the enhanced RetinaNet, the learning process of DML-Net involves both common mode learning (detecting the common prohibited items across datasets) and unique mode learning (detecting the unique prohibited items in each dataset). For common mode learning, we introduce an adversarial prototype alignment module to align the feature prototypes from different datasets in the domain-invariant feature space. For unique mode learning, we take advantage of feature distillation to enforce the student model to mimic the features extracted by multiple pre-trained teacher models. By tightly combining and jointly training the dual modes, our DML-Net method successfully eliminates the domain discrepancy and exhibits superior model capacity on the multi-dataset. Extensive experimental results on several combined X-ray image datasets demonstrate the effectiveness of our method against several state-of-the-art methods. Our code is available at https://github.com/vampirename/dmlnet. Fenghong Yang, Runqing Jiang, Yan Yan 0001, Jing-Hao Xue, Hanzi Wang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Relationship-Guided Knowledge Transfer for Class-Incremental Facial Expression RecognitionabstractHuman emotions contain both basic and compound facial expressions. In many practical scenarios, it is difficult to access all the compound expression categories at one time. In this paper, we investigate comprehensive facial expression recognition (FER) in the class-incremental learning paradigm, where we define well-studied and easily-accessible basic expressions as initial classes and learn new compound expressions incrementally. To alleviate the stability-plasticity dilemma in our incremental task, we propose a novel Relationship-Guided Knowledge Transfer (RGKT) method for class-incremental FER. Specifically, we develop a multi-region feature learning (MFL) module to extract fine-grained features for capturing subtle differences in expressions. Based on the MFL module, we further design a basic expression-oriented knowledge transfer (BET) module and a compound expression-oriented knowledge transfer (CET) module, by effectively exploiting the relationship across expressions. The BET module initializes the new compound expression classifiers based on expression relevance between basic and compound expressions, improving the plasticity of our model to learn new classes. The CET module transfers expression-generic knowledge learned from new compound expressions to enrich the feature set of old expressions, facilitating the stability of our model against forgetting old classes. Extensive experiments on three facial expression databases show that our method achieves superior performance in comparison with several state-of-the-art methods. Yuanling Lv, Yan Yan 0001, Jing-Hao Xue, Si Chen 0002, Hanzi Wang |
IEEE Trans. Image Process. | 2 |
| 2024 | Cross-Modal Contrastive Learning Network for Few-Shot Action RecognitionabstractFew-shot action recognition aims to recognize new unseen categories with only a few labeled samples of each class. However, it still suffers from the limitation of inadequate data, which easily leads to the overfitting and low-generalization problems. Therefore, we propose a cross-modal contrastive learning network (CCLN), consisting of an adversarial branch and a contrastive branch, to perform effective few-shot action recognition. In the adversarial branch, we elaborately design a prototypical generative adversarial network (PGAN) to obtain synthesized samples for increasing training samples, which can mitigate the data scarcity problem and thereby alleviate the overfitting problem. When the training samples are limited, the obtained visual features are usually suboptimal for video understanding as they lack discriminative information. To address this issue, in the contrastive branch, we propose a cross-modal contrastive learning module (CCLM) to obtain discriminative feature representations of samples with the help of semantic information, which can enable the network to enhance the feature learning ability at the class-level. Moreover, since videos contain crucial sequences and ordering information, thus we introduce a spatial-temporal enhancement module (SEM) to model the spatial context within video frames and the temporal context across video frames. The experimental results show that the proposed CCLN outperforms the state-of-the-art few-shot action recognition methods on four challenging benchmarks, including Kinetics, UCF101, HMDB51 and SSv2. Xiao Wang 0072, Yan Yan 0001, Hai-Miao Hu, Bo Li 0006, Hanzi Wang |
IEEE Trans. Image Process. | 2 |
| 2024 | Visual-Textual Attribute Learning for Class-Incremental Facial Expression RecognitionabstractIn this paper, we study facial expression recognition (FER) in the class-incremental learning (CIL) setting, which defines the classification of well-studied and easily-accessible basic expressions as an initial task while learning new compound expressions gradually. Motivated by the fact that compound expressions are meaningful combinations of basic expressions, we treat basic expressions as attributes (i.e., semantic descriptors), and thus compound expressions are represented in terms of attributes. To this end, we propose a novel visual-textual attribute learning network (VTA-Net), mainly consisting of a textual-guided visual module (TVM) and a textual compositional module (TCM), for class-incremental FER. Specifically, TVM extracts textual-aware visual features and classifies expressions by incorporating the textual information into visual attribute learning. Meanwhile, TCM generates visual-aware textual features and predicts expressions by exploiting the dependency between textual attributes and category names of old and new expressions based on a textual compositional graph. In particular, a visual-textual distillation loss is introduced to calibrate TVM and TCM during incremental learning. Finally, the outputs from TVM and TCM are fused to make a final prediction. On the one hand, at each incremental task, the representations of visual attributes are enhanced since visual attributes are shared across old and new expressions. This increases the stability of our method. On the other hand, the textual modality, which involves rich prior knowledge of the relevance between expressions, facilitates our model to identify subtle visual distinctions between compound expressions, improving the plasticity of our method. Experimental results on both in-the-lab and in-the-wild facial expression databases show the superiority of our method against several state-of-the-art methods for class-incremental FER. Yuanling Lv, Guangyu Huang, Yan Yan 0001, Jing-Hao Xue, Si Chen 0002, Hanzi Wang |
IEEE Trans. Multim. | 3 |
| 2024 | Class-Aware Dual-Supervised Aggregation Network for Video Object DetectionabstractVideo object detection has attracted increasing attention in recent years. Although great success has been achieved by off-the-shelf video object detection methods through delicately designing various types of feature aggregation, they overlook the class-aware supervision and thus still suffer from the problem of classification incapability, which means the classification between objects with deteriorated or similar appearances is error-prone. In this article, we propose a novel class-aware dual-supervised aggregation network (CDANet) for video object detection, including three substantial improvements to effectively alleviate the classification incapability problem of previous methods. First, we develop a class-aware cross-modality distillation supervision that transfers the semantic knowledge of label data to the features of video data, effectively enhancing the semantic representations of features. Second, we design a graph-guided feature aggregation module that effectively models the structural relations between features by leveraging the dynamic residual graph convolutional network, enabling our CDANet to perform more effective feature aggregation in the temporal domain. Third, we present a class-aware proposal contrastive supervision to maximize the intra-class agreement and inter-class disagreement, which is conducive to improving the semantic discriminability of features. The class-aware dual supervision and feature aggregation are tightly tied into a unified end-to-end framework to make our CDANet fully exploit class-specific semantic knowledge and inter-frame temporal dependencies to enhance object appearance representations, which facilitates the classification of detected objects. We conduct experiments on the challenging ImageNet VID dataset, and the results demonstrate the superiority of our CDANet against state-of-the-art methods. More remarkably, our CDANet achieves 85.4% mAP with ResNet-101 or 86.5% mAP with ResNeXt-101. Qiang Qi, Yan Yan 0001, Hanzi Wang |
IEEE Trans. Multim. | 2 |
| 2024 | When Sparse Neural Network Meets Label Noise Learning: A Multistage Learning FrameworkabstractRecent methods in network pruning have indicated that a dense neural network involves a sparse subnetwork (called a winning ticket), which can achieve similar test accuracy to its dense counterpart with much fewer network parameters. Generally, these methods search for the winning tickets on well-labeled data. Unfortunately, in many real-world applications, the training data are unavoidably contaminated with noisy labels, thereby leading to performance deterioration of these methods. To address the above-mentioned problem, we propose a novel two-stream sample selection network (TS3-Net), which consists of a sparse subnetwork and a dense subnetwork, to effectively identify the winning ticket with noisy labels. The training of TS3-Net contains an iterative procedure that switches between training both subnetworks and pruning the smallest magnitude weights of the sparse subnetwork. In particular, we develop a multistage learning framework including a warm-up stage, a semisupervised alternate learning stage, and a label refinement stage, to progressively train the two subnetworks. In this way, the classification capability of the sparse subnetwork can be gradually improved at a high sparsity level. Extensive experimental results on both synthetic and real-world noisy datasets (including MNIST, CIFAR-10, CIFAR-100, ANIMAL-10N, Clothing1M, and WebVision) demonstrate that our proposed method achieves state-of-the-art performance with very small memory consumption for label noise learning. Code is available at https://github.com/Runqing-forMost/TS3-Net/tree/master. Runqing Jiang, Yan Yan 0001, Jing-Hao Xue, Hanzi Wang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | MRCN: A Novel Modality Restitution and Compensation Network for Visible-Infrared Person Re-identificationabstractVisible-infrared person re-identification (VI-ReID), which aims to search identities across different spectra, is a challenging task due to large cross-modality discrepancy between visible and infrared images. The key to reduce the discrepancy is to filter out identity-irrelevant interference and effectively learn modality-invariant person representations. In this paper, we propose a novel Modality Restitution and Compensation Network (MRCN) to narrow the gap between the two modalities. Specifically, we first reduce the modality discrepancy by using two Instance Normalization (IN) layers. Next, to reduce the influence of IN layers on removing discriminative information and to reduce modality differences, we propose a Modality Restitution Module (MRM) and a Modality Compensation Module (MCM) to respectively distill modality-irrelevant and modality-relevant features from the removed information. Then, the modality-irrelevant features are used to restitute to the normalized visible and infrared features, while the modality-relevant features are used to compensate for the features of the other modality. Furthermore, to better disentangle the modality-relevant features and the modality-irrelevant features, we propose a novel Center-Quadruplet Causal (CQC) loss to encourage the network to effectively learn the modality-relevant features and the modality-irrelevant features. Extensive experiments are conducted to validate the superiority of our method on the challenging SYSU-MM01 and RegDB datasets. More remarkably, our method achieves 95.1% in terms of Rank-1 and 89.2% in terms of mAP on the RegDB dataset. Yan Yan 0001, Jie Li 0001, Hanzi Wang |
AAAI | 2 |
| 2023 | A Dual-Path Transformer Network for Scene Text DetectionabstractThe prosperity of deep learning contributes to the rapid progress of scene text detection. Among all the methods, segmentation-based methods have drawn extensive attention due to their superiority in detecting text instances of arbitrary shapes and extreme aspect ratios. However, the bottom-up methods are limited to the performance of their segmentation models. In this paper, we propose DPTNet (Dual-Path Transformer Network), a simple yet effective network to utilize both global and local information for the scene text detection task. Moreover, we propose a parallel design that integrates the convolutional network with a powerful self-attention mechanism to provide complementary clues. In addition, a bi-directional interaction module across two paths is developed to provide complementary clues along the channel and spatial dimensions. Our DPTNet achieves state-of-the-art results on several standard benchmarks in terms of both detection accuracy and speed. Yan Yan 0001, Hanzi Wang |
ICASSP | 2 |
| 2023 | DF-Net: Diversity-Focused Network for Video Object DetectionabstractVideo object detection is a challenging task due to deteriorated object appearances. To enhance per-frame features, one way is to aggregate features from several support frames. However, proposals generated by the region proposal network may not be precise and diverse due to the fixed anchors, limiting the detection performance. We propose a novel architecture called Diversity-Focused Network (DF-Net), which consists of three modules: 1) An affine transform module (ATM), which is proposed to model the deblurring process and fuse the feature maps of different receptive fields by a multi-level attention block; 2) A label assignment module (LAM), which is proposed to assign the labels to the proposals used in a fine-grained aggregation manner; 3) A regression-guided diffusion module (RGDM), which is proposed to obtain the features of diversity and higher quality. Experiments show that DF-Net achieves favorable results on the most representative large-scale ImageNet VID dataset. Remarkably, the DF-Net achieves 84.8% mAP with ResNet-101 without post-processing steps. Zhenyu Qiu, Qiang Qi, Yan Yan 0001, Hanzi Wang |
ICIP | 3 |
| 2023 | Joint Relation Modeling and Feature Learning for Class-Incremental Facial Expression Recognition
Yuanling Lv, Yan Yan 0001, Hanzi Wang |
PRCV (5) | 2 |
| 2023 | Spatio-Temporal Self-supervision for Few-Shot Action Recognition
Wanchuan Yu, Hanyu Guo, Yan Yan 0001, Jie Li 0001, Hanzi Wang |
PRCV (1) | 3 |
| 2023 | SPL-Net: Spatial-Semantic Patch Learning Network for Facial Attribute Recognition with Limited Labeled Data
Yan Yan 0001, Ying Shu, Si Chen 0002, Jing-Hao Xue, Chunhua Shen, Hanzi Wang |
Int. J. Comput. Vis. | 1 |
| 2023 | Learning an attention-aware parallel sharing network for facial attribute recognition
Si Chen 0002, Xinyu Lai, Yan Yan 0001, Dahan Wang, Shunzhi Zhu |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | MTNet: Mutual tri-training network for unsupervised domain adaptation on person re-identification
Si Chen 0002, Liuxiang Qiu, Zimin Tian, Yan Yan 0001, Dahan Wang, Shunzhi Zhu |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | HC-GCN: hierarchical contrastive graph convolutional network for unsupervised domain adaptation on person re-identification
Si Chen 0002, Bolun Xu, Yan Yan 0001, Xia Du, Weiwei Zhuang, Yun Wu 0001 |
Multim. Syst. | 4 |
| 2023 | Identity-Aware Contrastive Knowledge Distillation for Facial Attribute RecognitionabstractFacial attribute recognition (FAR) is an important and yet challenging multi-label learning task in computer vision. Existing FAR methods have achieved promising performance with the development of deep learning. However, they usually suffer from prohibitive computational and memory costs. In this paper, we propose an identity-aware contrastive knowledge distillation method, termed ICKD, to compress the FAR model. A nonlinear weight-sharing mapping (NWSM) mechanism is firstly designed to avoid the difficulty of directly matching features of the teacher and student networks due to the lower representation ability of the student network. Furthermore, an identity-aware contrastive distillation (ICD) loss is employed to guide the student network to effectively learn the mutual relations between samples with multiple attributes. In addition, an adjustable ladder distillation (ALD) loss is developed to automatically adjust the importance of different distillation points with the progress of training. Extensive experiments demonstrate that our method can significantly improve the performance of student networks and outperforms the existing FAR methods on the public challenging datasets. Si Chen 0002, Xueyan Zhu, Yan Yan 0001, Shunzhi Zhu, Shaozi Li, Dahan Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | TCNet: A Novel Triple-Cooperative Network for Video Object DetectionabstractVideo object detection aims at accurately localizing the objects in videos and correctly recognizing their categories. Off-the-shelf video object detection methods have made some progress in recent years but they still suffer from the problems of inaccurate object localization, incorrect object recognition or insufficient relation learning, resulting in limited detection performance. In this paper, we propose a novel triple-cooperative network (TCNet) for high-performance video object detection, with three substantial improvements to ameliorate the problems of existing methods. First, we develop a context-aware proposal refinement module to generate high-quality proposals, enabling our TCNet to achieve more accurate object localization. Second, we present a similarity-aware semantic distillation module that innovatively leverages the semantic knowledge of class labels as additional supervisory signals to enhance the object recognition ability of our TCNet. Third, we design a structure-aware relation learning module to effectively model the structural relations between features with an adaptive-pruning residual graph convolutional network, making our TCNet perform more effective feature aggregation. We conduct extensive experiments on the challenging ImageNet VID dataset and the experimental results demonstrate that our TCNet outperforms current state-of-the-art methods. More remarkably, our TCNet achieves 85.2% mAP and 86.3% mAP with ResNet-101 and ResNeXt-101, respectively. Qiang Qi, Tianxiang Hou, Yan Yan 0001, Yang Lu 0009, Hanzi Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Drop Loss for Person Attribute Recognition With Imbalanced Noisy-Labeled SamplesabstractPerson attribute recognition (PAR) aims to simultaneously predict multiple attributes of a person. Existing deep learning-based PAR methods have achieved impressive performance. Unfortunately, these methods usually ignore the fact that different attributes have an imbalance in the number of noisy-labeled samples in the PAR training datasets, thus leading to suboptimal performance. To address the above problem of imbalanced noisy-labeled samples, we propose a novel and effective loss called drop loss for PAR. In the drop loss, the attributes are treated differently in an easy-to-hard way. In particular, the noisy-labeled candidates, which are identified according to their gradient norms, are dropped with a higher drop rate for the harder attribute. Such a manner adaptively alleviates the adverse effect of imbalanced noisy-labeled samples on model learning. To illustrate the effectiveness of the proposed loss, we train a simple ResNet-50 model based on the drop loss and term it DropNet. Experimental results on two representative PAR tasks (including facial attribute recognition and pedestrian attribute recognition) demonstrate that the proposed DropNet achieves comparable or better performance in terms of both balanced accuracy and classification accuracy over several state-of-the-art PAR methods. Yan Yan 0001, Youze Xu, Jing-Hao Xue, Yang Lu 0009, Hanzi Wang, Wentao Zhu 0002 |
IEEE Trans. Cybern. | 1 |
| 2023 | DGRNet: A Dual-Level Graph Relation Network for Video Object DetectionabstractVideo object detection is a fundamental and important task in computer vision. One mainstay solution for this task is to aggregate features from different frames to enhance the detection on the current frame. Off-the-shelf feature aggregation paradigms for video object detection typically rely on inferring feature-to-feature (Fea2Fea) relations. However, most existing methods are unable to stably estimate Fea2Fea relations due to the appearance deterioration caused by object occlusion, motion blur or rare poses, resulting in limited detection performance. In this paper, we study Fea2Fea relations from a new perspective, and propose a novel dual-level graph relation network (DGRNet) for high-performance video object detection. Different from previous methods, our DGRNet innovatively leverages the residual graph convolutional network to simultaneously model Fea2Fea relations at two different levels including frame level and proposal level, which facilitates performing better feature aggregation in the temporal domain. To prune unreliable edge connections in the graph, we introduce a node topology affinity measure to adaptively evolve the graph structure by mining the local topological information of pairwise nodes. To the best of our knowledge, our DGRNet is the first video object detection method that leverages dual-level graph relations to guide feature aggregation. We conduct experiments on the ImageNet VID dataset and the results demonstrate the superiority of our DGRNet against state-of-the-art methods. Especially, our DGRNet achieves 85.0% mAP and 86.2% mAP with ResNet-101 and ResNeXt-101, respectively. Qiang Qi, Tianxiang Hou, Yang Lu 0009, Yan Yan 0001, Hanzi Wang |
IEEE Trans. Image Process. | 4 |
| 2022 | When Facial Expression Recognition Meets Few-Shot Learning: A Joint and Alternate Learning FrameworkabstractHuman emotions involve basic and compound facial expressions. However, current research on facial expression recognition (FER) mainly focuses on basic expressions, and thus fails to address the diversity of human emotions in practical scenarios. Meanwhile, existing work on compound FER relies heavily on abundant labeled compound expression training data, which are often laboriously collected under the professional instruction of psychology. In this paper, we study compound FER in the cross-domain few-shot learning setting, where only a few images of novel classes from the target domain are required as a reference. In particular, we aim to identify unseen compound expressions with the model trained on easily accessible basic expression datasets. To alleviate the problem of limited base classes in our FER task, we propose a novel Emotion Guided Similarity Network (EGS-Net), consisting of an emotion branch and a similarity branch, based on a two-stage learning framework. Specifically, in the first stage, the similarity branch is jointly trained with the emotion branch in a multi-task fashion. With the regularization of the emotion branch, we prevent the similarity branch from overfitting to sampled base classes that are highly overlapped across different episodes. In the second stage, the emotion branch and the similarity branch play a “two-student game” to alternately learn from each other, thereby further improving the inference ability of the similarity branch on unseen compound expressions. Experimental results on both in-the-lab and in-the-wild compound expression datasets demonstrate the superiority of our proposed method against several state-of-the-art methods. Xinyi Zou, Yan Yan 0001, Jing-Hao Xue, Si Chen 0002, Hanzi Wang |
AAAI | 2 |
| 2022 | Learn-to-Decompose: Cascaded Decomposition Network for Cross-Domain Few-Shot Facial Expression Recognition
Xinyi Zou, Yan Yan 0001, Jing-Hao Xue, Si Chen 0002, Hanzi Wang |
ECCV (19) | 2 |
| 2022 | Learning Correlation for Online Multiple Object TrackingabstractExisting multiple object tracking methods usually strengthen data association by discriminative identity embeddings. However, many works treat object detection and association as two individual tasks, thus gaining limited benefits. In this paper, we follow the joint detection and tracking paradigm to learn correlation for online multiple object tracking. The proposed method, named LCTrack, links the two tasks by an attention mechanism. Specifically, for robust feature representations, we introduce an identity-aware attention module to extract reliable identity embeddings and model their correlation between two consecutive frames. Furthermore, for effective correlation learning, we design a target-aware loss to train the identity embedding extraction, which is well compatible with the detection task. Therefore, LCTrack can boost the position prediction and data association by the enhanced feature representation. Experimental results on the MOTChallenge benchmarks demonstrate the effectiveness and favorable performance of the proposed LCTrack in comparison with state-of-the-art methods. Chihui Zhuang, Haihui Ye, Yan Yan 0001, Hanzi Wang |
ICASSP | 4 |
| 2022 | Multi-Focus Guided Semantic Aggregation for Video Object DetectionabstractFor the task of video object detection, it is useful to aggregate semantic information from supporting frames. However, existing methods only focus on the current frame during the semantic aggregation, called Single-Focus methods. They neglect semantic information among supporting frames and deteriorate overall performance. In this work, we propose a method called Multi-Focus guided Semantic Aggregation (MFSA) for video object detection. We introduce a novel Relation Propagation Module (RPM) to capture and propagate proposal-to-proposal semantic dependencies. Moreover, we propose a simple yet effective Multi-Focus strategy to leverage captured dependencies to guide feature enhancement at a batch level. Aided by this strategy, our method can greatly improve aggregation efficiency of Single-Focus methods and enhance the accuracy of a per-frame detector significantly with negligible computing overhead. We perform extensive experiments on the ImageNet VID dataset. The results show that MFSA achieves excellent performance and a superior speed-accuracy tradeoff among the competing methods. Haihui Ye, Guangge Wang, Yang Lu 0009, Yan Yan 0001, Hanzi Wang |
ICASSP | 4 |
| 2022 | Bounding Box Distribution Learning and Center Point Calibration for Robust Visual TrackingabstractVisual tracking aims at both robust target classification and accurate localization. However, the reliability of the target bounding box and classification score are not properly addressed by most existing trackers, resulting in inaccurate tracking performance. In this paper, we propose to learn bounding box distribution in training and calibrate the center point response in inference for robust online tracking. Specifically, we propose a simple yet effective bounding box distribution learning (BDL) module to model the target bounding box distribution and enhance the localization ability of our network. Furthermore, we propose a center point calibration (CPC) module to calibrate the origin classification score with the predicted localization uncertainty and generate an accurate target center point. The proposed tracking method is referred to as DLPC. The experimental results on four challenging datasets (i.e., OTB100, VOT2019, LaSOT, and TrackingNet) show that DLPC performs favorably against several state-of-the-art trackers while running in real-time at 60 fps. Chihui Zhuang, Yan Yan 0001, Yang Lu 0009, Hanzi Wang |
ICASSP | 3 |
| 2022 | MSFL-Net: Multi-Semantic Feature Learning Network for Occluded Person Re-IdentificationabstractRecently, occluded person re-identification (Re-ID) has received significant interest due to its widespread real-world applications. However, most existing occluded person Re-ID methods ignore semantic granularities that indicate different levels of occluded information of the human body, leading to sub-optimal performance. To address this, we propose a Multi-Semantic Feature Learning Network (MSFL-Net) for occluded person Re-ID. Specifically, MSFL-Net involves a backbone network and a Multi-branch Feature Learning sub-network (MFL). MFL consists of two local-global branches and a global branch to learn multisemantic features in a multi-branch deep network architecture. In each local-global branch, we design a local subbranch and a semantic-guided global sub-branch to extract discriminative features at a certain level of feature granularity and semantic granularity. In the global branch, we learn global features at the largest level of semantic granularity. In particular, a patch contrastive loss is developed to explicitly encourage the semantic feature maps to capture the information from specific body parts. By extracting multi-semantic features, our method is effective in dealing with person Re-ID at different occlusion levels. Experimental results on an occluded person Re-ID dataset (Occluded-REID) and two partial person Re-ID datasets (Partial-iLIDS and Partial-REID) show the superiority of our method against state-of-the-art person Re-ID methods. Guangyu Huang, Yan Yan 0001, Si Chen 0002, Wentao Zhu 0002, Hanzi Wang |
IJCB | 3 |
| 2022 | MDNet: Motion Distinction Network for Effective Action RecognitionabstractMotion information is critical for action recognition. Most existing methods perform motion enhancement only from the channel or spatiotemporal dimension. They may fail to achieve fine-grained motion modeling, leading to suboptimal performance. In this paper, we propose a novel motion distinction network (MDNet) to address this challenging problem. Specifically, we first propose a channel-wise motion enhancement (CME) module, which aims to emphasize motion-related channels by leveraging a channel-wise gating mechanism. Then, we propose a cascaded spatiotemporal enhancement (CSTE) module to enhance motion features along the spatiotemporal dimension. Moreover, we design a multi-attention fusion strategy to further refine the enhanced motion features in a moderate manner, enabling the network to focus on discriminative motion regions. The proposed two modules and fusion strategy are complementary for fine-grained motion enhancement. Extensive experiments are conducted on the challenging Something-Something V1 and Kinetics-400 datasets to show the effectiveness of the proposed MDNet. Rongrong Ji, Weirong Ye, Xiao Wang 0072, Yan Yan 0001, Hanzi Wang |
ICIP | 4 |
| 2022 | Guided Sampling Based Feature Aggregation for Video Object DetectionabstractVideo object detection is a challenging task due to the presence of appearance deterioration in video frames. Recently, feature aggregation based methods which aggregate context information from object proposals in different frames to improve the performance, have dominated the task. However, much invalid information may be introduced during feature aggregation since frames and proposals are usually selected at random. In this paper, we propose a guided sampling based feature aggregation network (GSFA) to perform more effective feature aggregation. Specifically, we introduce a frame-level sampling module and a proposal-level sampling module to sample informative frames and proposals from a video sequence adaptively. As a result, the proposed GSFA can effectively aggregate context information from the semantically rich frames and proposals to boost the performance. Experimental results on the ImageNet VID dataset show the proposed GSFA achieves the state-of-the-art performance of 84.8% mAP with ResNet-101 and 85.8% mAP with ResNeXt-101. Haosheng Chen 0001, Yan Yan 0001, Yang Lu 0009, Hanzi Wang |
ICIP | 3 |
| 2022 | Dualfeat: Dual Feature Aggregation for Video Object DetectionabstractVideo object detection aims to detect and track each object in a given video. However, due to the problem of appearance deterioration in the video, it is still challenging to obtain good results when we apply traditional image object detection methods to videos. In this paper, we propose a new feature aggregation method, called Dual Feature Aggregation (DualFeat) for video object detection. By effectively combining the temporal and spatial attention mechanisms, we make full use of the temporal and spatial information in videos. Meanwhile, we leverage a real-time tracker to track detected objects in video frames, where features are aggregated again with previously obtained features. Such a way helps to obtain more comprehensive and richer features, greatly improving the accuracy of video object detection. We perform experiments on the ILSVRC2017 dataset, and the experimental results also verify the effectiveness of our method. Kaiwen Du, Yan Yan 0001, Hanzi Wang |
ICIP | 3 |
| 2022 | An End-to-End Scene Text Detector with Dynamic AttentionabstractDetecting the arbitrarily oriented text in natural images is a challenging task in multimedia due to variations in text curvatures, orientations, and aspect ratios of natural scenes. Most previous scene text detectors often fail to locate the text instances which have a peculiar shape (an extreme aspect ratio) precisely. In this paper, we propose a dynamic end-to-end framework (DEF) which includes a convolution-based dynamic encoder (CDE) with various attention types to generate a deformable and dynamic view for multi-oriented text instances and curve ones. Different from previous methods that apply time-consuming post-processing steps like NMS, our method uses a Transformer-based decoder (TD) with a bipartite matching loss to model the relationship of corresponding queries and ground truths. As a result, by leveraging such a well-designed architecture, the receptive field will not be limited to a fixed shape, and a combination of global attention and local features provides a better representation for texts in natural scenes. We conduct extensive experiments qualitatively and quantitatively on several popular datasets. Experimental results show that the proposed method achieves superior performance compared with several state-of-the-art scene text detectors. Yan Yan 0001, Hanzi Wang |
MMAsia | 2 |
| 2022 | Adaptive Deep Disturbance-Disentangled Learning for Facial Expression Recognition
Delian Ruan, Rongyun Mo, Yan Yan 0001, Si Chen 0002, Jing-Hao Xue, Hanzi Wang |
Int. J. Comput. Vis. | 3 |
| 2022 | Learning meta-adversarial features via multi-stage adaptation network for robust visual object tracking
Si Chen 0002, Yan Yan 0001, Dahan Wang, Shunzhi Zhu |
Neurocomputing | 4 |
| 2022 | Dimension-aware attention for efficient mobile networks
Rongyun Mo, Shenqi Lai, Yan Yan 0001, Zhenhua Chai, Xiaolin Wei |
Pattern Recognit. | 3 |
| 2022 | TRL: Transformer based refinement learning for hybrid-supervised semantic segmentation
Pengfei Fang, Yan Yan 0001, Yang Lu 0009, Hanzi Wang |
Pattern Recognit. Lett. | 3 |
| 2022 | Deep Multi-Task Multi-Label CNN for Effective Facial Attribute ClassificationabstractFacial Attribute Classification (FAC) has attracted increasing attention in computer vision and pattern recognition. However, state-of-the-art FAC methods perform face detection/alignment and FAC independently. The inherent dependencies between these tasks are not fully exploited. In addition, most methods predict all facial attributes using the same CNN network architecture, which ignores the different learning complexities of facial attributes. To address the above problems, we propose a novel deep multi-task multi-label CNN, termed DMM-CNN, for effective FAC. Specifically, DMM-CNN jointly optimizes two closely-related tasks (i.e., facial landmark detection and FAC) to improve the performance of FAC by taking advantage of multi-task learning. To deal with the diverse learning complexities of facial attributes, we divide the attributes into two groups: objective attributes and subjective attributes. Two different network architectures are respectively designed to extract features for two groups of attributes, and a novel dynamic weighting scheme is proposed to automatically assign the loss weight to each facial attribute during training. Furthermore, an adaptive thresholding strategy is developed to effectively alleviate the problem of class imbalance for multi-label learning. Experimental results on the challenging CelebA and LFWA datasets show the superiority of the proposed DMM-CNN method compared with several state-of-the-art FAC methods. Longbiao Mao, Yan Yan 0001, Jing-Hao Xue, Hanzi Wang |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Stage-Aware Feature Alignment Network for Real-Time Semantic Segmentation of Street ScenesabstractOver the past few years, deep convolutional neural network-based methods have made great progress in semantic segmentation of street scenes. Some recent methods align feature maps to alleviate the semantic gap between them and achieve high segmentation accuracy. However, they usually adopt the feature alignment modules with the same network configuration in the decoder and thus ignore the different roles of stages of the decoder during feature aggregation, leading to a complex decoder structure. Such a manner greatly affects the inference speed. In this paper, we present a novel Stage-aware Feature Alignment Network (SFANet) based on the encoder-decoder structure for real-time semantic segmentation of street scenes. Specifically, a Stage-aware Feature Alignment module (SFA) is proposed to align and aggregate two adjacent levels of feature maps effectively. In the SFA, by taking into account the unique role of each stage in the decoder, a novel stage-aware Feature Enhancement Block (FEB) is designed to enhance spatial details and contextual information of feature maps from the encoder. In this way, we are able to address the misalignment problem with a very simple and efficient multi-branch decoder structure. Moreover, an auxiliary training strategy is developed to explicitly alleviate the multi-scale object problem without bringing additional computational costs during the inference phase. Experimental results show that the proposed SFANet exhibits a good balance between accuracy and speed for real-time semantic segmentation of street scenes. In particular, based on ResNet-18, SFANet respectively obtains 78.1% and 74.7% mean of class-wise Intersection-over-Union (mIoU) at inference speeds of 37 FPS and 96 FPS on the challenging Cityscapes and CamVid test datasets by using only a single GTX 1080Ti GPU. Xi Weng, Yan Yan 0001, Si Chen 0002, Jing-Hao Xue, Hanzi Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Co-Clustering on Bipartite Graphs for Robust Model FittingabstractRecently, graph-based methods have been widely applied to model fitting. However, in these methods, association information is invariably lost when data points and model hypotheses are mapped to the graph domain. In this paper, we propose a novel model fitting method based on co-clustering on bipartite graphs (CBG) to estimate multiple model instances in data contaminated with outliers and noise. Model fitting is reformulated as a bipartite graph partition behavior. Specifically, we use a bipartite graph reduction technique to eliminate some insignificant vertices (outliers and invalid model hypotheses), thereby improving the reliability of the constructed bipartite graph and reducing the computational complexity. We then use a co-clustering algorithm to learn a structured optimal bipartite graph with exact connected components for partitioning that can directly estimate the model instances (i.e., post-processing steps are not required). The proposed method fully utilizes the duality of data points and model hypotheses on bipartite graphs, leading to superior fitting performance. Exhaustive experiments show that the proposed CBG method performs favorably when compared with several state-of-the-art fitting methods. Shuyuan Lin, Hailing Luo, Yan Yan 0001, Guobao Xiao, Hanzi Wang |
IEEE Trans. Image Process. | 3 |
| 2022 | Deep Correlation Filter Tracking With Shepherded Instance-Aware ProposalsabstractVisual tracking is a core component of intelligent transportation systems and it is crucial to reduce or avoid traffic accidents. Recently, deep correlation filter (DCF) based trackers have exhibited good tracking performance. However, existing DCF based trackers are still ineffective to cope with large scale variations and severe distortions (e.g., heavy occlusions, significant deformations, large rotations, etc.), leading to the inferior performance. To address these issues, we develop a novel DeepCFIAP++ tracker, which incorporates effective shepherded instance-aware proposals into DCFs. DeepCFIAP++ can not only estimate the target scale at every frame but also re-detect the target in the case of severe distortions. Firstly, we propose to exploit both color and edge cues to generate complementary detection proposals to effectively handle various challenging scenarios. Then, we propose to utilize multi-layer target-specific deep features to rank the generated detection proposals and choose the instance-aware proposals, which will result in more robust tracking performance. Finally, we propose to use the DCFs to shepherd the instance-aware proposals toward their best locations, which will result in more accurate tracking results. Experimental results on five challenging datasets (i.e., OTB2013, OTB2015, VOT2016, VOT2017 and UAV20L) demonstrate that DeepCFIAP++ performs competitively with several other state-of-the-art DCF based trackers. Qiangqiang Wu, Yan Yan 0001, Hanzi Wang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | FastVOD-Net: A Real-Time and High-Accuracy Video Object DetectorabstractVideo object detection is a tough task due to the severe appearance degradation caused by rapid motion, sudden occlusion or rare poses. The great challenge facing video object detection is the simultaneous requirements on both accuracy and speed because the pursuit of one aspect usually causes significant expense to the other. Most existing methods mainly focus on improving detection accuracy with little attention to computationally efficient solutions, and thus they are impractical for many real-world applications. This motivates us to develop a real-time and high-accuracy video object detection method. In this paper, we propose a novel video object detector, called FastVOD-Net, which can yield highly accurate detection results at real-time speed. Specifically, we first develop a temporally-cascaded deformable alignment (TCDA) module to model the object displacements induced by video motion. Then, we introduce another two modules, namely spatially-refined temporal aggregation (SRTA) and attention-guided semantic distillation (AGSD), to improve the appearance feature of the currently processed frame and enhance the semantic representation of non-keyframes, respectively. For keyframe scheduling, we design an adaptive keyframe selection scheduler (AKSS) to adjust the keyframe interval online, making the keyframe usage more rational. On one hand, the characteristics of our FastVOD-Net enable it to sparsely perform expensive feature extraction, which significantly reduces the computational cost and thus guarantees real-time speed. On the other hand, the collaboration of the above tightly-coupled modules and adaptive keyframe scheduler makes FastVOD-Net fully exploit inter-frame temporal dependencies and thus guarantees high accuracy. Experiments on the ImageNet VID dataset show that our FastVOD-Net achieves 79.3% mAP at 29.6 fps or 81.2% mAP at 23.0 fps on an Nvidia RTX 2080 Ti GPU, which is the state-of-the-art performance in real time. Qiang Qi, Xiao Wang 0072, Tianxiang Hou, Yan Yan 0001, Hanzi Wang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Deep Multi-Branch Aggregation Network for Real-Time Semantic Segmentation in Street ScenesabstractReal-time semantic segmentation, which aims to achieve high segmentation accuracy at real-time inference speed, has received substantial attention over the past few years. However, many state-of-the-art real-time semantic segmentation methods tend to sacrifice some spatial details or contextual information for fast inference, thus leading to degradation in segmentation quality. In this paper, we propose a novel Deep Multi-branch Aggregation Network (called DMA-Net) based on the encoder-decoder structure to perform real-time semantic segmentation in street scenes. Specifically, we first adopt ResNet-18 as the encoder to efficiently generate various levels of feature maps from different stages of convolutions. Then, we develop a Multi-branch Aggregation Network (MAN) as the decoder to effectively aggregate different levels of feature maps and capture the multi-scale information. In MAN, a lattice enhanced residual block is designed to enhance feature representations of the network by taking advantage of the lattice structure. Meanwhile, a feature transformation block is introduced to explicitly transform the feature map from the neighboring branch before feature aggregation. Moreover, a global context block is used to exploit the global contextual information. These key components are tightly combined and jointly optimized in a unified network. Extensive experimental results on the challenging Cityscapes and CamVid datasets demonstrate that our proposed DMA-Net respectively obtains 77.0% and 73.6% mean Intersection over Union (mIoU) at the inference speed of 46.7 FPS and 119.8 FPS by only using a single NVIDIA GTX 1080Ti GPU. This shows that DMA-Net provides a good tradeoff between segmentation quality and speed for semantic segmentation in street scenes. Xi Weng, Yan Yan 0001, Genshun Dong, Hanzi Wang, Ji Zhang 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Feature Decomposition and Reconstruction Learning for Effective Facial Expression RecognitionabstractIn this paper, we propose a novel Feature Decomposition and Reconstruction Learning (FDRL) method for effective facial expression recognition. We view the expression information as the combination of the shared information (expression similarities) across different expressions and the unique information (expression-specific variations) for each expression. More specifically, FDRL mainly consists of two crucial networks: a Feature Decomposition Network (FDN) and a Feature Reconstruction Network (FRN). In particular, FDN first decomposes the basic features extracted from a backbone network into a set of facial action-aware latent features to model expression similarities. Then, FRN captures the intra-feature and inter-feature relationships for la-tent features to characterize expression-specific variations, and reconstructs the expression feature. To this end, two modules including an intra-feature relation modeling module and an inter-feature relation modeling module are developed in FRN. Experimental results on both the in-the-lab databases (including CK+, MMI, and Oulu-CASIA) and the in-the-wild databases (including RAF-DB and SFEW) show that the proposed FDRL method consistently achieves higher recognition accuracy than several state-of-the-art methods. This clearly highlights the benefit of feature decomposition and reconstruction for classifying expressions. Delian Ruan, Yan Yan 0001, Shenqi Lai, Zhenhua Chai, Chunhua Shen, Hanzi Wang |
CVPR | 2 |
| 2021 | Learning Spatial-Semantic Relationship for Facial Attribute Recognition With Limited Labeled DataabstractRecent advances in deep learning have demonstrated excellent results for Facial Attribute Recognition (FAR), typically trained with large-scale labeled data. However, in many real-world FAR applications, only limited labeled data are available, leading to remarkable deterioration in performance for most existing deep learning-based FAR methods. To address this problem, here we propose a method termed Spatial-Semantic Patch Learning (SSPL). The training of SSPL involves two stages. First, three auxiliary tasks, consisting of a Patch Rotation Task (PRT), a Patch Segmentation Task (PST), and a Patch Classification Task (PCT), are jointly developed to learn the spatial-semantic relationship from large-scale unlabeled facial data. We thus obtain a powerful pre-trained model. In particular, PRT exploits the spatial information of facial images in a self-supervised learning manner. PST and PCT respectively capture the pixel-level and image-level semantic information of facial images based on a facial parsing model. Second, the spatial-semantic knowledge learned from auxiliary tasks is transferred to the FAR task. By doing so, it enables that only a limited number of labeled data are required to fine-tune the pre-trained model. We achieve superior performance compared with state-of-the-art methods, as substantiated by extensive experiments and studies. Ying Shu, Yan Yan 0001, Si Chen 0002, Jing-Hao Xue, Chunhua Shen, Hanzi Wang |
CVPR | 2 |
| 2021 | D³Net: Dual-Branch Disturbance Disentangling Network for Facial Expression RecognitionabstractOne of the main challenges in facial expression recognition (FER) is to address the disturbance caused by various disturbing factors, including common ones (such as identity, pose, and illumination) and potential ones (such as hairstyle, accessory, and occlusion). Recently, a number of FER methods have been developed to explicitly or implicitly alleviate the disturbance involved in facial images. However, these methods either consider only a few common disturbing factors or neglect the prior information of these disturbing factors, thus resulting in inferior recognition performance. In this paper, we propose a novel Dual-branch Disturbance Disentangling Network (D3Net), mainly consisting of an expression branch and a disturbance branch, to perform effective FER. In the disturbance branch, a label-aware sub-branch (LAS) and a label-free sub-branch (LFS) are elaborately designed to cope with different types of disturbing factors. On the one hand, LAS explicitly captures the disturbance due to some common disturbing factors by transfer learning on a pretrained model. On the other hand, LFS implicitly encodes the information of potential disturbing factors in an unsupervised manner. In particular, we introduce an Indian buffet process (IBP) prior to model the distribution of potential disturbing factors in LFS. Moreover, we leverage adversarial training to increase the differences between disturbance features and expression features, thereby enhancing the disentanglement of disturbing factors. By disentangling the disturbance from facial images, we are able to extract discriminative expression features. Extensive experiments demonstrate that our proposed method performs favorably against several state-of-the-art FER methods on both in-the-lab and in-the-wild databases. Rongyun Mo, Yan Yan 0001, Jing-Hao Xue, Si Chen 0002, Hanzi Wang |
ACM Multimedia | 2 |
| 2021 | Towards a Unified Middle Modality Learning for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) aims to search identities of pedestrians across different spectra. In this task, one of the major challenges is the modality discrepancy between the visible (VIS) and infrared (IR) images. Some state-of-the-art methods try to design complex networks or generative methods to mitigate the modality discrepancy while ignoring the highly non-linear relationship between the two modalities of VIS and IR. In this paper, we propose a non-linear middle modality generator (MMG), which helps to reduce the modality discrepancy. Our MMG can effectively project VIS and IR images into a unified middle modality image (UMMI) space to generate middle-modality (M-modality) images. The generated M-modality images and the original images are fed into the backbone network to reduce the modality discrepancy.Furthermore, in order to pull together the two types of M-modality images generated from the VIS and IR images in the UMMI space, we propose a distribution consistency loss (DCL) to make the modality distribution of the generated M-modalities images as consistent as possible. Finally, we propose a middle modality network (MMN) to further enhance the discrimination and richness of features in an explicit manner. Extensive experiments have been conducted to validate the superiority of MMN for VI-ReID over some state-of-the-art methods on two challenging datasets. The gain of MMN is more than 11.1% and 8.4% in terms of Rank-1 and mAP, respectively, even compared with the latest state-of-the-art methods on the SYSU-MM01 dataset. Yan Yan 0001, Yang Lu 0009, Hanzi Wang |
ACM Multimedia | 2 |
| 2021 | Small-Vote Sample Selection for Label-Noise Learning
Youze Xu, Yan Yan 0001, Jing-Hao Xue, Yang Lu 0009, Hanzi Wang |
ECML/PKDD (3) | 2 |
| 2021 | Guest Editorial: Special issue on deep learning with small samples
Jing-Hao Xue, Jufeng Yang, Yan Yan 0001, Yujiu Yang 0001, Zongqing Lu 0001, Zhanyu Ma |
Neurocomputing | 4 |
| 2021 | Robust visual tracking via spatio-temporal adaptive and channel selective correlation filters
Yan Yan 0001, Liming Zhang 0002, Hanzi Wang |
Pattern Recognit. | 3 |
| 2021 | Recurrent Context Aggregation Network for Single Image DehazingabstractExisting learning-based dehazing methods are prone to cause excessive dehazing and failure to dense haze, mainly because that the global features of hazy images are not fully utilized, while the local features of hazy images are not enough discriminative. In this letter, we propose a Recurrent Context Aggregation Network (RCAN) to effectively dehaze images and restore color fidelity. In RCAN, an efficient and generic module, called Context Aggression Block (CAB), is designed to improve the feature representation by taking advantage of both global and local features, which are complementary for robust dehazing because that local features can capture different levels of haze, and global features can focus on textures and object edges of a whole image. In addition, RCAN adopts a deep recurrent mechanism to improve the dehazing performance without introducing additional network parameters. Extensive experimental results on both synthetic and real-world datasets show that the proposed RCAN performs better than other state-of-the-art dehazing methods. Runqing Chen, Yang Lu 0009, Yan Yan 0001, Hanzi Wang |
IEEE Signal Process. Lett. | 4 |
| 2021 | Semantic-Aware Occlusion-Robust Network for Occluded Person Re-IdentificationabstractIn recent years, deep learning-based person re-identification (Re-ID) methods have made significant progress. However, the performance of these methods substantially decreases when dealing with occlusion, which is ubiquitous in realistic scenarios. In this article, we propose a novel semantic-aware occlusion-robust network (SORN) that effectively exploits the intrinsic relationship between the tasks of person Re-ID and semantic segmentation for occluded person Re-ID. Specifically, the SORN is composed of three branches, including a local branch, a global branch, and a semantic branch. In particular, the local branch extracts part-based local features, and the global branch leverages a novel spatial-patch contrastive loss (SPC) to extract occlusion-robust global features. Meanwhile, the semantic branch generates a foreground-background mask for a pedestrian image, which indicates the non-occluded areas of the human body. The three branches are jointly trained in a unified multi-task learning network. Finally, pedestrian matching is performed based on the local features extracted from the non-occluded areas and the global features extracted from the whole pedestrian image. Extensive experimental results on a large-scale occluded person Re-ID dataset (i.e., Occluded-DukeMTMC) and two partial person Re-ID datasets (i.e., Partial-REID and Partial-iLIDS) show the superiority of the proposed method compared with several state-of-the-art methods for occluded and partial person Re-ID. We also demonstrate the effectiveness of the proposed method on two general person Re-ID datasets (i.e., Market-1501 and DukeMTMC-reID). Yan Yan 0001, Jing-Hao Xue, Yang Hua 0001, Hanzi Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Real-Time High-Performance Semantic Image Segmentation of Urban Street ScenesabstractDeep Convolutional Neural Networks (DCNNs) have recently shown outstanding performance in semantic image segmentation. However, state-of-the-art DCNN-based semantic segmentation methods usually suffer from high computational complexity due to the use of complex network architectures. This greatly limits their applications in the real-world scenarios that require real-time processing. In this paper, we propose a real-time high-performance DCNN-based method for robust semantic segmentation of urban street scenes, which achieves a good trade-off between accuracy and speed. Specifically, a Lightweight Baseline Network with Atrous convolution and Attention (LBN-AA) is firstly used as our baseline network to efficiently obtain dense feature maps. Then, the Distinctive Atrous Spatial Pyramid Pooling (DASPP), which exploits the different sizes of pooling operations to encode the rich and distinctive semantic information, is developed to detect objects at multiple scales. Meanwhile, a Spatial detail-Preserving Network (SPN) with shallow convolutional layers is designed to generate high-resolution feature maps preserving the detailed spatial information. Finally, a simple but practical Feature Fusion Network (FFN) is used to effectively combine both deep and shallow features from the semantic branch (DASPP) and the spatial branch (SPN), respectively. Extensive experimental results show that the proposed method respectively achieves the accuracy of 73.6% and 68.0% mean Intersection over Union (mIoU) at the inference speeds of 51.0 fps and 39.3 fps on the challenging Cityscapes and CamVid test datasets (by only using a single NVIDIA TITAN X card). This demonstrates that the proposed method offers excellent performance at the real-time speed for semantic segmentation of urban street scenes. Genshun Dong, Yan Yan 0001, Chunhua Shen, Hanzi Wang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Multi-Stream Siamese and Faster Region-Based Neural Network for Real-Time Object TrackingabstractObject tracking is a challenging task in computer vision based intelligent transportation systems. Recently, Siamese based object tracking methods have attracted significant attention due to their highly efficient performance. These tracking methods usually train a Siamese network to match the initial target patch of the first frame with candidates in a new frame. In these methods, the offline training of the deep neural network and the online instance searching are effectively combined. However, these methods usually do not include template update or object re-identification, which easily results in the drift problem. In this paper, we propose a novel real-time object tracking method to overcome the above problems by effectively combining a multi-stream Siamese network and a region-based convolutional neural network. Specifically, a novel multi-stream Siamese network is proposed to search the target and update the instance template in a new frame. In addition, a faster region-based convolutional neural network detector is used to perform object re-identification in order to improve the tracking performance by making full use of the object category information. These two networks are tightly coupled to ensure that the proposed tracking method has high efficiency and strong discriminative capability. Experimental results on several object tracking benchmarks show that our tracking method can effectively track vehicles and pedestrians in video sequences by exploiting the object category information. The proposed tracking method achieves real-time operations and outperforms several other state-of-the-art methods. Liming Zhang 0002, Yan Yan 0001, Hanzi Wang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2020 | Learning Target-Specific Response Attention for Siamese Network Based Visual Tracking
Penghui Zhao, Haosheng Chen 0001, Yan Yan 0001, Hanzi Wang |
ACIVS | 4 |
| 2020 | Deep Disturbance-Disentangled Learning for Facial Expression RecognitionabstractTo achieve effective facial expression recognition (FER), it is of great importance to address various disturbing factors, including pose, illumination, identity, and so on. However, a number of FER databases merely provide the labels of facial expression, identity, and pose, but lack the label information for other disturbing factors. As a result, many methods are only able to cope with one or two disturbing factors, ignoring the heavy entanglement between facial expression and multiple disturbing factors. In this paper, we propose a novel Deep Disturbance-disentangled Learning (DDL) method for FER. DDL is capable of simultaneously and explicitly disentangling multiple disturbing factors by taking advantage of multi-task learning and adversarial transfer learning. The training of DDL involves two stages. First, a Disturbance Feature Extraction Model (DFEM) is pre-trained to perform multi-task learning for classifying multiple disturbing factors on the large-scale face database (which has the label information for various disturbing factors). Second, a Disturbance-Disentangled Model (DDM), which contains a global shared sub-network and two task-specific (i.e., expression and disturbance) sub-networks, is learned to encode the disturbance-disentangled information for expression recognition. The expression sub-network adopts a multi-level attention mechanism to extract expression-specific features, while the disturbance sub-network leverages adversarial transfer learning to extract disturbance-specific features based on the pre-trained DFEM. Experimental results on both the in-the-lab FER databases (including CK+, MMI, and Oulu-CASIA) and the in-the-wild FER databases (including RAF-DB and SFEW) demonstrate the superiority of our proposed method compared with several state-of-the-art methods. Delian Ruan, Yan Yan 0001, Si Chen 0002, Jing-Hao Xue, Hanzi Wang |
ACM Multimedia | 2 |
| 2020 | Learning intra-inter semantic aggregation for video object detectionabstractVideo object detection is a challenging task due to the appearance deterioration problems in video frames. Thus, object features extracted from different frames of a video are usually deteriorated in varying degrees. Currently, some state-of-the-art methods enhance the deteriorated object features in a reference frame by aggregating the undeteriorated object features extracted from other frames, simply based on their learned appearance relation among object features. In this paper, we propose a novel intra-inter semantic aggregation method (ISA) to learn more effective intra and inter relations for semantically aggregating object features. Specifically, in the proposed ISA, we first introduce an intra semantic aggregation module (Intra-SAM) to enhance the deteriorated spatial features based on the learned intra relation among the features at different positions of an individual object. Then, we present an inter semantic aggregation module (Inter-SAM) to enhance the deteriorated object features in the temporal domain based on the learned inter relation among object features. As a result, by leveraging Intra-SAM and Inter-SAM, the proposed ISA can generate discriminative features from the novel perspective of intra-inter semantic aggregation for robust video object detection. We conduct extensive experiments on the ImageNet VID dataset to evaluate ISA. The proposed ISA obtains 84.5% mAP and 85.2% mAP with ResNet-101 and ResNeXt-101, and it achieves superior performance compared with several state-of-the-art video object detectors. Haosheng Chen 0001, Kaiwen Du, Yan Yan 0001, Hanzi Wang |
MMAsia | 4 |
| 2020 | Robust visual tracking via scale-aware localization and peak response strengthabstractExisting regression-based deep trackers usually localize a target based on a response map, where the highest peak response corresponds to the predicted target location. Nevertheless, when the background distractors appear or the target scale changes frequently, the response map is prone to produce multiple sub-peak responses to interfere with model prediction. In this paper, we propose a robust online tracking method via Scale-Aware localization and Peak Response strength (SAPR), which can learn a discriminative model predictor to estimate a target state accurately. Specifically, to cope with large scale variations, we propose a Scale-Aware Localization (SAL) module to provide multi-scale response maps based on the scale pyramid scheme. Furthermore, to focus on the target response, we propose a simple yet effective Peak Response Strength (PRS) module to fuse the multi-scale response maps and the response maps generated by a correlation filter. According to the response map with the maximum classification score, the model predictor iteratively updates its filter weights for accurate target state estimation. Experimental results on three benchmark datasets, including OTB100, VOT2018 and LaSOT, demonstrate that the proposed SAPR accurately estimates the target state, achieving the favorable performance against several state-of-the-art trackers. Luo Xiong, Kaiwen Du, Yan Yan 0001, Hanzi Wang |
MMAsia | 4 |
| 2020 | Large margin deep embedding for aesthetic image classification
Guanjun Guo, Hanzi Wang, Yan Yan 0001, Liming Zhang 0002, Bo Li 0006 |
Sci. China Inf. Sci. | 3 |
| 2020 | Object-adaptive LSTM network for real-time visual tracking with adversarial data augmentation
Yihan Du, Yan Yan 0001, Si Chen 0002, Yang Hua 0001 |
Neurocomputing | 2 |
| 2020 | A fast face detection method via convolutional neural network
Guanjun Guo, Hanzi Wang, Yan Yan 0001, Bo Li 0006 |
Neurocomputing | 3 |
| 2020 | A two-step hypergraph reduction based fitting method for unbalanced data
Guobao Xiao, Yan Yan 0001, Hanzi Wang |
Pattern Recognit. Lett. | 3 |
| 2020 | Low-resolution facial expression recognition: A filter learning perspective
Yan Yan 0001, Zizhao Zhang 0002, Si Chen 0002, Hanzi Wang |
Signal Process. | 1 |
| 2020 | Accelerated Guided Sampling for Multistructure Model FittingabstractThe performance of many robust model fitting techniques is largely dependent on the quality of the generated hypotheses. In this paper, we propose a novel guided sampling method, called accelerated guided sampling (AGS), to efficiently generate the accurate hypotheses for multistructure model fitting. Based on the observations that residual sorting can effectively reveal the data relationship (i.e., determine whether two data points belong to the same structure), and keypoint matching scores can be used to distinguish inliers from gross outliers, AGS effectively combines the benefits of residual sorting and keypoint matching scores to efficiently generate accurate hypotheses via information theoretic principles. Moreover, we reduce the computational cost of residual sorting in AGS by designing a new residual sorting strategy, which only sorts the top-ranked residuals of input data, rather than all input data. Experimental results demonstrate the effectiveness of the proposed method in computer vision tasks, such as homography matrix and fundamental matrix estimation. Taotao Lai, Hanzi Wang, Yan Yan 0001, Tat-Jun Chin, Bo Li 0006 |
IEEE Trans. Cybern. | 3 |
| 2020 | Joint Deep Learning of Facial Expression Synthesis and RecognitionabstractRecently, deep learning based facial expression recognition (FER) methods have attracted considerable attention and they usually require large-scale labelled training data. Nonetheless, the publicly available facial expression databases typically contain a small amount of labelled data. In this paper, to overcome the above issue, we propose a novel joint deep learning of facial expression synthesis and recognition method for effective FER. More specifically, the proposed method involves a two-stage learning procedure. Firstly, a facial expression synthesis generative adversarial network (FESGAN) is pre-trained to generate facial images with different facial expressions. To increase the diversity of the training images, FESGAN is elaborately designed to generate images with new identities from a prior distribution. Secondly, an expression recognition network is jointly learned with the pre-trained FESGAN in a unified framework. In particular, the classification loss computed from the recognition network is used to simultaneously optimize the performance of both the recognition network and the generator of FESGAN. Moreover, in order to alleviate the problem of data bias between the real images and the synthetic images, we propose an intra-class loss with a novel real data-guided back-propagation (RDBP) algorithm to reduce the intra-class variations of images from the same class, which can significantly improve the final performance. Extensive experimental results on public facial expression databases demonstrate the superiority of the proposed method compared with several state-of-the-art FER methods. Yan Yan 0001, Si Chen 0002, Chunhua Shen, Hanzi Wang |
IEEE Trans. Multim. | 1 |
| 2019 | Hypergraph Optimization for Multi-Structural Geometric Model FittingabstractRecently, some hypergraph-based methods have been proposed to deal with the problem of model fitting in computer vision, mainly due to the superior capability of hypergraph to represent the complex relationship between data points. However, a hypergraph becomes extremely complicated when the input data include a large number of data points (usually contaminated with noises and outliers), which will significantly increase the computational burden. In order to overcome the above problem, we propose a novel hypergraph optimization based model fitting (HOMF) method to construct a simple but effective hypergraph. Specifically, HOMF includes two main parts: an adaptive inlier estimation algorithm for vertex optimization and an iterative hyperedge optimization algorithm for hyperedge optimization. The proposed method is highly efficient, and it can obtain accurate model fitting results within a few iterations. Moreover, HOMF can then directly apply spectral clustering, to achieve good fitting performance. Extensive experimental results show that HOMF outperforms several state-of-the-art model fitting methods on both synthetic data and real images, especially in sampling efficiency and in handling data with severe outliers. Shuyuan Lin, Guobao Xiao, Yan Yan 0001, David Suter, Hanzi Wang |
AAAI | 3 |
| 2019 | A Cascaded Noise-Robust Deep CNN for Face RecognitionabstractState-of-the-art face recognition methods have achieved excellent performance on the clean datasets. However, in real-world applications, the captured face images are usually contaminated with noise, which significantly decreases the performance of these face recognition methods. In this paper, we propose a cascaded noise-robust deep convolutional neural network (CNR-CNN) method, consisting of two sub-networks, i.e., a denoising sub-network and a face recognition sub-network, for face recognition under noise. Instead of separately training the two sub-networks, we jointly train them in a cascaded manner. As a result, the images generated from the denoising sub-network are beneficial to the training of the face recognition sub-network. Furthermore, the dense connectivity is used to concatenate the feature maps layer-by-layer in the denoising sub-network, which can effectively exploit the shallow and deep features of CNN. Experimental results on public face datasets demonstrate the superior performance of the proposed method over several state-of-the-art methods. Xiangbang Meng, Yan Yan 0001, Si Chen 0002, Hanzi Wang |
ICIP | 2 |
| 2019 | Correlation Filter Tracking with Adaptive Proposal Selection for Accurate Scale EstimationabstractRecently, some correlation filter based trackers with detection proposals have achieved state-of-the-art tracking results. However, a large number of redundant proposals given by the proposal generator may degrade the performance and speed of these trackers. In this paper, we propose an adaptive proposal selection algorithm which can generate a small number of high-quality proposals to handle the problem of scale variations for visual object tracking. Specifically, we firstly utilize the color histograms in the HSV color space to represent the instances (i.e., the initial target in the first frame and the predicted target in the previous frame) and proposals. Then, an adaptive strategy based on the color similarity is formulated to select high-quality proposals. We further integrate the proposed adaptive proposal selection algorithm with coarse-to-fine deep features to validate the generalization and efficiency of the proposed tracker. Experiments on two benchmark datasets demonstrate that the proposed algorithm performs favorably against several state-of-the-art trackers. Luo Xiong, Yan Yan 0001, Hanzi Wang |
ICME | 3 |
| 2019 | Robust Visual Tracking via Statistical Positive Sample Generation and Gradient Aware LearningabstractIn recent years, Convolutional Neural Network (CNN) based trackers have achieved state-of-the-art performance on multiple benchmark datasets. Most of these trackers train a binary classifier to distinguish the target from its background. However, they suffer from two limitations. Firstly, these trackers cannot effectively handle significant appearance variations due to the limited number of positive samples. Secondly, there exists a significant imbalance of gradient contributions between easy and hard samples, where the easy samples usually dominate the computation of gradient. In this paper, we propose a robust tracking method via Statistical Positive sample generation and Gradient Aware learning (SPGA) to address the above two limitations. To enrich the diversity of positive samples, we present an effective and efficient statistical positive sample generation algorithm to generate positive samples in the feature space. Furthermore, to handle the issue of imbalance between easy and hard samples, we propose a gradient sensitive loss to harmonize the gradient contributions between easy and hard samples. Extensive experiments on three challenging benchmark datasets including OTB50, OTB100 and VOT2016 demonstrate that the proposed SPGA performs favorably against several state-of-the-art trackers. Lijian Lin, Haosheng Chen 0001, Yan Yan 0001, Hanzi Wang |
MMAsia | 4 |
| 2019 | Superpixel-Guided Two-View Deterministic Geometric Model Fitting
Guobao Xiao, Hanzi Wang, Yan Yan 0001, David Suter |
Int. J. Comput. Vis. | 3 |
| 2019 | Robust geometric model fitting based on iterative Hypergraph Construction and Partition
Guobao Xiao, Hanzi Wang, Yan Yan 0001, Liming Zhang 0002 |
Neurocomputing | 3 |
| 2019 | Adaptive deep metric embeddings for person re-identification under occlusions
Wanxiang Yang, Yan Yan 0001, Si Chen 0002 |
Neurocomputing | 2 |
| 2019 | Searching for Representative Modes on Hypergraphs for Robust Geometric Model FittingabstractIn this paper, we propose a simple and effective geometric model fitting method to fit and segment multi-structure data even in the presence of severe outliers. We cast the task of geometric model fitting as a representative mode-seeking problem on hypergraphs. Specifically, a hypergraph is first constructed, where the vertices represent model hypotheses and the hyperedges denote data points. The hypergraph involves higher-order similarities (instead of pairwise similarities used on a simple graph), and it can characterize complex relationships between model hypotheses and data points. In addition, we develop a hypergraph reduction technique to remove "insignificant" vertices while retaining as many "significant" vertices as possible in the hypergraph. Based on the simplified hypergraph, we then propose a novel mode-seeking algorithm to search for representative modes within reasonable time. Finally, the proposed mode-seeking algorithm detects modes according to two key elements, i.e., the weighting scores of vertices and the similarity analysis between vertices. Overall, the proposed fitting method is able to efficiently and effectively estimate the number and the parameters of model instances in the data simultaneously. Experimental results demonstrate that the proposed method achieves significant superiority over several state-of-the-art model fitting methods on both synthetic data and real images. Hanzi Wang, Guobao Xiao, Yan Yan 0001, David Suter |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Learning Object Scale With Click Supervision for Object DetectionabstractWeakly-supervised object detection has recently attracted increasing attention since it only requires image-level annotations. However, the performance obtained by existing methods is still far from being satisfactory compared with fully-supervised object detection methods. To achieve a good trade-off between annotation cost and object detection performance, we propose a simple yet effective method which incorporates CNN visualization with click supervision to generate the pseudo ground-truths (i.e., bounding boxes). These pseudo ground-truths can be used to train a fully-supervised detector. To estimate the object scale, we firstly adopt a proposal selection algorithm to preserve high-quality proposals, and then generate Class Activation Maps (CAMs) for these preserved proposals by the proposed CNN visualization algorithm called Spatial Attention CAM. Finally, we fuse these CAMs together to generate pseudo ground-truths and train a fully-supervised object detector with these ground-truths. Experimental results on the PASCAL VOC 2007 and VOC 2012 datasets show that the proposed method can obtain much higher accuracy for estimating the object scale, compared with the state-of-the-art image-level based methods and the center-click based method. Liao Zhang, Yan Yan 0001, Hanzi Wang |
IEEE Signal Process. Lett. | 2 |
| 2018 | DSNet: Deep and Shallow Feature Learning for Efficient Visual Tracking
Qiangqiang Wu, Yan Yan 0001, Hanzi Wang |
ACCV (5) | 2 |
| 2018 | Object-Adaptive LSTM Network for Visual TrackingabstractConvolutional Neural Networks (CNNs) have shown outstanding performance in visual object tracking. However, most of classification-based tracking methods using CNNs are time-consuming due to expensive computation of complex online fine-tuning and massive feature extractions. Besides, these methods suffer from the problem of over-fitting since the training and testing stages of CNN models are based on the videos from the same domain. Recently, matching-based tracking methods (such as Siamese networks) have shown remarkable speed superiority, while they cannot well address target appearance variations and complex scenes for inherent lack of online adaptability and background information. In this paper, we propose a novel object-adaptive LSTM network, which can effectively exploit sequence dependencies and dynamically adapt to the temporal object variations via constructing an intrinsic model for object appearance and motion. In addition, we develop an efficient strategy for proposal selection, where the densely sampled proposals are firstly pre-evaluated using the fast matching-based method and then the well-selected high-quality proposals are fed to the sequence-specific learning LSTM network. This strategy enables our method to adaptively track an arbitrary object and operate faster than conventional CNN-based classification tracking methods. To the best of our knowledge, this is the first work to apply an LSTM network for classification in visual object tracking. Experimental results on OTB and TC-128 benchmarks show that the proposed method achieves state-of-the-art performance, which exhibits great potentials of recurrent structures for visual object tracking. Yihan Du, Yan Yan 0001, Si Chen 0002, Yang Hua 0001, Hanzi Wang |
ICPR | 2 |
| 2018 | Improved Correlation Filter Tracking with Hard Negative MiningabstractRecently, the correlation filter based trackers have achieved very good tracking performance. However, due to the boundary effects of the circulant matrix and the usage of cosine window, the lack of effective negative samples becomes a challenging problem for the correlation filter based trackers. This problem may cause overfitting so that these trackers become very sensitive to deformation and occlusion. In this paper, we propose a novel object tracker (i.e., STAPLE_HNM), which can effectively select hard negative samples and assign adaptive weights to these samples to train the correlation filter. Experimental results demonstrate that the proposed STAPLE_HNM tracker effectively improves the performance of the baseline STAPLE_CA tracker on the OTB-50 and OTB-100 datasets. Moreover, the proposed STAPLE_HNM tracker also achieves superior performance among several state-of-the-art trackers. Chunguang Qie, Guanjun Guo, Yan Yan 0001, Liming Zhang 0002, Hanzi Wang |
ICPR | 3 |
| 2018 | Multi-task Learning of Cascaded CNN for Facial Attribute ClassificationabstractRecently, facial attribute classification (FAC) has attracted significant attention in the computer vision community. Great progress has been made along with the availability of challenging FAC datasets. However, conventional FAC methods usually firstly pre-process the input images (i.e., perform face detection and alignment) and then predict facial attributes. These methods ignore the inherent dependencies among these tasks (i.e., face detection, facial landmark localization and FAC). Moreover, some methods using convolutional neural network are trained based on the fixed loss weights without considering the differences between facial attributes. In order to address the above problems, we propose a novel multi-task learning of cascaded convolutional neural network method, termed MCFA, for predicting multiple facial attributes simultaneously. Specifically, the proposed method takes advantage of three cascaded sub-networks (i.e., S_Net, M_Net and L_Net corresponding to the neural networks under different scales) to jointly train multiple tasks in a coarse-to-fine manner, which can achieve end-to-end optimization. Furthermore, the proposed method automatically assigns the loss weight to each facial attribute based on a novel dynamic weighting scheme, thus making the proposed method concentrate on predicting the more difficult facial attributes. Experimental results show that the proposed method outperforms several state-of-the-art FAC methods on the challenging CelebA and LFWA datasets. Ni Zhuang, Yan Yan 0001, Si Chen 0002, Hanzi Wang |
ICPR | 2 |
| 2018 | Robust Correlation Filter Tracking with Shepherded Instance-Aware ProposalsabstractIn recent years, convolutional neural network (CNN) based correlation filter trackers have achieved state-of-the-art results on the benchmark datasets. However, the CNN based correlation filters cannot effectively handle large scale variation and distortion (such as fast motion, background clutter, occlusion, etc.), leading to the sub-optimal performance. In this paper, we propose a novel CNN based correlation filter tracker with shepherded instance-aware proposals, namely DeepCFIAP, which automatically estimates the target scale in each frame and re-detects the target when distortion happens. DeepCFIAP is proposed to take advantage of the merits of both instance-aware proposals and CNN based correlation filters. Compared with the CNN based correlation filter trackers, DeepCFIAP can successfully solve the problems of large scale variation and distortion via the shepherded instance-aware proposals, resulting in more robust tracking performance. Specifically, we develop a novel proposal ranking algorithm based on the similarities between proposals and instances. In contrast to the detection proposal based trackers, DeepCFIAP shepherds the instance-aware proposals towards their optimal positions via the CNN based correlation filters, resulting in more accurate tracking results. Extensive experiments on two challenging benchmark datasets demonstrate that the proposed DeepCFIAP performs favorably against state-of-the-art trackers and it is especially feasible for long-term tracking. Qiangqiang Wu, Yan Yan 0001, Hanzi Wang |
ACM Multimedia | 4 |
| 2018 | Conceptual space based model fitting for multi-structure data
Guobao Xiao, Hailing Luo, Bo Li 0006, Yan Yan 0001, Hanzi Wang |
Neurocomputing | 6 |
| 2018 | Expression-targeted feature learning for effective facial expression recognition
Yan Yan 0001, Si Chen 0002, Hanzi Wang |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Revisiting graph construction for fast image segmentation
Zizhao Zhang 0002, Fuyong Xing, Hanzi Wang, Yan Yan 0001, Xiaoshuang Shi, Lin Yang 0002 |
Pattern Recognit. | 4 |
| 2018 | Multi-label learning based deep transfer neural network for facial attribute classification
Ni Zhuang, Yan Yan 0001, Si Chen 0002, Hanzi Wang, Chunhua Shen |
Pattern Recognit. | 2 |
| 2018 | Object Discovery via Cohesion MeasurementabstractColor and intensity are two important components in an image. Usually, groups of image pixels, which are similar in color or intensity, are an informative representation for an object. They are therefore particularly suitable for computer vision tasks, such as saliency detection and object proposal generation. However, image pixels, which share a similar real-world color, may be quite different since colors are often distorted by intensity. In this paper, we reinvestigate the affinity matrices originally used in image segmentation methods based on spectral clustering. A new affinity matrix, which is robust to color distortions, is formulated for object discovery. Moreover, a cohesion measurement (CM) for object regions is also derived based on the formulated affinity matrix. Based on the new CM, a novel object discovery method is proposed to discover objects latent in an image by utilizing the eigenvectors of the affinity matrix. Then we apply the proposed method to both saliency detection and object proposal generation. Experimental results on several evaluation benchmarks demonstrate that the proposed CM-based method has achieved promising performance for these two tasks. Guanjun Guo, Hanzi Wang, Wanlei Zhao, Yan Yan 0001, Xuelong Li 0001 |
IEEE Trans. Cybern. | 4 |
| 2018 | Automatic Image Cropping for Visual Aesthetic Enhancement Using Deep Neural Networks and Cascaded RegressionabstractDespite recent progress, computational visual aesthetic is still challenging. Image cropping, which refers to the removal of unwanted scene areas, is an important step to improve the aesthetic quality of an image. However, it is challenging to evaluate whether cropping leads to aesthetically pleasing results because the assessment is typically subjective. In this paper, we propose a novel cascaded cropping regression (CCR) method to perform image cropping by learning the knowledge from professional photographers. The proposed CCR method improves the convergence speed of the cascaded method, which directly uses random-ferns regressors. In addition, a two-step learning strategy is proposed and used in the CCR method to address the problem of lacking labelled cropping data. Specifically, a deep convolutional neural network (CNN) classifier is first trained on large-scale visual aesthetic datasets. The deep CNN model is then designed to extract features from several image cropping datasets, upon which the cropping bounding boxes are predicted by the proposed CCR method. Experimental results on public image cropping datasets demonstrate that the proposed method significantly outperforms several state-of-the-art image cropping methods. Guanjun Guo, Hanzi Wang, Chunhua Shen, Yan Yan 0001, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 4 |
| 2017 | A Hierarchical Voting Scheme for Robust Geometric Model Fitting
Guobao Xiao, Yan Yan 0001, Hanzi Wang |
ICIG (1) | 5 |
| 2017 | An efficient deep neural networks training framework for robust face recognitionabstractIn recent years, the triplet loss-based deep neural networks (DNN) are widely used in the task of face recognition and achieve the state-of-the-art performance. However, the complexity of training the triplet loss-based DNN is significantly high due to the difficulty in generating high-quality training samples. In this paper, we propose a novel DNN training framework to accelerate the training process of the triplet loss-based DNN and meanwhile to improve the performance of face recognition. More specifically, the proposed framework contains two stages: 1) The DNN initialization. A deep architecture based on the softmax loss function is designed to initialize the DNN. 2) The adaptive fine-tuning. Based on the trained model, a set of high-quality triplet samples is generated and used to fine-tune the network, where an adaptive triplet loss function is introduced to improve the discriminative ability of DNN. Experimental results show that, the model obtained by the proposed DNN training framework achieves 97.3% accuracy on the LFW benchmark with low training complexity, which verifies the efficiency and effectiveness of the proposed framework. Canping Su, Yan Yan 0001, Si Chen 0002, Hanzi Wang |
ICIP | 2 |
| 2017 | Weighted median-shift on graphs for geometric model fittingabstractIn this paper, we deal with geometric model fitting problems on graphs, where each vertex represents a model hypothesis, and each edge represents the similarity between two model hypotheses. Conventional median-shift methods are very efficient and they can automatically estimate the number of clusters. However, they assign the same weighting scores to all vertices of a graph, which can not show the discriminability on different vertices. Therefore, we propose a novel weighted median-shift on graphs method (WMSG) to fit and segment multiple-structure data. Specifically, we assign a weighting score to each vertex according to the distribution of the corresponding inliers. After that, we shift vertices towards the weighted median vertices iteratively to detect modes. The proposed method can adaptively estimate the number of model instances and deal with data contaminated with a large number of outliers. Experimental results on both synthetic data and real images show the advantages of the proposed method over several state-of-the-art model fitting methods. Hanzi Wang, Guobao Xiao, Yan Yan 0001, Liming Zhang 0002 |
ICIP | 5 |
| 2017 | Efficient guided hypothesis generation for multi-structure epipolar geometry estimation
Taotao Lai, Hanzi Wang, Yan Yan 0001, Guobao Xiao, David Suter |
Comput. Vis. Image Underst. | 3 |
| 2017 | A unified hypothesis generation framework for multi-structure model fitting
Taotao Lai, Hanzi Wang, Yan Yan 0001, Liming Zhang 0002 |
Neurocomputing | 3 |
| 2017 | A novel robust model fitting approach towards multiple-structure data segmentation
Yan Yan 0001, Min Liu 0006, Si Chen 0002 |
Neurocomputing | 1 |
| 2017 | Motion Segmentation Via a Sparsity ConstraintabstractMotion segmentation is an important task for intelligent transportation systems. In this paper, inspired by the fact that a feature point trajectory can be sparsely represented as a combination of several feature point trajectories that share coherent transformations, an efficient and effective motion segmentation method with a sparsity constraint is proposed. Specifically, we first propose an accumulated scheme to efficiently integrate motion information from all the frames of a video sequence to construct a correlation matrix. Then, a sparse affinity matrix is built on the correlation matrix by using information-theoretic principles, where the nonzero elements in the same row of the sparse affinity matrix correspond to the feature point trajectories more likely belonging to the same motion. Thereafter, a segment and merge procedure is proposed to effectively estimate the number of motions via the sparse affinity matrix. Finally, by applying spectral clustering on the sparse affinity matrix, different motions in the video sequence are accurately segmented based on the estimated number of motions. Experimental results on theHopkins 155and the62-clipdatasets demonstrate that the proposed method achieves superior performance compared with several state-of-the-art methods. Taotao Lai, Hanzi Wang, Yan Yan 0001, Tat-Jun Chin, Wanlei Zhao |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2017 | Local Co-Occurrence Selection via Partial Least Squares for Pedestrian DetectionabstractChannel feature detectors are the most popular approaches for pedestrian detection recently. However, most of these approaches train the boosted decision trees by selecting a single feature at each node, which does not effectively exploit the multi-feature cues and spatial information. To address this issue, this paper proposes to construct the co-occurrence of multiple channel features in local image neighborhoods for pedestrian detection. In our approach, a binary pattern of feature co-occurrence is represented by combining the binary variables quantized from each channel feature, and the spatial information is incorporated by selecting the neighbors to jointly represent the feature co-occurrence in a local image block. However, feature co-occurrence selection leads to many possible feature combinations, which significantly increase the computational cost at the training stage. Therefore, in order to reduce the number of candidate features and obtain the most discriminative features effectively, a partial least squares-based feature selection approach called variable importance on projection is exploited. Comprehensive experiments are conducted on several challenging pedestrian data sets, and superior performances are achieved by the proposed approach in comparison with some state-of-the-art pedestrian detection approaches. Hanzi Wang, Yan Yan 0001, Bo Li 0006, Chang Wen Chen |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2016 | Message Passing on the Two-Layer Network for Geometric Model Fitting
Guobao Xiao, Yan Yan 0001, Hanzi Wang |
ACCV (1) | 3 |
| 2016 | Superpixel-Based Two-View Deterministic Fitting for Multiple-Structure Data
Guobao Xiao, Hanzi Wang, Yan Yan 0001, David Suter |
ECCV (6) | 3 |
| 2016 | Conceptual space based gross outlier removal for geometric model fittingabstractIn this paper, we propose an efficient and robust gross outlier removal method, called the Conceptual Space based Gross Outlier Removal (CSGOR) method, to remove gross outliers for geometric model fitting. In the proposed method, each data point is mapped to a conceptual space by computing the preference of "good" model hypotheses. In the conceptual space, the distributions of inliers and gross outliers are significantly different. Specifically, inliers of each model instance are distributed in a subspace and they are far away from the origin of the conceptual space, while gross outliers are distributed near the origin. In this manner, the problem of densely gross outlier removal is formulated as a binary classification problem. The main advantage of the proposed method is that it can handle data with a large proportion of outliers and effectively remove gross outliers in data. Experimental results on both synthetic and real data have demonstrated the efficiency and effectiveness of the proposed method. Guobao Xiao, Bo Li 0006, Yan Yan 0001, Hanzi Wang |
ICARCV | 5 |
| 2016 | Sparse similarity metric learning for kinship verificationabstractMetric learning technique learns a linear transformation of the given training data which can significantly promote the performance of a prediction task, such as kinship verification. However, many of the existing metric learning methods do not explicitly regularize for sparsity or low-rank, which in practice usually results in high-rank solutions that are not only time-consuming but also tend to overfitting. In addition, some methods simply neglect the positive semidefinite (PSD) constraint causing the learned metric to be potentially noisy. In this paper, we propose an effective sparse similarity metric learning (SSML) method which enforces both the group sparsity and the PSD constraints on the learned similarity matrix for kinship verification. In order to solve the proposed optimization problem efficiently, we successfully apply the alternating direction method of multipliers (ADMM) to obtain the optimal solution. Experimental results demonstrate that the proposed method achieves competitive results compared with other state-of-the-art metric learning methods on widely used kinship datasets. Yan Yan 0001, Si Chen 0002, Hanzi Wang |
VCIP | 2 |
| 2016 | Visual saliency detection based on homology similarity and an experimental evaluation
Hanzi Wang, Liming Zhang 0002, Yan Yan 0001, Hong-Yuan Mark Liao |
J. Vis. Commun. Image Represent. | 4 |
| 2016 | Discriminative local collaborative representation for online object tracking
Si Chen 0002, Shaozi Li, Rongrong Ji, Yan Yan 0001, Shunzhi Zhu |
Knowl. Based Syst. | 4 |
| 2016 | Robust visual tracking via online semi-supervised co-boosting
Si Chen 0002, Shunzhi Zhu, Yan Yan 0001 |
Multim. Syst. | 3 |
| 2016 | Rapid hypothesis generation by combining residual sorting with local constraints
Taotao Lai, Hanzi Wang, Yan Yan 0001, Dahan Wang, Guobao Xiao |
Multim. Tools Appl. | 3 |
| 2016 | Quadratic projection based feature extraction with its application to biometric recognition
Yan Yan 0001, Hanzi Wang, Si Chen 0002, Xiaochun Cao, David Zhang 0001 |
Pattern Recognit. | 1 |
| 2016 | Mode seeking on graphs for geometric model fitting via preference analysis
Guobao Xiao, Hanzi Wang, Yan Yan 0001, Liming Zhang 0002 |
Pattern Recognit. Lett. | 3 |
| 2016 | Discriminative Weighted Sparse Partial Least Squares for Human DetectionabstractChannel feature detectors have shown great advantages in human detection. However, a large pool of channel features extracted for human detection usually contains many redundant and irrelevant features. To address this issue, we propose a robust discriminative weighted sparse partial least square approach for feature selection and apply it to human detection. Unlike partial least squares (PLS), which is a straightforward dimensionality reduction technique, we propose using sparse PLS to achieve feature selection. Furthermore, in order to obtain a robust latent matrix, we formulate a discriminative regularized weighted least square problem, where a discriminative term is incorporated to effectively distinguish positive samples from negative samples. A robust sparse weight matrix is trained based on the latent matrix and used for feature selection. Finally, we use the selected channel features to train the boosted decision trees and incorporate the weights of selected features with each tree. The human detector trained by the selected features can preserve high robustness and discriminativeness. Experimental results on some challenging human data sets demonstrate that the proposed approach is effective and achieves state-of-the-art performance. Yan Yan 0001, Hanzi Wang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2015 | Mode-Seeking on Hypergraphs for Robust Geometric Model FittingabstractIn this paper, we propose a novel geometric model fitting method, called Mode-Seeking on Hypergraphs (MSH), to deal with multi-structure data even in the presence of severe outliers. The proposed method formulates geometric model fitting as a mode seeking problem on a hypergraph in which vertices represent model hypotheses and hyperedges denote data points. MSH intuitively detects model instances by a simple and effective mode seeking algorithm. In addition to the mode seeking algorithm, MSH includes a similarity measure between vertices on the hypergraph and a "weight-aware sampling" technique. The proposed method not only alleviates sensitivity to the data distribution, but also is scalable to large scale problems. Experimental results further demonstrate that the proposed method has significant superiority over the state-of-the-art fitting methods on both synthetic data and real images. Hanzi Wang, Guobao Xiao, Yan Yan 0001, David Suter |
ICCV | 3 |
| 2015 | Learning Hough regression models via bridge partial least squares for object detection
Jianyu Tang, Hanzi Wang, Yan Yan 0001 |
Neurocomputing | 3 |
| 2015 | Robust visual tracking by metric learning with weighted histogram representations
Jun Wang 0131, Hanzi Wang, Yan Yan 0001 |
Neurocomputing | 3 |
| 2014 | Structured partial least squares for simultaneous object tracking and segmentation
Bineng Zhong 0001, Xiao-Tong Yuan, Rongrong Ji, Yan Yan 0001, Zhen Cui 0001, Xiaopeng Hong, Yan Chen 0017, Tian Wang 0001, Duansheng Chen |
Neurocomputing | 4 |
| 2014 | Multi-subregion based correlation filter bank for robust face recognition
Yan Yan 0001, Hanzi Wang, David Suter |
Pattern Recognit. | 1 |
| 2014 | Efficient Semidefinite Spectral Clustering via Lagrange DualityabstractWe propose an efficient approach to semidefinite spectral clustering (SSC), which addresses the Frobenius normalization with the positive semidefinite (p.s.d.) constraint for spectral clustering. Compared with the original Frobenius norm approximation-based algorithm, the proposed algorithm can more accurately find the closest doubly stochastic approximation to the affinity matrix by considering the p.s.d. constraint. In this paper, SSC is formulated as a semidefinite programming (SDP) problem. In order to solve the high computational complexity of SDP, we present a dual algorithm based on the Lagrange dual formalization. Two versions of the proposed algorithm are proffered: one with less memory usage and the other with faster convergence rate. The proposed algorithm has much lower time complexity than that of the standard interior-point-based SDP solvers. Experimental results on both the UCI data sets and real-world image data sets demonstrate that: 1) compared with the state-of-the-art spectral clustering methods, the proposed algorithm achieves better clustering performance and 2) our algorithm is much more efficient and can solve larger-scale SSC problems than those standard interior-point SDP solvers. Yan Yan 0001, Chunhua Shen, Hanzi Wang |
IEEE Trans. Image Process. | 1 |
| 2013 | Robust Modular Linear Regression Based Classification for Face Recognition with OcclusionabstractFace recognition with occlusion is a challenging problem. Recently, the modular representation based method, i.e., modular linear regression based classification (MLRC) was proposed to deal with this problem. However, MLRC just simply combines the individual decision of each block within an image (based on the min rule) to make final decision. Therefore, the block distance information is not fully exploited. In this paper, we propose a robust modular linear regression based classification (RMLRC) method to overcome the above problem. RMLRC can effectively fuse the information provided by all the blocks and thus alleviate the limiations of the MLRC method. Experimental results show that the RMLRC method can achieve promising results for face recognition with occlusion. Guanglu Liu, Yan Yan 0001, Hanzi Wang |
ICIG | 2 |
| 2013 | Discriminative filter based regression learning for facial expression recognitionabstractIn this paper, we propose a novel discriminative filter based regression learning (DFRL) method, which can effectively remove irrelevant information while preserving useful information for facial expression recognition. DFRL integrates the filter technique and the linear analysis techniques (i.e., Linear Discriminant Analysis-LDA and Linear Ridge Regression-LRR) to obtain an effective image representation. Two steps are involved in DFRL: 1) The discriminative filters corresponding to different facial expressions are separately trained by optimizing the cost function of the two-class LDA, 2) LRR is used to extract valuable expressional information with high discriminability from the combined filtered images. Experimental results on several challenging datasets demonstrate the superior effectiveness and generalization ability of the proposed DFRL compared with other competing methods. Zizhao Zhang 0002, Yan Yan 0001, Hanzi Wang |
ICIP | 2 |
| 2013 | An effective unconstrained correlation filter and its kernelization for face recognition
Yan Yan 0001, Hanzi Wang, Cuihua Li, Chenhui Yang, Bineng Zhong 0001 |
Neurocomputing | 1 |
| 2012 | Accelerated robust sparse coding for fast face recognition
Guanglu Liu, Yan Yan 0001, Hanzi Wang |
ICPR | 2 |
| 2012 | Robust visual tracking with the cross-bin metric
Chaoxin Lyu, Yan Yan 0001, Hanzi Wang |
ICPR | 2 |