VLDB 2026 Research / reviewers in the wild / expert
Liyan Zhang 0001
dblp:54/4596-1
· DBLP profile ↗
71ranked-venue papers
9as first author
50since 2021 · last 2026
0000-0002-1549-3317ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 4 first-author · 26 since 2021Artificial intelligence and machine learning · 26 · 3 first-author · 19 since 2021Databases, data management, data science and information retrieval · 11 · 3 first-author · 2 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Domain-Adaptive Hashing via Evidential Learning and Progressive AlignmentabstractDomain-adaptive hashing enhances discriminative hash representations by transferring knowledge from a label-rich source domain to a label-scarce target domain. It has attracted significant attention due to its ability to enable efficient cross-domain retrieval without requiring target domain labels. However, existing methods generally assume that source domain labels are completely accurate. In practice, labels obtained via web crawling or crowdsourcing often contain varying degrees of noise, which hampers semantic alignment and aggravates domain shift. To tackle these issues, we propose a novel method termed Evidential Learning and Progressive Alignment (ELPA) for domain-adaptive hashing. This method comprises two key modules: the Uncertainty-aware Noise Separation (UNS) and the Progressive Cross-domain Alignment (PCA). In the UNS, we exploit the belief and uncertainty masses obtained from the evidential learning model and utilize the posterior probabilities of a Gaussian Mixture Model to effectively distinguish clean samples from noisy ones. In PCA, we introduce a progressive partial optimal transport mechanism that prioritizes pseudo-label generation for well-aligned target samples, thereby gradually achieving class-level and global-level cross-domain alignment. Extensive experiments across multiple benchmark datasets with various noise ratios demonstrate that ELPA consistently surpasses existing state-of-the-art methods, exhibiting superior robustness and generalization capability. Tiantian Gong, Yeyun Wu, Liyan Zhang 0001 |
WWW | 4 |
| 2026 | BadIQA: Backdoor Attack Against No-Reference Image Quality Assessment Models
Yinghao Wu, Liyan Zhang 0001 |
IEEE Internet Things J. | 2 |
| 2026 | Multi-Scale Adaptive Clustering and Local Consistency Learning for Unsupervised Clothing-Changing Person Re-IdentificationabstractClothing-changing person re-identification (CC-ReID) aims to address cross-camera person identification challenges caused by variations in pedestrian attire, making it a highly valuable research area within computer vision. Due to the high cost of labeling data, unsupervised learning methods have gained significant attention for CC-ReID tasks. However, existing unsupervised methods frequently suffer from high pseudo-label noise and reliance on complex preprocessing (e.g., human parsing) or multi-encoder architectures, resulting in increased computational overhead and deployment difficulties. To tackle these challenges, this paper proposes an end-to-end unsupervised CC-ReID framework named Multi-Scale Adaptive Clustering and Local Consistency Learning (MALC). Utilizing a single CLIP Vision-Encoder as its backbone, MALC discards such intricate procedures. Its core innovations include a multi-scale adaptive density clustering (MS-ADC) strategy to improve pseudo-label quality, and a local consistency learning (LCL) approach that imposes constraints on local region features to enhance robustness against clothing variations. Through the joint optimization of global and local losses, the model learns a highly discriminative and robust feature representation. Experimental results demonstrate that MALC, employing only RGB images, substantially outperforms comparable unsupervised approaches that rely on additional parsing information, showcasing notable advantages in identification accuracy and ease of deployment. The code will be made available at: https://github.com/ykding666/MALC. Yongkang Ding, Ivonne Xu, Shuangquan Lyu, Liyan Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Semantic Consistency And Integrity Network For Cloth-changing Person Re-identificationabstractCloth-changing Person Re-identification aims to retrieve target pedestrians across different cameras under clothing-changing scenarios. In recent years, many scholars have made significant explorations in this field. However, existing methods often overlook the semantic consistency and integrity of features. To address this issue, we design a Semantic Consistency and Integrity Network (SCI-Net) to learn semantically invariant features and strip clothing bias from identity features while maintaining their semantic integrity. The network consists of three branches: clothing branch, raw image branch, and head feature enhancement branch. Specifically, we first propose a Head Soft Attention Generation Module to produce head soft attention, thereby obtaining enhanced head features. Then, to ensure that raw features can effectively learn invariant semantic information from head-enhanced features, Semantic Consistency Constraint is proposed to facilitate mutual learning between the two branches. Finally, we leverage knowledge transfer to enable clothing branch to perceive clothing bias entangled with raw features and simulate causal intervention to quantify and remove clothing bias. Experiments on the LTCC-ReID and PRCC datasets demonstrate that our model outperforms other state-of-the-art methods. Anqi Wang 0010, Liyan Zhang 0001 |
ICASSP | 2 |
| 2025 | Frequency-Enhanced Part Feature Mining and Cross-Modality Alignment for Visible-Infrared Person Re-Identification
Yuqing Wu, Yongkang Ding, Liyan Zhang 0001 |
ICIC (1) | 3 |
| 2025 | Richer Semantics, Better Alignment: Aligning Visual Features with Explicit and Enriched Semantics for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VIReID) retrieves pedestrian images with the same identity across different modalities. Existing methods learn visual features solely from images, failing to align them into the modality-invariant semantic space. In this paper, we propose a novel framework, termed Richer Semantics, Better Alignment (RSBA), to align visual features with explicit and enriched semantics. Specifically, we first develop an Explicit Semantics-Guided Feature Alignment (ESFA) module, which supplements textual descriptions for cross-modality images and aligns image-text pairs within each modality, alleviating the distribution discrepancy of visual features. We then devise a Consistent Similarity-Guided Indirect Alignment (CSIA) module, which constrains the similarity between intra-modality image-text pairs to be consistent with that between inter-modality text-text pairs, indirectly aligning visual features with cross-modality semantics. Furthermore, we design a Cross-View Semantics Compensation (CVSC) module, which integrates multi-view texts and improves the image-text matching of one-to-one in ESFA and CSIA to one-to-many, further strengthening the alignment of visual features within the semantic space. Extensive experimental results on three public datasets demonstrate the effectiveness and superiority of our proposed RSBA. Neng Dong, Shuanglin Yan, Liyan Zhang 0001, Jinhui Tang 0001 |
IJCAI | 3 |
| 2025 | Generalized Person Re-identification via Hierarchical Style Mixing: Camera-Aware and Domain Fusion
Yongkang Ding, Liyan Zhang 0001 |
PRCV (16) | 3 |
| 2025 | Graph-based Consistent Reconstruction and Alignment for imbalanced text-image person re-identification
Guodong Du 0005, Tiantian Gong, Liyan Zhang 0001 |
Expert Syst. Appl. | 3 |
| 2025 | Person Parsing-Driven and Text-Guided for Cloth-Changing Person Re-IdentificationabstractWith the rapid development of the Visual Internet of Things (VIoT), person re-Identification (ReID) technology has made significant strides in research, particularly in the domains of urban security and intelligent surveillance. However, cloth-changing person re-identification (CC-ReID) poses significant challenges to traditional appearance-based methods due to frequent changes in clothing. To address this issue, this paper proposes a Person Parsing-Driven and Text-Guided for Cloth-Changing Person Re-identification (PT-ReID). This approach optimizes the image encoder of the CLIP model through a multi-branch design and leverages person parsing techniques to extract stable, clothing-invariant biological features. Additionally, the method incorporates Context Optimization (CoOp) technology, using learnable text descriptions as soft supervision to enhance cross-modal alignment. A saliency region loss mechanism is also introduced to minimize discrepancies between different feature representations, further improving the model’s ability to learn discriminative pedestrian features. Experimental results on three public datasets—PRCC, LTCC, and VC-Clothes—demonstrate that PT-ReID significantly outperforms state-of-the-art methods in Rank-1 accuracy and mean Average Precision (mAP) under cloth-changing scenarios, proving its effectiveness and robustness in handling complex clothing variations. Yongkang Ding, Yuqing Wu, Chenwei Wu 0006, Meina Qu, Liyan Zhang 0001 |
IEEE Internet Things J. | 5 |
| 2025 | Diverse Semantics-Guided Feature Alignment and Decoupling for Visible-Infrared Person Re-IdentificationabstractVisible-Infrared Person Re-Identification (VI-ReID) is a challenging task due to the large modality discrepancy between visible and infrared images, which complicates the alignment of their features into a suitable common space. Moreover, style noise, such as illumination and color contrast, reduces the identity discriminability and modality invariance of features. To address these challenges, we propose a novel Diverse Semantics-guided Feature Alignment and Decoupling (DSFAD) network to align identity-relevant features from different modalities into a textual embedding space and disentangle identity-irrelevant features within each modality. Specifically, we develop a Diverse Semantics-guided Feature Alignment (DSFA) module, which generates pedestrian descriptions with diverse sentence structures to guide the cross-modality alignment of visual features. Furthermore, to filter out style information, we propose a Semantic Margin-guided Feature Decoupling (SMFD) module, which decomposes visual features into pedestrian-related and style-related components, and then constrains the similarity between the former and the textual embeddings to be at least a margin higher than that between the latter and the textual embeddings. Additionally, to prevent the loss of pedestrian semantics during feature decoupling, we design a Semantic Consistency-guided Feature Restitution (SCFR) module, which further excavates useful information for identification from the style-related features and restores it back into the pedestrian-related features, and then constrains the similarity between the features after restitution and the textual embeddings to be consistent with that between the features before decoupling and the textual embeddings. Extensive experiments on three VI-ReID datasets demonstrate the superiority of our DSFAD. The code will be made publicly available at https://github.com/nengdong96/DSFAD. Neng Dong, Shuanglin Yan, Liyan Zhang 0001, Jinhui Tang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | CrossNet: Cross-Scene Background Subtraction Network via 3D Optical FlowabstractThis paper investigates an intriguing yet unsolved problem of cross-scene background subtraction for training only one deep model to process large-scale video streaming. We propose an end-to-end cross-scene background subtraction network via 3D optical flow, dubbed CrossNet. First, we design a new motion descriptor, hierarchical 3D optical flows (3D-HOP), to observe fine-grained motion. Then, we build a cross-modal dynamic feature filter (CmDFF) to enable the motion and appearance feature interaction. CrossNet exhibits better generalization since the proposed modules are encouraged to learn more discriminative semantic information between the foreground and the background. Furthermore, we design a loss function to balance the size diversity of foreground instances since small objects are usually missed due to training bias. Our whole background subtraction model is called Hierarchical Optical Flow Attention Model (HOFAM). Unlike most of the existing stochastic-process-based and CNN-based background subtraction models, HOFAM will avoid inaccurate online model updating, not heavily rely on scene-specific information, and well represent ambient motion in the open world. Experimental results on several well-known benchmarks demonstrate that it outperforms state-of-the-art by a large margin. The proposed framework can be flexibly integrated into arbitrary streaming media systems in a plug-and-play form. Codes are available athttps://github.com/dongzhang89/HOFAM. Dong Liang 0008, Qiong Wang 0001, Zongqi Wei, Liyan Zhang 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Graph Attention Network for Context-Aware Visual TrackingabstractSiamese-network-based trackers convert the general object tracking as a similarity matching task between a template and a search region. Using convolutional feature cross correlation (Xcorr) for similarity matching, a large number of Siamese trackers are proposed and achieved great success. However, due to the predefined size of the target feature, these trackers suffer from either retaining much background information or losing important foreground information. Moreover, the global matching between the target and search region also largely neglects the part-level structural information and the contextual information of the target. To tackle the aforementioned obstacles, in this article, we propose a simple context-aware Siamese graph attention network, which establishes part-to-part correspondence between the Siamese branches with a complete bipartite graph. The object information from the template is propagated to the search region via a graph attention mechanism. With such a design, a target-aware template input is enabled to replace the prefixed template region, which can adaptively fit the size and aspect ratio variations in different objects. Based on it, we further construct a context-aware feature matching mechanism to embed both the target and the contextual information in the search region. Experiments on challenging benchmarks including GOT-10k, TrackingNet, LaSOT, VOT2020, and OTB-100 demonstrate that the proposed SiamGAT* outperforms many state-of-the-art trackers and achieves leading performance. Code is available at: https://git.io/SiamGAT. Yanyan Shao, Dongyan Guo, Zhenhua Wang 0003, Liyan Zhang 0001, Jianhua Zhang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Global-Local Multiple Granularity Learning for Cross-Modality Visible-Infrared Person ReidentificationabstractCross-modality visible-infrared person reidentification (VI-ReID), which aims to retrieve pedestrian images captured by both visible and infrared cameras, is a challenging but essential task for smart surveillance systems. The huge barrier between visible and infrared images has led to the large cross-modality discrepancy and intraclass variations. Most existing VI-ReID methods tend to learn discriminative modality-sharable features based on either global or part-based representations, lacking effective optimization objectives. In this article, we propose a novel global-local multichannel (GLMC) network for VI-ReID, which can learn multigranularity representations based on both global and local features. The coarse- and fine-grained information can complement each other to form a more discriminative feature descriptor. Besides, we also propose a novel center loss function that aims to simultaneously improve the intraclass cross-modality similarity and enlarge the interclass discrepancy to explicitly handle the cross-modality discrepancy issue and avoid the model fluctuating problem. Experimental results on two public datasets have demonstrated the superiority of the proposed method compared with state-of-the-art approaches in terms of effectiveness. Liyan Zhang 0001, Guodong Du 0005, Fan Liu 0003, Huawei Tu, Xiangbo Shu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Improvement of the Clothes-Changing Person Re-identification with Multiple Loss FunctionsabstractPerson re-identification plays a crucial role in various domains, such as intelligent transportation, public safety, and smart cities. However, the current person re-identification methods are unable to address the challenges posed by long-term clothes-changing scenarios. As a result, clothes-changing person re-identification has gained increasing attention and is considered a highly challenging task. Most existing works in this area mainly focus on learning from multi-modal information such as outlines and sketches, while overlooking valuable features in the original RGB images. In this paper, we propose a novel method that conducts in-depth research on RGB images. We introduce the Scaled Triplet Loss (STL) for metric learning, which helps similar samples be closer in the feature space and encourages the model to pay more attention to the distances between samples that are difficult to distinguish. In addition, over an existing loss function, we propose Augmentation Clothing-irrelevant Features Loss (ACL), which better exploits features unrelated to clothing in RGB images, resulting in more advanced model performance. Extensive experiments have demonstrated the performance of our model surpasses other state-of-the-art methods in two real-world datasets. Yongkang Ding, Rui Mao 0014, Hanyue Zhu, Liyan Zhang 0001 |
CSCWD | 4 |
| 2024 | Discriminative Pedestrian Features and Gated Channel Attention for Clothes-Changing Person Re-IdentificationabstractIn public safety and social life, the task of Clothes-Changing Person Re-Identification (CC-ReID) has become increasingly significant. However, this task faces considerable challenges due to appearance changes caused by clothing alterations. Addressing this issue, this paper proposes an innovative method for disentangled feature extraction, effectively extracting discriminative features from pedestrian images that are invariant to clothing. This method leverages pedestrian parsing techniques to identify and retain features closely associated with individual identity while disregarding the variable nature of clothing attributes. Furthermore, this study introduces a gated channel attention mechanism, which, by adjusting the network’s focus, aids the model in more effectively learning and emphasizing features critical for pedestrian identity recognition. Extensive experiments conducted on two standard CC-ReID datasets validate the effectiveness of the proposed approach, with performance surpassing current leading solutions. The Top-1 accuracy under clothing change scenarios on the PRCC and VC-Clothes datasets reached 64.8% and 83.7%, respectively. Yongkang Ding, Rui Mao 0014, Hanyue Zhu, Anqi Wang 0010, Liyan Zhang 0001 |
ICME | 5 |
| 2024 | Multidimensional Semantic Disentanglement Network for Clothes-Changing Person Re-IdentificationabstractThis study focuses on the Clothes-Changing Person Re-Identification (CC-ReID) problem, aiming to achieve precise recognition of the same pedestrian despite changes in attire. Despite some progress in this field, challenges persist in maintaining pedestrian identity consistency due to variations in clothing, leading to recognition disturbances. To address this, we propose a novel Multidimensional Semantic Disentanglement Network (MSD-Net). This network enhances the recognition capability for non-clothing areas by reducing reliance on clothing features and integrating discriminative and global features. Specifically, we employ semantic segmentation maps for pedestrian feature disentanglement, combined with RGB images, to effectively erase clothing features and consequently enhance focus on non-clothing areas. Additionally, we introduce a method to convert pedestrian semantic segmentation maps into dual-precision feature maps, utilizing a spatial attention mechanism to proactively learn distinctive pedestrian features, thereby further improving model performance. Extensive experiments on two standard CC-ReID datasets validate the effectiveness of our approach, outperforming existing state-of-the-art solutions. On the PRCC and VC-Clothes datasets, our model achieves Top-1 accuracies of 65.3% and 84.1%, respectively, in clothes-changing scenarios. Yongkang Ding, Anqi Wang 0010, Liyan Zhang 0001 |
ICMR | 3 |
| 2024 | Bottom-up color-independent alignment learning for text-image person re-identification
Guodong Du 0001, Hanyue Zhu, Liyan Zhang 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Contrastive completing learning for practical text-image person ReID: Robuster and cheaper
Guodong Du 0005, Tiantian Gong, Liyan Zhang 0001 |
Expert Syst. Appl. | 3 |
| 2024 | Work like a doctor: Unifying scan localizer and dynamic generator for automated computed tomography report generation
Yuhao Tang, Haichen Yang, Liyan Zhang 0001 |
Expert Syst. Appl. | 3 |
| 2024 | NDAM-YOLOseg: a real-time instance segmentation model based on multi-head attention mechanism
Chengang Dong, Yuhao Tang, Liyan Zhang 0001 |
Multim. Syst. | 3 |
| 2024 | Disentangled body features for clothing change person re-identification
Yongkang Ding, Yinghao Wu, Anqi Wang 0010, Tiantian Gong, Liyan Zhang 0001 |
Multim. Tools Appl. | 5 |
| 2024 | Higher efficient YOLOv7: a one-stage method for non-salient object detection
Chengang Dong, Yuhao Tang, Liyan Zhang 0001 |
Multim. Tools Appl. | 3 |
| 2024 | Fashion item captioning via grid-relation self-attention and gated-enhanced decoder
Yuhao Tang, Liyan Zhang 0001 |
Multim. Tools Appl. | 2 |
| 2024 | Multi-Scale Explicit Matching and Mutual Subject Teacher Learning for Generalizable Person Re-IdentificationabstractDomain generalization in person re-identification (DG-ReID) stands out as the most challenging task and practically important branch in the ReID field, which enables the direct deployment of pre-trained models in unseen and real scenarios. Recent works have made significant efforts in this task via the image-matching paradigm, which searches for the local correspondences in the feature maps. A common practice of employing pixel-wise matching is typically used to ensure efficient matching. This, however, makes the matching susceptible to deviations caused by identity-irrelevant pixel features. On the other hand, patch-wise matching also demonstrates that it will disregard the spatial orientation of pedestrians and amplify the impact of noise. To address the mentioned issues, this paper proposes the Multi-Scale Query-Adaptive Convolution (QAConv-MS) framework, which encodes patches in the feature maps to pixels using template kernels of various scales. This enables the matching process to enjoy broader receptive fields and robustness to orientations and noises. To stabilize the matching process and facilitate the independent learning of each sub-kernel within the template kernels to capture diverse local patterns, we propose the OrthoGonal Norm (OGNorm), which consists of two orthogonal normalizations. We also present Mutual Subject Teacher Learning (MSTL) to address the potential issues of overconfidence and overfitting in the model. MSTL allows two models to individually select the most challenging data for training, resulting in more dependable soft labels that can provide mutual supervision. Extensive experiments conducted in both single-source and multi-source setups offer compelling evidence of our framework’s generalization and competitiveness. Kaixiang Chen, Pengfei Fang, Liyan Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Camera-Aware Recurrent Learning and Earth Mover's Test-Time Adaption for Generalizable Person Re-IdentificationabstractDomain generalization in person re-identification (ReID) aims to design a generalizable model, which is trained under the supervision of a set of labeled source domains and can be directly deployed on unknown domains. Existing approaches simply treat each identity as a distinct class and ignore the differences among cameras. We argue that the camera information is crucial for learning discriminative representations, as people’s behavior usually varies between cameras. In this paper, we present Multi-Centroid Memory (MCM) to capture different camera information for each identity and Soft Triple Hard (ST-Hard) loss to align the information of the same identity across cameras. Furthermore, in contrast to the traditional approaches of training a single model using a parallel training mechanism, we propose the Recurrent Implicit Lifelong Learning (RILL) that feeds the source domains into the model in a continuous loop to train an expert for each domain. To make each expert further generalized to other source domains, during the training on the current domain, RILL adopts a style replay-based method to simulate the training of the previous domain, encouraging each domain’s expert to extract generalizable features. We also present Earth Mover’s Test-time Adaption (EMTA) to be used in conjunction with RILL, which enables source domains that are more similar to the test domain to play a more significant role in the test. This is achieved by our proposed Earth Mover’s Similarity (EMS), which helps model the similarities between the source and test domains. Extensive experiments on two evaluation protocols fully demonstrate our framework’s generalization and competitiveness. Kaixiang Chen, Tiantian Gong, Liyan Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Erasing, Transforming, and Noising Defense Network for Occluded Person Re-IdentificationabstractOcclusion perturbation presents a significant challenge in person re-identification (re-ID), and existing methods that rely on external visual cues require additional computational resources and only consider the issue of missing information caused by occlusion. In this paper, we propose a simple yet effective framework, termed Erasing, Transforming, and Noising Defense Network (ETNDNet), which treats occlusion as a noise disturbance and solves occluded person re-ID from the perspective of adversarial defense. In the proposed ETNDNet, we introduce three strategies: Firstly, we randomly erase the feature map to create an adversarial representation with incomplete information, enabling adversarial learning of identity loss to protect the re-ID system from the disturbance of missing information. Secondly, we introduce random transformations to simulate the position misalignment caused by occlusion, training the extractor and classifier adversarially to learn robust representations immune to misaligned information. Thirdly, we perturb the feature map with random values to address noisy information introduced by obstacles and non-target pedestrians, and employ adversarial gaming in the re-ID system to enhance its resistance to occlusion noise. Without bells and whistles, ETNDNet has three key highlights: (i) it does not require any external modules with parameters, (ii) it effectively handles various issues caused by occlusion from obstacles and non-target pedestrians, and (iii) it designs the first GAN-based adversarial defense paradigm for occluded person re-ID. Extensive experiments on six public datasets fully demonstrate the effectiveness, superiority, and practicality of the proposed ETNDNet. The code will be released at https://github.com/nengdong96/ETNDNet. Neng Dong, Liyan Zhang 0001, Shuanglin Yan, Hao Tang 0007, Jinhui Tang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Enhanced Invariant Feature Joint Learning via Modality-Invariant Neighbor Relations for Cross-Modality Person Re-IdentificationabstractCross-modality visible-Infrared person re-identification (cm-ReID) is extremely challenging due to the huge modality discrepancy between RGB and IR modalities. Existing methods focus on the sample features themselves, trying to learn modality-invariant features and perform alignment to reduce the modality discrepancy in dataset-level, while the negative impact of specific features and the identity optimization are not specifically addressed. Moreover, most methods that only extracts modality-invariant appearance features cannot acquire enough discriminative matching information for identifying different persons since the information of invariant features is limited compared with original features. Accordingly, in this paper, we propose a Enhanced Invariant Feature Joint Learning Framework (EIFJLF) for cm-ReID to handle the above problems. First, we propose a specific feature confusion baseline with a novel channel-blended transformation, which confuses the visible color and infrared spectrum to alleviate the influence of specific features, so that model pays more attention to other discriminative invariant features. Second, we present an adaptive heterogeneous center loss for better identity optimization. The adaptive margin of the loss makes samples not too close to the center, avoiding losing effectiveness too early and overfitting meantime further boosting performance. Finally, we design a novel similarity feature refinement module to utilize intra-modality relations and achieve invariant information compensation. Intra-modality relations are valuable built-in invariant features and we model these relations with similarity between samples into affinities and then update the original features to achieve information compensation. EIFJLF works for more informative invariant feature learning and more stable alignment. For cm-ReID, our work is a brand new attempt. Extensive experimental results on two standard benchmarks have demonstrated superiority of the proposed method compared with state-of-the-art methods. Guodong Du 0005, Liyan Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Deblurring Videos Using Spatial-Temporal Contextual Transformer With Feature PropagationabstractWe present a simple and effective approach to explore both local spatial-temporal contexts and non-local temporal information for video deblurring. First, we develop an effective spatial-temporal contextual transformer to explore local spatial-temporal contexts from videos. As the features extracted by the spatial-temporal contextual transformer does not model the non-local temporal information of video well, we then develop a feature propagation method to aggregate useful features from the long-range frames so that both local spatial-temporal contexts and non-local temporal information can be better utilized for video deblurring. Finally, we formulate the spatial-temporal contextual transformer with the feature propagation into a unified deep convolutional neural network (CNN) and train it in an end-to-end manner. We show that using the spatial-temporal contextual transformer with the feature propagation is able to generate useful features and makes the deep CNN model more compact and effective for video deblurring. Extensive experimental results show that the proposed method performs favorably against state-of-the-art ones on the benchmark datasets in terms of accuracy and model parameters. Liyan Zhang 0001, Boming Xu, Zhongbao Yang, Jinshan Pan |
IEEE Trans. Image Process. | 1 |
| 2024 | Coupling Global Context and Local Contents for Weakly-Supervised Semantic SegmentationabstractThanks to the advantages of the friendly annotations and the satisfactory performance, weakly-supervised semantic segmentation (WSSS) approaches have been extensively studied. Recently, the single-stage WSSS (SS-WSSS) was awakened to alleviate problems of the expensive computational costs and the complicated training procedures in multistage WSSS. However, the results of such an immature model suffer from problems of background incompleteness and object incompleteness. We empirically find that they are caused by the insufficiency of the global object context and the lack of local regional contents, respectively. Under these observations, we propose an SS-WSSS model with only the image-level class label supervisions, termed weakly supervised feature coupling network (WS-FCN), which can capture the multiscale context formed from the adjacent feature grids, and encode the fine-grained spatial information from the low-level features into the high-level ones. Specifically, a flexible context aggregation (FCA) module is proposed to capture the global object context in different granular spaces. Besides, a semantically consistent feature fusion (SF2) module is proposed in a bottom-up parameter-learnable fashion to aggregate the fine-grained local contents. Based on these two modules, WS-FCN lies in a self-supervised end-to-end training fashion. Extensive experimental results on the challenging PASCAL VOC 2012 and MS COCO 2014 demonstrate the effectiveness and efficiency of WS-FCN, which can achieve state-of-the-art results by 65.02% and 64.22% mIoU on PASCAL VOC 2012 val set and test set, 34.12% mIoU on MS COCO 2014 val set, respectively. The code and weight have been released at:WS-FCN. Liyan Zhang 0001, Jinhui Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Image-Specific Information Suppression and Implicit Local Alignment for Text-Based Person SearchabstractText-based person search (TBPS) is a challenging task that aims to search pedestrian images with the same identity from an image gallery given a query text. In recent years, TBPS has made remarkable progress, and state-of-the-art (SOTA) methods achieve superior performance by learning local fine-grained correspondence between images and texts. However, most existing methods rely on explicitly generated local parts to model fine-grained correspondence between modalities, which is unreliable due to the lack of contextual information or the potential introduction of noise. Moreover, the existing methods seldom consider the information inequality problem between modalities caused by image-specific information. To address these limitations, we propose an efficient joint multilevel alignment network (MANet) for TBPS, which can learn aligned image/text feature representations between modalities at multiple levels, and realize fast and effective person search. Specifically, we first design an image-specific information suppression (ISS) module, which suppresses image background and environmental factors by relation-guided localization (RGL) and channel attention filtration (CAF), respectively. This module effectively alleviates the information inequality problem and realizes the alignment of information volume between images and texts. Second, we propose an implicit local alignment (ILA) module to adaptively aggregate all pixel/word features of image/text to a set of modality-shared semantic topic centers and implicitly learn the local fine-grained correspondence between modalities without additional supervision and cross-modal interactions. Also, a global alignment (GA) is introduced as a supplement to the local perspective. The cooperation of global and local alignment modules enables better semantic alignment between modalities. Extensive experiments on multiple databases demonstrate the effectiveness and superiority of our MANet. Shuanglin Yan, Hao Tang 0007, Liyan Zhang 0001, Jinhui Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Prototype-guided Cross-modal Completion and Alignment for Incomplete Text-based Person Re-identificationabstractTraditional text-based person re-identification (ReID) techniques heavily rely on fully matched multi-modal data, which is an ideal scenario. However, due to inevitable data missing and corruption during the collection and processing of cross-modal data, the incomplete data issue is usually met in real-world applications. Therefore, we consider a more practical task termed the incomplete text-based ReID task, where person images and text descriptions are not completely matched and contain partially missing modality data. To this end, we propose a novel Prototype-guided Cross-modal Completion and Alignment (PCCA) framework to handle the aforementioned issues for incomplete text-based ReID. Specifically, we cannot directly retrieve person images based on a text query on missing modality data. Therefore, we propose the cross-modal nearest neighbor construction strategy for missing data by computing the cross-modal similarity between existing images and texts, which provides key guidance for the completion of missing modal features. Furthermore, to efficiently complete the missing modal features, we construct the relation graphs with the aforementioned cross-modal nearest neighbor sets of missing modal data and the corresponding prototypes, which can further enhance the generated missing modal features. Additionally, for tighter fine-grained alignment between images and texts, we raise a prototype-aware cross-modal alignment loss that can effectively reduce the modality heterogeneity gap for better fine-grained alignment in common space. Extensive experimental results on several benchmarks with different missing ratios amply demonstrate that our method can consistently outperform state-of-the-art text-image ReID approaches. Tiantian Gong, Guodong Du 0005, Yongkang Ding, Liyan Zhang 0001 |
ACM Multimedia | 5 |
| 2023 | Improving fashion captioning via attribute-based alignment and multi-level language model
Yuhao Tang, Liyan Zhang 0001 |
Appl. Intell. | 2 |
| 2023 | Augmented FCN: rethinking context modeling for semantic segmentation
Liyan Zhang 0001, Jinhui Tang 0001 |
Sci. China Inf. Sci. | 2 |
| 2023 | Multi-Granularity Anchor-Contrastive Representation Learning for Semi-Supervised Skeleton-Based Action RecognitionabstractIn the semi-supervised skeleton-based action recognition task, obtaining more discriminative information from both labeled and unlabeled data is a challenging problem. As the current mainstream approach, contrastive learning can learn more representations of augmented data, which can be considered as the pretext task of action recognition. However, such a method still confronts three main limitations: 1) It usually learns global-granularity features that cannot well reflect the local motion information. 2) The positive/negative pairs are usually pre-defined, some of which are ambiguous. 3) It generally measures the distance between positive/negative pairs only within the same granularity, which neglects the contrasting between the cross-granularity positive and negative pairs. Toward these limitations, we propose a novel Multi-granularity Anchor-Contrastive representation Learning (dubbed as MAC-Learning) to learn multi-granularity representations by conducting inter- and intra-granularity contrastive pretext tasks on the learnable and structural-link skeletons among three types of granularities covering local, context, and global views. To avoid the disturbance of ambiguous pairs from noisy and outlier samples, we design a more reliable Multi-granularity Anchor-Contrastive Loss (dubbed as MAC-Loss) that measures the agreement/disagreement between high-confidence soft-positive/negative pairs based on the anchor graph instead of the hard-positive/negative pairs in the conventional contrastive loss. Extensive experiments on both NTU RGB+D and Northwestern-UCLA datasets show that the proposed MAC-Learning outperforms existing competitive methods in semi-supervised skeleton-based action recognition tasks. Xiangbo Shu, Binqian Xu, Liyan Zhang 0001, Jinhui Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Progressive Instance-Aware Feature Learning for Compositional Action RecognitionabstractIn order to enable the model to generalize to unseen "action-objects" (compositional action), previous methods encode multiple pieces of information (i.e., the appearance, position, and identity of visual instances) independently and concatenate them for classification. However, these methods ignore the potential supervisory role of instance information (i.e., position and identity) in the process of visual perception. To this end, we present a novel framework, namely Progressive Instance-aware Feature Learning (PIFL), to progressively extract, reason, and predict dynamic cues of moving instances from videos for compositional action recognition. Specifically, this framework extracts features from foreground instances that are likely to be relevant to human actions (Position-aware Appearance Feature Extraction in Section III-B1), performs identity-aware reasoning among instance-centric features with semantic-specific interactions (Identity-aware Feature Interaction in Section III-B2), and finally predicts instances' position from observed states to force the model into perceiving their movement (Semantic-aware Position Prediction in Section III-B3). We evaluate our approach on two compositional action recognition benchmarks, namely, Something-Else and IKEA-Assembly. Our approach achieves consistent accuracy gain beyond off-the-shelf action recognition algorithms in terms of both ground truth and detected position of instances. Rui Yan 0010, Lingxi Xie, Xiangbo Shu, Liyan Zhang 0001, Jinhui Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Debiased Contrastive Curriculum Learning for Progressive Generalizable Person Re-IdentificationabstractDomain generalization (DG) in person re-identification (ReID) is an extremely challenging but essential task, which aims to learn a generalizable model over multiple labeled source domains that can perform well on unseen target domains. Most existing DG strategies in ReID directly aggregate multiple source data together for training, incurring a large inter-domain bias and unstable model optimization that lead the model apt to overfitting domain bias and the model training more time-consuming respectively, thus hampering the generalization and convergence speed of the model. To tackle the aforementioned issues, inspired by Curriculum Learning that mimics the process of human lifelong DG learning (from easy to hard), we put forward a novel Debiased Contrastive Curriculum Learning (DCCL) strategy for DG ReID, which is designed to incrementally enhance generalization in an easy-to-hard training way that can continuously accumulate learning experience to make learning in unknown domains easier and effectively eliminate the domain bias to help the model learn rich domain-invariant discriminative features, thereby strengthening generalization and accelerating convergence for the model. In addition, to simultaneously learn class-level and instance-level discriminative representations, we raise a non-parametric hybrid contrastive loss to equip the DCCL model. We also particularly design an inter-domain mix module to variegate the features of the newly added source domain at each stage of DCCL, further establishing the advantages of DCCL. Extensive experimental results on four public ReID benchmarks fully demonstrate that our DCCL can effectively strengthen the generalization capacity of the model to unseen domains and outperform the state-of-the-art methods. Tiantian Gong, Kaixiang Chen, Liyan Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Describe Fashion Products via Local Sparse Self-Attention Mechanism and Attribute-Based Re-Sampling StrategyabstractIn this paper, we convert conventional image captioning task into a new paradigm in the fashion domain, fashion item captioning. This task requires the ability of associating a group of item angles and generating a longer and more fine-grained description. To link several images in a more efficient way, we propose a local sparse self-attention mechanism(LSAM), which only enables one region to interact with those in the adjacent area and allows a lighter architecture with fewer layers. To capture more subtle details of the item, an attribute-based re-sampling strategy(ARS) is introduced to enhance the learning on those low-frequency but content-related attribute words. Besides, existing fashion datasets are limited in the quality of annotations and single fashion style. To bridge this gap, we propose a novel dataset for fashion item captioning, termed Fashion Item Captioning Dataset(FICD). FICD provides a meaningful complement to the existing fashion datasets, which comprises 294K images, 62K real-world product descriptions with diverse linguistic styles, rich attributes and categories. It is worth mentioning that we also annotate an attribute-level sentence for each item, which enables the FICD to be used not only for fashion item captioning but also for other fashion-related tasks. Moreover, the complex structure is even more restrictive for application to real-world scenarios. Without bells and whistles, our framework is simply designed with an end-to-end manner. Extensive experiments demonstrate the effectiveness of our LSAM-ARS. More remarkably, LSAM-ARS achieves state-of-the-art performance on FACAD and FICD datasets, with the CIDEr-D score being increased from 65.4% to 81.8%, 69.8% to 77.4%, respectively. Yuhao Tang, Liyan Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Centralized Feature Pyramid for Object DetectionabstractThe visual feature pyramid has shown its superiority in both effectiveness and efficiency in a variety of applications. However, current methods overly focus on inter-layer feature interactions while disregarding the importance of intra-layer feature regulation. Despite some attempts to learn a compact intra-layer feature representation with the use of attention mechanisms or vision transformers, they overlook the crucial corner regions that are essential for dense prediction tasks. To address this problem, we propose a Centralized Feature Pyramid (CFP) network for object detection, which is based on a globally explicit centralized feature regulation. Specifically, we first propose a spatial explicit visual center scheme, where a lightweight MLP is used to capture the globally long-range dependencies, and a parallel learnable visual center mechanism is used to capture the local corner regions of the input images. Based on this, we then propose a globally centralized regulation for the commonly-used feature pyramid in a top-down fashion, where the explicit visual center information obtained from the deepest intra-layer feature is used to regulate frontal shallow features. Compared to the existing feature pyramids, CFP not only has the ability to capture the global long-range dependencies but also efficiently obtain an all-round yet discriminative feature representation. Experimental results on the challenging MS-COCO validate that our proposed CFP can achieve consistent performance gains on the state-of-the-art YOLOv5 and YOLOX object detection baselines. Yu Quan, Liyan Zhang 0001, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | CLIP-Driven Fine-Grained Text-Image Person Re-IdentificationabstractText-Image Person Re-identification (TIReID) aims to retrieve the image corresponding to the given text query from a pool of candidate images. Existing methods employ prior knowledge from single-modality pre-training to facilitate learning, but lack multi-modal correspondence information. Vision-Language Pre-training, such as CLIP (Contrastive Language-Image Pretraining), can address the limitation. However, CLIP falls short in capturing fine-grained information, thereby not fully leveraging its powerful capacity in TIReID. Besides, the popular explicit local matching paradigm for mining fine-grained information heavily relies on the quality of local parts and cross-modal inter-part interaction/guidance, leading to intra-modal information distortion and ambiguity problems. Accordingly, in this paper, we propose a CLIP-driven Fine-grained information excavation framework (CFine) to fully utilize the powerful knowledge of CLIP for TIReID. To transfer the multi-modal knowledge effectively, we conduct fine-grained information excavation to mine modality-shared discriminative details for global alignment. Specifically, we propose a multi-level global feature learning (MGF) module that fully mines the discriminative local information within each modality, thereby emphasizing identity-related discriminative clues through enhanced interaction between global image (text) and informative local patches (words). MGF generates a set of enhanced global features for later inference. Furthermore, we design cross-grained feature refinement (CFR) and fine-grained correspondence discovery (FCD) modules to establish cross-modal correspondence at both coarse and fine-grained levels (image-word, sentence-patch, word-patch), ensuring the reliability of informative local patches/words. CFR and FCD are removed during inference to optimize computational efficiency. Extensive experiments on multiple benchmarks demonstrate the superior performance of our method in TIReID. Shuanglin Yan, Neng Dong, Liyan Zhang 0001, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | Joint Classification and Regression for Visual Tracking with Fully Convolutional Siamese NetworksabstractAbstract Visual tracking of generic objects is one of the fundamental but challenging problems in computer vision. Here, we propose a novel fully convolutional Siamese network to solve visual tracking by directly predicting the target bounding box in an end-to-end manner. We first reformulate the visual tracking task as two subproblems: a classification problem for pixel category prediction and a regression task for object status estimation at this pixel. With this decomposition, we design a simple yet effective Siamese architecture based classification and regression framework, termed SiamCAR, which consists of two subnetworks: a Siamese subnetwork for feature extraction and a classification-regression subnetwork for direct bounding box prediction. Since the proposed framework is both proposal- and anchor-free, SiamCAR can avoid the tedious hyper-parameter tuning of anchors, considerably simplifying the training. To demonstrate that a much simpler tracking framework can achieve superior tracking results, we conduct extensive experiments and comparisons with state-of-the-art trackers on a few challenging benchmarks. Without bells and whistles, SiamCAR achieves leading performance with a real-time speed. Furthermore, the ablation study validates that the proposed framework is effective with various backbone networks, and can benefit from deeper networks. Code is available at https://github.com/ohhhyeahhh/SiamCAR . Dongyan Guo, Yanyan Shao, Zhenhua Wang 0003, Chunhua Shen, Liyan Zhang 0001, Shengyong Chen |
Int. J. Comput. Vis. | 6 |
| 2022 | CTNet: Context-Based Tandem Network for Semantic SegmentationabstractContextual information has been shown to be powerful for semantic segmentation. This work proposes a novel Context-based Tandem Network (CTNet) by interactively exploring the spatial contextual information and the channel contextual information, which can discover the semantic context for semantic segmentation. Specifically, the Spatial Contextual Module (SCM) is leveraged to uncover the spatial contextual dependency between pixels by exploring the correlation between pixels and categories. Meanwhile, the Channel Contextual Module (CCM) is introduced to learn the semantic features including the semantic feature maps and class-specific features by modeling the long-term semantic dependence between channels. The learned semantic features are utilized as the prior knowledge to guide the learning of SCM, which can make SCM obtain more accurate long-range spatial dependency. Finally, to further improve the performance of the learned representations for semantic segmentation, the results of the two context modules are adaptively integrated to achieve better results. Extensive experiments are conducted on four widely-used datasets, i.e., PASCAL-Context, Cityscapes, ADE20K and PASCAL VOC2012. The results demonstrate the superior performance of the proposed CTNet by comparison with several state-of-the-art methods. The source code and models are available at https://github.com/syp2ysy/CTNet. Zechao Li, Yanpeng Sun, Liyan Zhang 0001, Jinhui Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Spatiotemporal Co-Attention Recurrent Neural Networks for Human-Skeleton Motion PredictionabstractHuman motion prediction aims to generate future motions based on the observed human motions. Witnessing the success of Recurrent Neural Networks (RNN) in modeling sequential data, recent works utilize RNNs to model human-skeleton motions on the observed motion sequence and predict future human motions. However, these methods disregard the existence of the spatial coherence among joints and the temporal evolution among skeletons, which reflects the crucial characteristics of human motions in spatiotemporal space. To this end, we propose a novel Skeleton-Joint Co-Attention Recurrent Neural Networks (SC-RNN) to capture the spatial coherence among joints, and the temporal evolution among skeletons simultaneously on a skeleton-joint co-attention feature map in spatiotemporal space. First, a skeleton-joint feature map is constructed as the representation of the observed motion sequence. Second, we design a new Skeleton-Joint Co-Attention (SCA) mechanism to dynamically learn a skeleton-joint co-attention feature map of this skeleton-joint feature map, which can refine the useful observed motion information to predict one future motion. Third, a variant of GRU embedded with SCA collaboratively models the human-skeleton motion and human-joint motion in spatiotemporal space by regarding the skeleton-joint co-attention feature map as the motion context. Experimental results of human motion prediction demonstrate that the proposed method outperforms the competing methods. Xiangbo Shu, Liyan Zhang 0001, Guo-Jun Qi, Wei Liu 0005, Jinhui Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Coherence Constrained Graph LSTM for Group Activity RecognitionabstractThis work aims to address the group activity recognition problem by exploring human motion characteristics. Traditional methods hold that the motions of all persons contribute equally to the group activity, which suppresses the contributions of some relevant motions to the whole activity while overstating some irrelevant motions. To address this problem, we present a Spatio-Temporal Context Coherence (STCC) constraint and a Global Context Coherence (GCC) constraint to capture the relevant motions and quantify their contributions to the group activity, respectively. Based on this, we propose a novel Coherence Constrained Graph LSTM (CCG-LSTM) with STCC and GCC to effectively recognize group activity, by modeling the relevant motions of individuals while suppressing the irrelevant motions. Specifically, to capture the relevant motions, we build the CCG-LSTM with a temporal confidence gate and a spatial confidence gate to control the memory state updating in terms of the temporally previous state and the spatially neighboring states, respectively. In addition, an attention mechanism is employed to quantify the contribution of a certain motion by measuring the consistency between itself and the whole activity at each time step. Finally, we conduct experiments on two widely-used datasets to illustrate the effectiveness of the proposed CCG-LSTM compared with the state-of-the-art methods. Jinhui Tang 0001, Xiangbo Shu, Rui Yan 0010, Liyan Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Causal Inference with Knowledge Distilling and Curriculum Learning for Unbiased VQAabstractRecently, many Visual Question Answering (VQA) models rely on the correlations between questions and answers yet neglect those between the visual information and the textual information. They would perform badly if the handled data distribute differently from the training data (i.e., out-of-distribution (OOD) data). Towards this end, we propose a two-stage unbiased VQA approach that addresses the unbiased issue from a causal perspective. In the causal inference stage, we mark the spurious correlation on the causal graph, explore the counterfactual causality, and devise a causal target based on the inherent correlations between the conventional and counterfactual VQA models. In the distillation stage, we introduce the causal target into the training process and leverages distilling as well as curriculum learning to capture the unbiased model. Since Causal Inference with Knowledge Distilling and Curriculum Learning (CKCL) reinforces the contribution of the visual information and eliminates the impact of the spurious correlation by distilling the knowledge in causal inference to the VQA model, it contributes to the good performance on both the standard data and out-of-distribution data. The extensive experimental results on VQA-CP v2 dataset demonstrate the superior performance of the proposed method compared to the state-of-the-art (SotA) methods. Yonghua Pan, Zechao Li, Liyan Zhang 0001, Jinhui Tang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Graph Attention TrackingabstractSiamese network based trackers formulate the visual tracking task as a similarity matching problem. Almost all popular Siamese trackers realize the similarity learning via convolutional feature cross-correlation between a target branch and a search branch. However, since the size of target feature region needs to be pre-fixed, these cross-correlation base methods suffer from either reserving much adverse background information or missing a great deal of foreground information. Moreover, the global matching be-tween the target and search region also largely neglects the target structure and part-level information.In this paper, to solve the above issues, we propose a simple target-aware Siamese graph attention network for general object tracking. We propose to establish part-to-part correspondence between the target and the search region with a complete bipartite graph, and apply the graph attention mechanism to propagate target information from the template feature to the search feature. Further, instead of using the pre-fixed region cropping for template-feature-area selection, we investigate a target-aware area selection mechanism to fit the size and aspect ratio variations of different objects. Experiments on challenging benchmarks including GOT-10k, UAV123, OTB-100 and LaSOT demonstrate that the proposed SiamGAT outperforms many state-of-the-art trackers and achieves leading performance. Code is available at: https://git.io/SiamGAT Dongyan Guo, Yanyan Shao, Zhenhua Wang 0003, Liyan Zhang 0001, Chunhua Shen |
CVPR | 5 |
| 2021 | Dense Face Detection via High-level Context MiningabstractThe appearance degradation caused by low resolution is the core problem of small face detection. Therefore, a natural approach is to assemble information from the context. This paper focuses on how to use high-level contextual information to improve the abilities of anchor-based detectors to detect dense and degenerate faces. We tap the spatial contextual information on the overall view based on the density map, and propose the prior of face co-occurrence for inferred bounding-boxes coordination. We also propose score-size-specific non-maximum suppression to replace the traditional non-maximum suppression at the end of anchor-based detectors. According to the inferred face boxes' quantity, score and size, the proposed synthetical solution reduces false positives and increases true positives. Our method does not require additional training, which is model-independent and can be embedded into existing face detectors. We also propose a dataset - Crowd Face for face detection, which is full of challenges. We expect to supply enough samples to highlight the difficulties of detecting dense and degenerate faces. We embed our proposed methods into state-of-the-art face detectors on massively benchmarked face datasets. Compared with the prior art on the WIDER FACE hard set, our method increase an Average Precision of 0.1 %-1.3%. On Crowd Face, it increases an Average Precision of 1 % – 6%. Dataset is available on: https://github.com/QxGeng/Crowd-Face. Qixiang Geng, Dong Liang 0008, Huiyu Zhou 0001, Liyan Zhang 0001, Ningzhong Liu |
FG | 4 |
| 2021 | Nlkd: Using Coarse Annotations For Semantic Segmentation Based on Knowledge DistillationabstractModern supervised learning relies on a large amount of training data, yet there are many noisy annotations in real datasets. For semantic segmentation tasks, pixel-level annotation noise is typically located at the edge of an object, while pixels within objects are fine-annotated. We argue the coarse annotations can provide instructive supervised information to guide model training rather than be discarded. This paper proposes a noise learning framework based on knowledge distillation NLKD, to improve segmentation performance on unclean data. It utilizes a teacher network to guide the student network that constitutes the knowledge distillation process. The teacher and student generate the pseudo-labels and jointly evaluate the quality of annotations to generate weights for each sample. Experiments demonstrate the effectiveness of NLKD, and we observe better performance with boundary-aware teacher networks and evaluation metrics. Furthermore, the proposed approach is model-independent and easy to implement, appropriate for integration with other tasks and models. Dong Liang 0008, Liyan Zhang 0001, Ningzhong Liu, Mingqiang Wei |
ICASSP | 4 |
| 2021 | Cross Scene Video Foreground Segmentation Via Co-Occurrence Probability Oriented Supervised and Unsupervised Model InteractionabstractUsing only one deep model for cross scene video foreground segmentation is still very challenging because existing methods are scene-dependent, which restricts the consistent segmentation. In this paper, we propose a cross scene video foreground segmentation framework to extend the generalization capability of those supervised model depending on scene-specific training. The proposed framework flexibly utilizes three well-trained supervised models as guidance to yield a coarse segmentation mask. The co-occurrence probability-based unsupervised background subtraction model is introduced to achieve scene adaptation in the plug and play style without any fine-tuning and labels. Experimental results on LIMU and CDNet2014 datasets validate our framework outperforms the state-of-the-art supervised/unsupervised approaches that participate in the comparison. Experiments also show the training efficiency-related improvements – when introducing the guidance models, the demand for quantity and quality of training samples to train the unsupervised model is reduced. Codes https://github.com/MeteoorLiu/Venus/tree/MeteoorLiu-SUMC Dong Liang 0008, Bin Kang, Liyan Zhang 0001, Ningzhong Liu |
ICASSP | 5 |
| 2021 | Rectifying Pseudo Label By Mutual Disagreement Learning For Unsupervised Domain Adaptation Person Re-IdentificationabstractUnsupervised domain adaptation person re-identification(re-ID) aims to transfer knowledge from the labeled source domain to unlabeled target domain, which is still a challenging task due to the large domain discrepancy. The clustering-based methods maintain advanced performance, generating pseudo-labels for unlabeled target domain images by clustering. However, not making full use of all valuable images and label noise derived from imperfect clustering results dramatically impact further performance improvement. To alleviate the above two problems, we propose a novel mutual disagreement learning(MDL) framework. We attach some of outliers(unclustered samples) with small loss to training process in an adversarial strategy manner. To rectify label noise, we train two networks and their momentum-based moving average models, making them teach each other and using prediction disagreement samples to update networks. Extensive experiments on four unsupervised domain tasks, Market-to-Duke, Duke-to-Market, Market-to-MSMT and Duke-to-MSMT, show that the advantages of our proposed MDL framework compared with other state-of-the-art methods. Liyan Zhang 0001 |
ICME | 2 |
| 2021 | Host-Parasite: Graph LSTM-in-LSTM for Group Activity RecognitionabstractThis article aims to tackle the problem of group activity recognition in the multiple-person scene. To model the group activity with multiple persons, most long short-term memory (LSTM)-based methods first learn the person-level action representations by several LSTMs and then integrate all the person-level action representations into the following LSTM to learn the group-level activity representation. This type of solution is a two-stage strategy, which neglects the "host-parasite" relationship between the group-level activity ("host") and person-level actions ("parasite") in spatiotemporal space. To this end, we propose a novel graph LSTM-in-LSTM (GLIL) for group activity recognition by modeling the person-level actions and the group-level activity simultaneously. GLIL is a "host-parasite" architecture, which can be seen as several person LSTMs (P-LSTMs) in the local view or a graph LSTM (G-LSTM) in the global view. Specifically, P-LSTMs model the person-level actions based on the interactions among persons. Meanwhile, G-LSTM models the group-level activity, where the person-level motion information in multiple P-LSTMs is selectively integrated and stored into G-LSTM based on their contributions to the inference of the group activity class. Furthermore, to use the person-level temporal features instead of the person-level static features as the input of GLIL, we introduce a residual LSTM with the residual connection to learn the person-level residual features, consisting of temporal features and static features. Experimental results on two public data sets illustrate the effectiveness of the proposed GLIL compared with state-of-the-art methods. Xiangbo Shu, Liyan Zhang 0001, Yunlian Sun, Jinhui Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | A treatment engine by multimodal EMR dataabstractIn recent years, with the development of electronic medical record (EMR) systems, it has become possible to mine patient clinical data to improve medical care quality. After the treatment engine learns knowledge from the EMR data, it can automatically recommend the next stage of prescriptions and provide treatment guidelines for doctors and patients. However, this task is always challenged by the multi-modality of EMR data. To more effectively predict the next stage of treatment prescription by using multimodal information and the connection between the modalities, we propose a cross-modal shared-specific feature complementary generation and attention fusion algorithm. In the feature extraction stage, specific information and shared information are obtained through a shared-specific feature extraction network. To obtain the correlation between the modalities, we propose a sorting network. We use the attention fusion network in the multimodal feature fusion stage to give different multimodal features at different stages with different weights to obtain a more prepared patient representation. Considering the redundant information of specific modal information and shared modal information, we introduce a complementary feature learning strategy, including modality adaptation for shared features, project adversarial learning for specific features, and reconstruction enhancement. The experimental results on the real EMR data set MIMIC-III prove its superiority and each part's effectiveness. Zhaomeng Huang, Liyan Zhang 0001 |
MMAsia | 2 |
| 2020 | Distilling knowledge in causal inference for unbiased visual question answeringabstractCurrent Visual Question Answering (VQA) models mainly explore the statistical correlations between answers and questions, which fail to capture the relationship between the visual information and answers. The performance dramatically decreases when the distribution of handled data is different from the training data. Towards this end, this paper proposes a novel unbiased VQA model by exploring the Casual Inference with Knowledge Distillation (CIKD) to reduce the influence of bias. Specifically, the causal graph is first constructed to explore the counterfactual causality and infer the casual target based on the causal effect, which well reduces the bias from questions and obtain answers without training. Then knowledge distillation is leveraged to transfer the knowledge of the inferred casual target to the conventional VQA model. It makes the proposed method enable to handle both the biased data and standard data. To address the problem of the bad bias from the knowledge distillation, the ensemble learning is introduced based on the hypothetical bias reason. Experiments are conducted to show the performance of the proposed method. The significant improvements over the state-of-the-art methods on the VQA-CP v2 dataset well validate the contributions of this work. Yonghua Pan, Zechao Li, Liyan Zhang 0001, Jinhui Tang 0001 |
MMAsia | 3 |
| 2020 | Hierarchical clustering via mutual learning for unsupervised person re-identificationabstractPerson re-identification (re-ID) aims to establish identity correspondence across different cameras. State-of-the-art re-ID approaches are mainly clustering-based Unsupervised Domain Adaptation (UDA) methods, which attempt to transfer the model trained on the source domain to target domain, by alternatively generating pseudo labels by clustering target-domain instances and training the network with generated pseudo labels to perform feature learning. However, these approaches suffer from the problem of inevitable label noise caused by the clustering procedure that dramatically impact the model training and feature learning of the target domain. To address this issue, we propose an unsupervised Hierarchical Clustering via Mutual Learning (HCML) framework, which can jointly optimize the dual training network and the clustering procedure to learn more discriminative features from the target domain. Specifically, the proposed HCML framework can effectively update the hard pseudo labels generated by clustering process and soft pseudo label generated by the training network both in on-line manner. We jointly adopt the repelled loss, triplet loss, soft identity loss and soft triplet loss to optimize the model. The experimental results on Market-to-Duke, Duke-to-Market, Market-to-MSMT and Duke-to-MSMT unsupervised domain adaptation tasks have demonstrated the superiority of our proposed HCML framework compared with other state-of-the-art methods. Liyan Zhang 0001, Zhaomeng Huang, Guodong Du 0005 |
MMAsia | 2 |
| 2020 | Weakly-supervised Semantic Guided Hashing for Social Image Retrieval
Zechao Li, Jinhui Tang 0001, Liyan Zhang 0001, Jian Yang 0003 |
Int. J. Comput. Vis. | 3 |
| 2019 | Image annotation refinement via 2P-KNN based group sparse reconstruction
Qian Ji, Liyan Zhang 0001, Xiangbo Shu, Jinhui Tang 0001 |
Multim. Tools Appl. | 2 |
| 2018 | A Feature Selection Method for Projection Twin Support Vector Machine
Rui Yan 0010, Qiaolin Ye, Liyan Zhang 0001, Xiangbo Shu |
Neural Process. Lett. | 3 |
| 2018 | Personalized Age Progression with Bi-Level Aging Dictionary LearningabstractAge progression is defined as aesthetically re-rendering the aging face at any future age for an individual face. In this work, we aim to automatically render aging faces in a personalized way. Basically, for each age group, we learn an aging dictionary to reveal its aging characteristics (e.g., wrinkles), where the dictionary bases corresponding to the same index yet from two neighboring aging dictionaries form a particular aging pattern cross these two age groups, and a linear combination of all these patterns expresses a particular personalized aging process. Moreover, two factors are taken into consideration in the dictionary learning process. First, beyond the aging dictionaries, each person may have extra personalized facial characteristics, e.g., mole, which are invariant in the aging process. Second, it is challenging or even impossible to collect faces of all age groups for a particular person, yet much easier and more practical to get face pairs from neighboring age groups. To this end, we propose a novel Bi-level Dictionary Learning based Personalized Age Progression (BDL-PAP) method. Here, bi-level dictionary learning is formulated to learn the aging dictionaries based on face pairs from neighboring age groups. Extensive experiments well demonstrate the advantages of the proposed BDL-PAP over other state-of-the-arts in term of personalized age progression, as well as the performance gain for cross-age face verification by synthesizing aging faces. Xiangbo Shu, Jinhui Tang 0001, Zechao Li, Hanjiang Lai, Liyan Zhang 0001, Shuicheng Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2017 | Computational face reader based on facial attribute estimation
Xiangbo Shu, Yunfei Cai, Liyan Zhang 0001, Jinhui Tang 0001 |
Neurocomputing | 4 |
| 2016 | Computational Face Reader
Xiangbo Shu, Liyan Zhang 0001, Jinhui Tang 0001, Guosen Xie, Shuicheng Yan |
MMM (1) | 2 |
| 2016 | Scalable attribute-driven face image retrieval
Changjian Zou, Liyan Zhang 0001, Bradley Denney |
Neurocomputing | 3 |
| 2016 | Sub-event recognition and summarization for structured scenario photos
Liyan Zhang 0001, Bradley Denney, Juwei Lu |
Multim. Tools Appl. | 1 |
| 2016 | Query-Driven Approach to Face Clustering and TaggingabstractIn the era of big data, a traditional offline setting to processing image data is simply not tenable. We simply do not have the computational power to process every image with every possible tag; moreover, we will not have the manpower to clean up the potentially noisy results. In this paper, we introduce a query-driven approach to visual tagging, focusing on the application of face tagging and clustering. We integrate active learning with query-driven probabilistic databases. Rather than asking a user to provide manual labels so as to minimize the uncertainty of labels (face tags) across the entire data set, we ask the user to provide labels that minimize the uncertainty of his/her query result (e.g., "How many times did Bob and Jim appear together?"). We use a data-driven Gaussian process model of facial appearance to write the probabilistic estimates of facial identity into a probabilistic database, which can then support inference through query answering. Importantly, the database is augmented with contextual constraints (faces in the same image cannot be the same identity, while faces in the same track must be identical). Experiments on the real-world photo collections demonstrate the effectiveness of the proposed method. Liyan Zhang 0001, Xikui Wang, Dmitri V. Kalashnikov, Sharad Mehrotra, Deva Ramanan |
IEEE Trans. Image Process. | 1 |
| 2015 | Semantic-aware Hashing for Social Image RetrievalabstractWith the proliferation of large-scale social images, recent years have witnessed the increasing amount of images with user-provided tags, which leads to considerable effort made on hashing based approximate nearest neighbor (ANN) search in huge databases. In this work, we propose a novel Semantic-aware Hashing method (SaH) by discovering knowledge from these social media resources to implement approximate similarity search. Different from the previous work, the proposed method learns semantic hashing codes by exploiting heterogeneous information from the textual and visual domains. The semantic structure in the textual domain is well preserved to learn the binary codes. To handle the noisy, incomplete, or subjective user-provided tags, the visual structure is also leveraged. On the other hand, an information theoretic regularization is exploited by using maximum entropy principle and a row-wise sparse model with l2,p (0 < p ≤ 1) mixed norm is introduced to filter certain noisy or redundant visual features. Experiments are conducted on a widely-used social image dataset and the comparison results demonstrate the outperforming performance of the proposed SaH method over state-of-the-art hashing techniques. Jinhui Tang 0001, Zechao Li, Liyan Zhang 0001, Qingming Huang |
ICMR | 3 |
| 2015 | Local Structure-Based Sparse Representation for Face RecognitionabstractThis article presents a simple yet effective face recognition method, called local structure-based sparse representation classification (LS_SRC). Motivated by the “divide-and-conquer” strategy, we first divide the face into local blocks and classify each local block, then integrate all the classification results to make the final decision. To classify each local block, we further divide each block into several overlapped local patches and assume that these local patches lie in a linear subspace. This subspace assumption reflects the local structure relationship of the overlapped patches, making sparse representation-based classification (SRC) feasible even when encountering the single-sample-per-person (SSPP) problem. To lighten the computing burden of LS_SRC, we further propose the local structure-based collaborative representation classification (LS_CRC). Moreover, the performance of LS_SRC and LS_CRC can be further improved by using the confusion matrix of the classifier. Experimental results on four public face databases show that our methods not only generalize well to SSPP problem but also have strong robustness to occlusion; little pose variation; and the variations of expression, illumination, and time. Fan Liu 0003, Jinhui Tang 0001, Yan Song 0005, Liyan Zhang 0001, Zhenmin Tang |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2015 | Real-Time System for Driver Fatigue Detection by RGB-D CameraabstractDrowsy driving is one of the major causes of fatal traffic accidents. In this article, we propose a real-time system that utilizes RGB-D cameras to automatically detect driver fatigue and generate alerts to drivers. By introducing RGB-D cameras, the depth data can be obtained, which provides extra evidence to benefit the task of head detection and head pose estimation. In this system, two important visual cues (head pose and eye state) for driver fatigue detection are extracted and leveraged simultaneously. We first present a real-time 3D head pose estimation method by leveraging RGB and depth data. Then we introduce a novel method to predict eye states employing the WLBP feature, which is a powerful local image descriptor that is robust to noise and illumination variations. Finally, we integrate the results from both head pose and eye states to generate the overall conclusion. The combination and collaboration of the two types of visual cues can reduce the uncertainties and resolve the ambiguity that a single cue may induce. The experiments were performed using an inside-car environment during the day and night, and theyfully demonstrate the effectiveness and robustness of our system as well as the proposed methods of predicting head pose and eye states. Liyan Zhang 0001, Fan Liu 0003, Jinhui Tang 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2014 | A collaborative approach for face verification and attributes refinement
Liyan Zhang 0001, Bradley Denney, Juwei Lu |
Inf. Sci. | 1 |
| 2014 | Context-based person identification framework for smart video surveillance
Liyan Zhang 0001, Dmitri V. Kalashnikov, Sharad Mehrotra, Ronen Vaisenberg |
Mach. Vis. Appl. | 1 |
| 2014 | Cross-Indexing of Binary SIFT Codes for Large-Scale Image SearchabstractIn recent years, there has been growing interest in mapping visual features into compact binary codes for applications on large-scale image collections. Encoding high-dimensional data as compact binary codes reduces the memory cost for storage. Besides, it benefits the computational efficiency since the computation of similarity can be efficiently measured by Hamming distance. In this paper, we propose a novel flexible scale invariant feature transform (SIFT) binarization (FSB) algorithm for large-scale image search. The FSB algorithm explores the magnitude patterns of SIFT descriptor. It is unsupervised and the generated binary codes are demonstrated to be dispreserving. Besides, we propose a new searching strategy to find target features based on the cross-indexing in the binary SIFT space and original SIFT space. We evaluate our approach on two publicly released data sets. The experiments on large-scale partial duplicate image retrieval system demonstrate the effectiveness and efficiency of the proposed algorithm. Houqiang Li, Liyan Zhang 0001, Wengang Zhou 0001, Qi Tian 0001 |
IEEE Trans. Image Process. | 3 |
| 2013 | A unified framework for context assisted face clusteringabstractAutomatic face clustering, which aims to group faces referring to the same people together, is a key component for face tagging and image management. Standard face clustering approaches that are based on analyzing facial features can already achieve high-precision results. However, they often suffer from low recall due to the large variation of faces in pose, expression, illumination, occlusion, etc. To improve the clustering recall without reducing the high precision, we leverage the heterogeneous context information to iteratively merge the clusters referring to same entities. We first investigate the appropriate methods to utilize the context information at the cluster level, including using of "common scene", people co-occurrence, human attributes, and clothing. We then propose a unified framework that employs bootstrapping to automatically learn adaptive rules to integrate this heterogeneous contextual information, along with facial features, together. Experimental results on two personal photo collections and one real-world surveillance dataset demonstrate the effectiveness of the proposed approach in improving recall while maintaining very high precision of face clustering. Liyan Zhang 0001, Dmitri V. Kalashnikov, Sharad Mehrotra |
ICMR | 1 |
| 2013 | A random-walk based recommendation algorithm considering item categories
Liyan Zhang 0001, Chunping Li |
Neurocomputing | 1 |
| 2013 | Cross-Space Affinity Learning with Its Application to Movie RecommendationabstractIn this paper, we propose a novel cross-space affinity learning algorithm over different spaces with heterogeneous structures. Unlike most of affinity learning algorithms on the homogeneous space, we construct a cross-space tensor model to learn the affinity measures on heterogeneous spaces subject to a set of order constraints from the training pool. We further enhance the model with a factorization form which greatly reduces the number of parameters of the model with a controlled complexity. Moreover, from the practical perspective, we show the proposed factorized cross-space tensor model can be efficiently optimized by a series of simple quadratic optimization problems in an iterative manner. The proposed cross-space affinity learning algorithm can be applied to many real-world problems, which involve multiple heterogeneous data objects defined over different spaces. In this paper, we apply it into the recommendation system to measure the affinity between users and the product items, where a higher affinity means a higher rating of the user on the product. For an empirical evaluation, a widely used benchmark movie recommendation data set—MovieLens—is used to compare the proposed algorithm with other state-of-the-art recommendation algorithms and we show that very competitive results can be obtained. Jinhui Tang 0001, Guo-Jun Qi, Liyan Zhang 0001, Changsheng Xu |
IEEE Trans. Knowl. Data Eng. | 3 |