Yimin Liu 0001

dblp:69/866-1 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0002-5543-4602ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Camera-aware Embedding Refinement for unsupervised person re-identification
Yimin Liu 0001, Meibin Qi, Yongle Zhang 0001, Wenbo Xu 0004, Qiang Wu 0001
Knowl. Based Syst.1
2025 A Temporal-Semantic Interaction Network for multi-frame infrared small target detection
Shuo Zhuang, Meibin Qi, Di Wang 0026, Kunyuan Li, Yimin Liu 0001
Knowl. Based Syst.6
2025 Consistent Image Inpainting With Pre-Perception and Cross-Perception Collaborative Processes
abstract
It has been proven that introducing multiple guidance sources boosts image inpainting performance. However, existing methods primarily focus on local relationships and neglect the holistic interplay between guidance and texture information. Moreover, they lack an effective feedback mechanism to adaptively update the guidance process as corrupted texture information is progressively restored, potentially resulting in inconsistent inpainting. To tackle this issue, we propose a novel scheme aligned with pre-perception and cross-perception collaborative processes in human drawing. To mimic the pre-perception process, we introduce a pre-perceptual transformer block that captures long-range contextual dependencies and activates meaningful information to individually optimize image structures, semantic layouts, and textures, thereby effectively controlling their respective generation. To mimic the cross-perception collaborative process, we propose a cyclic cross-perceptual interaction to maintain consistency across the entire image regarding structure, layout, and texture while progressively refining their details. This interaction accounts for the global attention relationship between texture and other guidance sources (including image structure and semantic layout) to enhance image texture, alongside integrating a dedicated feedback mechanism to update guidance information. The proposed components are alternately deployed in three-branch decoders of the new scheme from rough to fine-grained levels to achieve these two iterative processes of human drawing. Experimental results prove the superiority of the proposed scheme over state-of-the-art methods across three datasets.
Yongle Zhang 0001, Yimin Liu 0001, Hao Fan 0004, Ruotong Hu, Jian Zhang 0002, Qiang Wu 0001
IEEE Trans. Image Process.2
2024 Exploring reliable infrared object tracking with spatio-temporal fusion transformer
Meibin Qi, Qinxin Wang, Shuo Zhuang, Kunyuan Li, Yimin Liu 0001, Yanfang Yang
Knowl. Based Syst.6
2024 Polarized Prior Guided Fusion Network for Infrared Polarization Images
abstract
Typical infrared polarization image fusion aims to integrate background details in the infrared intensity and salient target in the degree of linear polarization (DoLP). Many fusion methods show advanced network architecture, but few works can form effective feature representations for the differences in prior distributions of the infrared intensity and DoLP, and the interference of DoLP with noise makes fusion more challenging. This paper employs a learned low-rank decomposition model to extract low-rank representations containing background details in infrared intensity and sparse features with salient targets in DoLP. To reduce noise interference, we design a fusion module based on an attention-guided filter, where the infrared intensity serves as a guide map to suppress the background in DoLP. Moreover, a novel loss constraint is proposed to improve the fusion performance. Specifically, the fusion network is trained by reconstructing polarized images in different directions from the fused image. Quantitative and qualitative experimental results validate the effectiveness of our approach. In comparison to existing methods, our fusion model can better preserve the polarization salient target and suppress the background interference with fewer parameters. The source code is available at https://github.com/lkyahpu/PIPFNet.
Kunyuan Li, Meibin Qi, Shuo Zhuang, Yimin Liu 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Improving Consistency of Proxy-Level Contrastive Learning for Unsupervised Person Re-Identification
abstract
Recently, contrastive learning-based unsupervised person re-identification (Re-ID) methods have garnered significant attention due to their effectiveness. These methods rely on predicted pseudo-labels to construct contrastive pairs, optimizing the network gradually. Some methods also utilize camera labels to explore intra-camera and inter-camera contrastive relations, achieving state-of-the-art results. However, these methods fail to address the issue of inconsistency in proxy-level contrastive learning, which arises from variations in the distribution of instances belonging to the same proxy. Specifically, they are sensitive to the distribution of instances in a mini-batch used for contrastive pair construction, and uncertainty or noise in the data distribution can lead to turbulence in the contrastive loss, degrading the effectiveness of contrastive learning. In this work, we first propose a dual-branch contrastive learning (DBCL) framework. The framework comprises a dual-branch structure with an identity discrimination branch and a camera view awareness branch. These branches are mutually trained to produce a jointly optimized model with both high person identification accuracy and cross-camera robustness. Moreover, to mitigate the proxy-level contrastive inconsistency issue in the camera view awareness branch, we design intra-camera and inter-camera consistent contrastive losses. Our DBCL has been extensively evaluated on several person Re-ID datasets and has demonstrated superior performance compared to state-of-the-art methods. Notably, on the challenging MSMT17 dataset with complex scenes, our method achieved an mAP of 45.3% and Rank-1 accuracy of 75.3%.
Yimin Liu 0001, Meibin Qi, Yongle Zhang 0001, Qiang Wu 0001, Jingjing Wu 0001, Shuo Zhuang
IEEE Trans. Inf. Forensics Secur.1
2024 Mutual Dual-Task Generator With Adaptive Attention Fusion for Image Inpainting
abstract
Image segmentation can reveal the semantic structure information in an image, which is helpful guidance information for image inpainting. Notably, it can help mitigate the artifacts on the boundaries of different semantic regions during the inpainting process. Existing semantic guidance-based image inpainting provides one-way guidance from the semantic segmentation task to the image inpainting task. There is no feedback from the inpainting results to adjust the guidance process, which causes inferior performance. To tackle this issue, this work proposes mutual dual-task generators to establish the interaction between image segmentation and image inpainting tasks. Thus, semantic segmentation guides image inpainting and also receives feedback from image inpainting. These two processes interact with each other and progressively improve the inpainting quality. The mutual dual-task generator consists of a shared encoder and mutual decoders with the bidirectional Cross-domain Feature DeNormalization (CFDN) module inside, which hierarchically models the Segmentation-guided image Texture (ST) generation and Texture-guided semantic Segmentation (TS) generation. At the end of mutual decoders, an Adaptive Attention Fusion (AAF) module is proposed to augment the texture and semantic class affinity between pixels, further refining the inpainted results. Experimental results demonstrate that the proposed mutual dual-task generator pipeline achieves superior inpainting performances over the state of the arts on three public datasets.
Yongle Zhang 0001, Yimin Liu 0001, Ruotong Hu, Qiang Wu 0001, Jian Zhang 0002
IEEE Trans. Multim.2
2023 Camera Proxy based Contrastive Learning with Hard Sampling for Unsupervised Person Re-identification
abstract
Because of the advantages of dealing with large-scale unlabelled data, unsupervised learning has recently attracted more attention for person re-identification. Particularly, the combination of the unsupervised learning paradigm with contrastive learning shows promising efficiency in network optimization. This work adopts the successful camera-aware contrastive learning approach and further explores its capability on the camera proxy level to improve the data pair consistency. Thus, it is more robust to the camera change, which still challenges the unsupervised person re-identification. This work proposed a Camera Proxy-based Contrastive Learning framework, which explicitly considers inter-camera scenario and intra-camera scenario. Moreover, this work is motivated by the strategy of selecting a hard negative sample in triplet loss learning and further extends it to contrastive learning for both negative and positive pair creation on the camera proxy level. Extensive experiments demonstrate the superiority of the proposed framework over state-of-the-art approaches on purely unsupervised re-identification.
Yimin Liu 0001, Meibin Qi, Qiang Wu 0001, Yanfang Yang, Xiaohong Li 0002, Jian Zhang 0002
ICME1
2022 A Local-Global Self-attention Interaction Network for RGB-D Cross-Modal Person Re-identification
Chuanlei Zhu, Xiaohong Li 0002, Meibin Qi, Yimin Liu 0001
PRCV (4)4
2022 Saliency and Granularity: Discovering Temporal Coherence for Video-Based Person Re-Identification
abstract
Video-based person re-identification (ReID) matches the same people across the video sequences with rich spatial and temporal information in complex scenes. It is highly challenging to capture discriminative information when occlusions and pose variations exist between frames. A key solution to this problem rests on extracting the temporal invariant features of video sequences. In this paper, we propose a novel method for discovering temporal coherence by designing a region-level saliency and granularity mining network (SGMN). Firstly, to address the varying noisy frame problem, we design a temporal spatial-relation module (TSRM) to locate frame-level salient regions, adaptively modeling the temporal relations on spatial dimension through a probe-buffer mechanism. It avoids the information redundancy between frames and captures the informative cues of each frame. Secondly, a temporal channel-relation module (TCRM) is proposed to further mine the small granularity information of each frame, which is complementary to TSRM by concentrating on discriminative small-scale regions. TCRM exploits a one-and-rest difference relation on channel dimension to enhance the granularity features, leading to stronger robustness against misalignments. Finally, we evaluate our SGMN with four representative video-based datasets, including iLIDS-VID, MARS, DukeMTMC-VideoReID, and LS-VID, and the results indicate the effectiveness of the proposed method.
Cuiqun Chen, Mang Ye, Meibin Qi, Jingjing Wu 0001, Yimin Liu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2022 Improving Feature Discrimination for Object Tracking by Structural-similarity-based Metric Learning
abstract
Existing approaches usually form the tracking task as an appearance matching procedure. However, the discrimination ability of appearance features is insufficient in these trackers, which is caused by their weak feature supervision constraints and inadequate exploitation of spatial contexts. To tackle this issue, this article proposes a novel appearance matching tracking (AMT) method to strengthen the feature restraints and capture discriminative spatial representations. Specifically, we first utilize a triplet structural loss function, which improves the learning capability of features by applying a structural similarity constraint with a triplet metric format on the features. It leverages feature statistics to capture the complex interactions of visual parts. Second, we put forward an adaptive matching module that exploits the dual spatial enhancement module to reinforce target feature discrimination. This not only boosts the representation ability of spatial context but also realizes spatially dynamic feature selection by attending to target deformation information. Moreover, this model introduces a simple but effective matching unit to intuitively evaluate the relative appearance differences between the target and the proposals. In addition, with the obtained discriminative features, AMT is capable of providing precise localization for the target. Therefore, the impact of spatial suppression imposed by window functions can be alleviated, allowing for effective tracking of high-speed moving objects. Extensive experiments prove that AMT outperforms state-of-the-art methods on six public datasets and demonstrate the effectiveness of each component in AMT.
Jingjing Wu 0001, Meibin Qi, Cuiqun Chen, Yimin Liu 0001
ACM Trans. Multim. Comput. Commun. Appl.5