VLDB 2026 Research / reviewers in the wild / expert
Meibin Qi
dblp:39/4774
· DBLP profile ↗
36ranked-venue papers
4as first author
22since 2021 · last 2025
0000-0001-9074-211XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 7 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Camera-aware Embedding Refinement for unsupervised person re-identification
Yimin Liu 0001, Meibin Qi, Yongle Zhang 0001, Wenbo Xu 0004, Qiang Wu 0001 |
Knowl. Based Syst. | 2 |
| 2025 | A Temporal-Semantic Interaction Network for multi-frame infrared small target detection
Shuo Zhuang, Meibin Qi, Di Wang 0026, Kunyuan Li, Yimin Liu 0001 |
Knowl. Based Syst. | 3 |
| 2025 | Contrastive learning-based joint pre-training for unsupervised domain adaptive person re-identification
Xiaohong Li 0002, Xuesong Dai, Shuo Zhuang, Meibin Qi |
Multim. Syst. | 5 |
| 2024 | Exploring reliable infrared object tracking with spatio-temporal fusion transformer
Meibin Qi, Qinxin Wang, Shuo Zhuang, Kunyuan Li, Yimin Liu 0001, Yanfang Yang |
Knowl. Based Syst. | 1 |
| 2024 | SketchTrans: Disentangled Prototype Learning With Transformer for Sketch-Photo RecognitionabstractMatching hand-drawn sketches with photos (a.k.a sketch-photo recognition or re-identification) faces the information asymmetry challenge due to the abstract nature of the sketch modality. Existing works tend to learn shared embedding spaces with CNN models by discarding the appearance cues for photo images or introducing GAN for sketch-photo synthesis. The former unavoidably loses discriminability, while the latter contains ineffaceable generation noise. In this paper, we start the first attempt to design an information-aligned sketch transformer (Sketch Trans+) viacross-modal disentangled prototype learning, while the transformer has shown great promise for discriminative visual modelling. Specifically, we design an asymmetric disentanglement scheme with a dynamic updatable auxiliary sketch (A-sketch) to align the modality representations without sacrificing information. The asymmetric disentanglement decomposes the photo representations into sketch-relevant and sketch-irrelevant cues, transferring sketch-irrelevant knowledge into the sketch modality to compensate for the missing information. Moreover, considering the feature discrepancy between the two modalities, we present a modality-aware prototype contrastive learning method that mines representative modality-sharing information using the modality-aware prototypes rather than the original feature representations. Extensive experiments on categoryand instance-level sketch-based datasets validate the superiority of our proposed method under various metrics. Code is available athttps://github.com/ccq195/SketchTrans Cuiqun Chen, Mang Ye, Meibin Qi, Bo Du 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Polarized Prior Guided Fusion Network for Infrared Polarization ImagesabstractTypical infrared polarization image fusion aims to integrate background details in the infrared intensity and salient target in the degree of linear polarization (DoLP). Many fusion methods show advanced network architecture, but few works can form effective feature representations for the differences in prior distributions of the infrared intensity and DoLP, and the interference of DoLP with noise makes fusion more challenging. This paper employs a learned low-rank decomposition model to extract low-rank representations containing background details in infrared intensity and sparse features with salient targets in DoLP. To reduce noise interference, we design a fusion module based on an attention-guided filter, where the infrared intensity serves as a guide map to suppress the background in DoLP. Moreover, a novel loss constraint is proposed to improve the fusion performance. Specifically, the fusion network is trained by reconstructing polarized images in different directions from the fused image. Quantitative and qualitative experimental results validate the effectiveness of our approach. In comparison to existing methods, our fusion model can better preserve the polarization salient target and suppress the background interference with fewer parameters. The source code is available at https://github.com/lkyahpu/PIPFNet. Kunyuan Li, Meibin Qi, Shuo Zhuang, Yimin Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Improving Consistency of Proxy-Level Contrastive Learning for Unsupervised Person Re-IdentificationabstractRecently, contrastive learning-based unsupervised person re-identification (Re-ID) methods have garnered significant attention due to their effectiveness. These methods rely on predicted pseudo-labels to construct contrastive pairs, optimizing the network gradually. Some methods also utilize camera labels to explore intra-camera and inter-camera contrastive relations, achieving state-of-the-art results. However, these methods fail to address the issue of inconsistency in proxy-level contrastive learning, which arises from variations in the distribution of instances belonging to the same proxy. Specifically, they are sensitive to the distribution of instances in a mini-batch used for contrastive pair construction, and uncertainty or noise in the data distribution can lead to turbulence in the contrastive loss, degrading the effectiveness of contrastive learning. In this work, we first propose a dual-branch contrastive learning (DBCL) framework. The framework comprises a dual-branch structure with an identity discrimination branch and a camera view awareness branch. These branches are mutually trained to produce a jointly optimized model with both high person identification accuracy and cross-camera robustness. Moreover, to mitigate the proxy-level contrastive inconsistency issue in the camera view awareness branch, we design intra-camera and inter-camera consistent contrastive losses. Our DBCL has been extensively evaluated on several person Re-ID datasets and has demonstrated superior performance compared to state-of-the-art methods. Notably, on the challenging MSMT17 dataset with complex scenes, our method achieved an mAP of 45.3% and Rank-1 accuracy of 75.3%. Yimin Liu 0001, Meibin Qi, Yongle Zhang 0001, Qiang Wu 0001, Jingjing Wu 0001, Shuo Zhuang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Asymmetric Deformable Spatio-temporal Framework for Infrared Object TrackingabstractThe Infrared Object Tracking (IOT) task aims to locate objects in infrared sequences. Since color and texture information is unavailable in infrared modality, most existing infrared trackers merely rely on capturing spatial contexts from the image to enhance feature representation, where other complementary information is rarely deployed. To fill this gap, we in this article propose a novel Asymmetric Deformable Spatio-Temporal Framework (ADSF) to fully exploit collaborative shape and temporal clues in terms of the objects. Firstly, an asymmetric deformable cross-attention module is designed to extract shape information, which attends to the deformable correlations between distinct frames in an asymmetric manner. Secondly, a spatio-temporal tracking framework is coined to learn the temporal variance trend of the object during the training process and store the template information closest to the tracking frame when testing. Comprehensive experiments demonstrate that ADSF outperforms state-of-the-art methods on three public datasets. Extensive ablation experiments further confirm the effectiveness of each component in ADSF. Furthermore, we conduct generalization validation to demonstrate that the proposed method also achieves performance gains in RGB-based tracking scenarios. Jingjing Wu 0001, Xi Zhou 0004, Xiaohong Li 0002, Hao Liu 0003, Meibin Qi, Richang Hong |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Camera Proxy based Contrastive Learning with Hard Sampling for Unsupervised Person Re-identificationabstractBecause of the advantages of dealing with large-scale unlabelled data, unsupervised learning has recently attracted more attention for person re-identification. Particularly, the combination of the unsupervised learning paradigm with contrastive learning shows promising efficiency in network optimization. This work adopts the successful camera-aware contrastive learning approach and further explores its capability on the camera proxy level to improve the data pair consistency. Thus, it is more robust to the camera change, which still challenges the unsupervised person re-identification. This work proposed a Camera Proxy-based Contrastive Learning framework, which explicitly considers inter-camera scenario and intra-camera scenario. Moreover, this work is motivated by the strategy of selecting a hard negative sample in triplet loss learning and further extends it to contrastive learning for both negative and positive pair creation on the camera proxy level. Extensive experiments demonstrate the superiority of the proposed framework over state-of-the-art approaches on purely unsupervised re-identification. Yimin Liu 0001, Meibin Qi, Qiang Wu 0001, Yanfang Yang, Xiaohong Li 0002, Jian Zhang 0002 |
ICME | 2 |
| 2023 | Signal separation and super-resolution DOA estimation based on multi-objective joint learning
Houhong Xiang, Meibin Qi, Baixiao Chen, Shuo Zhuang |
Appl. Intell. | 2 |
| 2022 | Sketch Transformer: Asymmetrical Disentanglement Learning from Dynamic SynthesisabstractSketch-photo recognition is a cross-modal matching problem whose query sets are sketch images drawn by artists or amateurs. Due to the significant modality difference between the two modalities, it is challenging to extract discriminative modality-shared feature representations. Existing works focus on exploring modality-invariant features to discover shared embedding space. However, they discard modality-specific cues, resulting in information loss and diminished discriminatory power of features. This paper proposes a novel asymmetrical disentanglement and dynamic synthesis learning method in the transformer framework (SketchTrans) to handle modality discrepancy by combining modality-shared information with modality-specific information. Specifically, an asymmetrical disentanglement scheme is introduced to decompose the photo features into sketch-relevant and sketch-irrelevant cues while preserving the original sketch structure. Using the sketch-irrelevant cues, we further translate the sketch modality component to photo representation through knowledge transfer, obtaining cross-modality representations with information symmetry. Moreover, we propose a dynamic updatable auxiliary sketch (A-sketch) modality generated from the photo modality to guide the asymmetrical disentanglement in a single framework. Under a multi-modality joint learning framework, this auxiliary modality increases the diversity of training samples and narrows the cross-modality gap. We conduct extensive experiments on three fine-grained sketch-based retrieval datasets, i.e., PKU-Sketch, QMUL-ChairV2, and QMUL-ShoeV2, outperforming the state-of-the-arts under various metrics. Cuiqun Chen, Mang Ye, Meibin Qi, Bo Du 0001 |
ACM Multimedia | 3 |
| 2022 | A Local-Global Self-attention Interaction Network for RGB-D Cross-Modal Person Re-identification
Chuanlei Zhu, Xiaohong Li 0002, Meibin Qi, Yimin Liu 0001 |
PRCV (4) | 3 |
| 2022 | Saliency and Granularity: Discovering Temporal Coherence for Video-Based Person Re-IdentificationabstractVideo-based person re-identification (ReID) matches the same people across the video sequences with rich spatial and temporal information in complex scenes. It is highly challenging to capture discriminative information when occlusions and pose variations exist between frames. A key solution to this problem rests on extracting the temporal invariant features of video sequences. In this paper, we propose a novel method for discovering temporal coherence by designing a region-level saliency and granularity mining network (SGMN). Firstly, to address the varying noisy frame problem, we design a temporal spatial-relation module (TSRM) to locate frame-level salient regions, adaptively modeling the temporal relations on spatial dimension through a probe-buffer mechanism. It avoids the information redundancy between frames and captures the informative cues of each frame. Secondly, a temporal channel-relation module (TCRM) is proposed to further mine the small granularity information of each frame, which is complementary to TSRM by concentrating on discriminative small-scale regions. TCRM exploits a one-and-rest difference relation on channel dimension to enhance the granularity features, leading to stronger robustness against misalignments. Finally, we evaluate our SGMN with four representative video-based datasets, including iLIDS-VID, MARS, DukeMTMC-VideoReID, and LS-VID, and the results indicate the effectiveness of the proposed method. Cuiqun Chen, Mang Ye, Meibin Qi, Jingjing Wu 0001, Yimin Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Structure-Aware Positional Transformer for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is a cross-modality retrieval problem, which aims at matching the same pedestrian between the visible and infrared cameras. Due to the existence of pose variation, occlusion, and huge visual differences between the two modalities, previous studies mainly focus on learning image-level shared features. Since they usually learn a global representation or extract uniformly divided part features, these methods are sensitive to misalignments. In this paper, we propose a structure-aware positional transformer (SPOT) network to learn semantic-aware sharable modality features by utilizing the structural and positional information. It consists of two main components: attended structure representation (ASR) and transformer-based part interaction (TPI). Specifically, ASR models the modality-invariant structure feature for each modality and dynamically selects the discriminative appearance regions under the guidance of the structure information. TPI mines the part-level appearance and position relations with a transformer to learn discriminative part-level modality features. With a weighted combination of ASR and TPI, the proposed SPOT explores the rich contextual and structural information, effectively reducing cross-modality difference and enhancing the robustness against misalignments. Extensive experiments indicate that SPOT is superior to the state-of-the-art methods on two cross-modal datasets. Notably, the Rank-1/mAP value on the SYSU-MM01 dataset has improved by 8.43%/6.80%. Cuiqun Chen, Mang Ye, Meibin Qi, Jingjing Wu 0001, Chia-Wen Lin |
IEEE Trans. Image Process. | 3 |
| 2022 | Audio Matters in Video Super-Resolution by Implicit Semantic GuidanceabstractVideo super-resolution (VSR) aims to use multiple consecutive low-resolution frames to recover the corresponding high-resolution frames. However, existing VSR methods only consider videos as image sequences, ignoring another essential timing informationaudio, while in fact, there is a semantic link between audio and vision, and extensive studies have shown that audio can provide supervisory information in visual networks. Meanwhile, the addition of semantic priors has been proven to be effective in super-resolution (SR) tasks, but a pretrained segmentation network is required to obtain semantic segmentation maps. By contrast, audio as the information contained in the video itself can be directly used. Therefore, in this study, we propose a novel and pluggable multiscale audiovisual fusion (MS-AVF) module to enhance VSR performance by exploiting the relevant audio information, which can be regarded as implicit semantic guidance compared with the kind of explicit segmentation priors. Specifically, we first fuse audiovisual features on the semantic feature maps of different granularities of the target frames, and then through a top-down multiscale fusion approach, feedback high-level semantics to the underlying global visual features layer by layer, thereby providing effective audio implicit semantic guidance for VSR. Experimental results show that audio can further improve the VSR effect. Moreover, by visualizing the learned attention mask, the proposed end-to-end model can automatically learn potential audiovisual semantic links, especially improving the accuracy and effectiveness of the SR of sound sources and their surrounding regions. Meibin Qi, Yang Zhao 0002, Wei Jia 0001, Ronggang Wang |
IEEE Trans. Multim. | 3 |
| 2022 | Improving Feature Discrimination for Object Tracking by Structural-similarity-based Metric LearningabstractExisting approaches usually form the tracking task as an appearance matching procedure. However, the discrimination ability of appearance features is insufficient in these trackers, which is caused by their weak feature supervision constraints and inadequate exploitation of spatial contexts. To tackle this issue, this article proposes a novel appearance matching tracking (AMT) method to strengthen the feature restraints and capture discriminative spatial representations. Specifically, we first utilize a triplet structural loss function, which improves the learning capability of features by applying a structural similarity constraint with a triplet metric format on the features. It leverages feature statistics to capture the complex interactions of visual parts. Second, we put forward an adaptive matching module that exploits the dual spatial enhancement module to reinforce target feature discrimination. This not only boosts the representation ability of spatial context but also realizes spatially dynamic feature selection by attending to target deformation information. Moreover, this model introduces a simple but effective matching unit to intuitively evaluate the relative appearance differences between the target and the proposals. In addition, with the obtained discriminative features, AMT is capable of providing precise localization for the target. Therefore, the impact of spatial suppression imposed by window functions can be alleviated, allowing for effective tracking of high-speed moving objects. Extensive experiments prove that AMT outperforms state-of-the-art methods on six public datasets and demonstrate the effectiveness of each component in AMT. Jingjing Wu 0001, Meibin Qi, Cuiqun Chen, Yimin Liu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | An End-to-end Heterogeneous Restraint Network for RGB-D Cross-modal Person Re-identificationabstractThe RGB-D cross-modal person re-identification (re-id) task aims to identify the person of interest across the RGB and depth image modes. The tremendous discrepancy between these two modalities makes this task difficult to tackle. Few researchers pay attention to this task, and the deep networks of existing methods still cannot be trained in an end-to-end manner. Therefore, this article proposes an end-to-end module for RGB-D cross-modal person re-id. This network introduces a cross-modal relational branch to narrow the gaps between two heterogeneous images. It models the abundant correlations between any cross-modal sample pairs, which are constrained by heterogeneous interactive learning. The proposed network also exploits a dual-modal local branch, which aims to capture the common spatial contexts in two modalities. This branch adopts shared attentive pooling and mutual contextual graph networks to extract the spatial attention within each local region and the spatial relations between distinct local parts, respectively. Experimental results on two public benchmark datasets, that is, the BIWI and RobotPKU datasets, demonstrate that our method is superior to the state-of-the-art. In addition, we perform thorough experiments to prove the effectiveness of each component in the proposed method. Jingjing Wu 0001, Meibin Qi, Cuiqun Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Self-Supervised Light Field Depth Estimation Using Epipolar Plane ImagesabstractExploiting light field data makes it possible to obtain dense and accurate depth map. However, synthetic scenes with limited disparity range cannot contain the diversity of real scenes. By training in synthetic data, current learning based methods do not perform well in real scenes. In this paper, we propose a self-supervised learning framework for light field depth estimation. Different from the existing end-to-end training methods using disparity label per pixel, our approach implements network training by estimating EPI disparity shift after refocusing, which extends the disparity range of epipolar lines. To reduce the sensitivity of EPI to noise, we propose a new input mode called EPI-Stack, which stacks EPIs in the view dimension. This method is less sensitive to noise scenes than traditional input mode and improves the efficiency of estimation. Compared with other state-of-the-art methods, the proposed method can also obtain higher quality results in real-world scenarios, especially in the complex occlusion and depth discontinuity. Kunyuan Li, Jun Zhang 0017, Jun Gao 0006, Meibin Qi |
3DV | 4 |
| 2021 | Towards accurate estimation for visual object tracking with multi-hierarchy feature aggregation
Jingjing Wu 0001, Meibin Qi, Xiaohong Li 0002 |
Neurocomputing | 3 |
| 2021 | Global-Local Graph Convolutional Network for cross-modality person re-identification
Xiaohong Li 0002, Cuiqun Chen, Meibin Qi, Jingjing Wu 0001 |
Neurocomputing | 4 |
| 2021 | Learning discriminative features with a dual-constrained guided network for video-based person re-identification
Cuiqun Chen, Meibin Qi, Guanghong Huang, Jingjing Wu 0001, Xiaohong Li 0002 |
Multim. Tools Appl. | 2 |
| 2021 | Mask-guided dual attention-aware network for visible-infrared person re-identification
Meibin Qi, Suzhi Wang, Guanghong Huang, Jingjing Wu 0001, Cuiqun Chen |
Multim. Tools Appl. | 1 |
| 2020 | A Cross-Modal Multi-granularity Attention Network for RGB-IR Person Re-identification
Meibin Qi, Jingjing Wu 0001, Cuiqun Chen |
Neurocomputing | 3 |
| 2020 | Person Attribute Recognition by Sequence Contextual Relation LearningabstractPerson attribute recognition aims to identify the attribute labels from the pedestrian images. Extracting contextual relation from the images and attributes, including the spatial-semantic relations, the spatial context and the semantic correlation, is beneficial to enhance the discrimination of the features for recognizing the attributes. Thus, this work proposes a sequence contextual relation learning (SCRL) method to capture these relations. It first embeds the images and attributes into sequences in two branches. Then SCRL flexibly learns the contextual relation from the sequences with the parallel attention model structure, which integrates the inter-attention and intra-attention models. The inter-attention module is utilized to extract the spatial-semantic relations, while the intra-attention is designed to gain the spatial context and the semantic correlation. Both attention modules are comprised of several parallel attention units and each unit can obtain the pairwise relations in one subspace. Therefore, they obtain the relations in multiple subspaces, which can improve the comprehensiveness of the relation learning. Additionally, for the sake of better extraction of spatial-semantic relations, this paper employs connectionist temporal classification (CTC) loss which is capable of driving the network to enforce monotonic alignment between the image and attribute. It can also accelerate the convergence of the network by the algorithm in it. Extensive experiments on five public datasets, i.e., Market-1501 attribute, Duke attribute, PETA, RAP and PA-100K datasets, demonstrate the effectiveness of the proposed method. Jingjing Wu 0001, Hao Liu 0003, Meibin Qi, Bo Ren 0002, Xiaohong Li 0002, Yashen Wang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | A Task-Oriented Heuristic for Repairing Infeasible Solutions to Overlapping Coalition Structure GenerationabstractOverlapping coalition formation (OCF), which provides a natural framework for modeling scenarios where each agent can join and allocate their resources to several completely different coalitions at the same time, has become a very active topic in multiagent systems. For OCF in resource-constrained and subadditive task oriented domains, an agent may not possess sufficient resources to meet the needs of multiple coalitions simultaneously. As a result, there may exist many potential resource conflicts among the rival overlapping coalitions. To tackle such situations, we first present a natural variation of the traditional OCF model and analyze the size of the solution space and the computational complexity of the overlapping coalition structure generation (OCSG) problem. Next, we develop a generic task-oriented heuristic (TOH) for individual repairs that can be used in binary meta-heuristic algorithms to generate overlapping coalitions in a parallel manner. Moreover, we show how the proposed TOH repairs a 2-D individual to resolve resource conflicts and discuss several basic properties. Finally, to evaluate the effectiveness of TOH, we compare it with the existing agent-oriented heuristic for the OCSG problem. The empirical results demonstrate that TOH is of high efficiency and effectiveness in harsh environments with fierce competition over scarce resources. Guofu Zhang, Zhaopin Su, Miqing Li, Meibin Qi, Xin Yao 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2019 | Local region partition for person re-identification
Huifang Chu, Meibin Qi, Hao Liu 0003 |
Multim. Tools Appl. | 2 |
| 2019 | Deep feature representation and multiple metric ensembles for person re-identification in security surveillance system
Meibin Qi, Jingxian Han, Hao Liu 0003 |
Multim. Tools Appl. | 1 |
| 2019 | Independent metric learning with aligned multi-part features for video-based person re-identification
Jingjing Wu 0001, Meibin Qi, Hao Liu 0003 |
Multim. Tools Appl. | 3 |
| 2018 | Feature Fusion and Ellipse Segmentation for Person Re-identification
Meibin Qi, Junxian Zeng, Cuiqun Chen |
PRCV (1) | 1 |
| 2018 | Video-Based Person Re-Identification With Accumulative Motion ContextabstractVideo-based person re-identification plays a central role in realistic security and video surveillance. In this paper, we propose a novel accumulative motion context (AMOC) network for addressing this important problem, which effectively exploits the long-range motion context for robustly identifying the same person under challenging conditions. Given a video sequence of the same or different persons, the proposed AMOC network jointly learns appearance representation and motion context from a collection of adjacent frames using a two-stream convolutional architecture. Then, AMOC accumulates clues from motion context by recurrent aggregation, allowing effective information flow among adjacent frames and capturing dynamic gist of the persons. The architecture of AMOC is end-to-end trainable, and thus, motion context can be adapted to complement appearance clues under unfavorable conditions (e.g., occlusions). Extensive experiments are conduced on three public benchmark data sets, i.e., the iLIDS-VID, PRID-2011, and MARS data sets, to investigate the performance of AMOC. The experimental results demonstrate that the proposed AMOC network outperforms state-of-the-arts for video-based re-identification significantly and confirm the advantage of exploiting long-range motion context for video-based person re-identification, validating our motivation evidently. Hao Liu 0003, Zequn Jie, Jayashree Karlekar, Meibin Qi, Shuicheng Yan, Jiashi Feng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Neural Person Search MachinesabstractWe investigate the problem of person search in the wild in this work. Instead of comparing the query against all candidate regions generated in a query-blind manner, we propose to recursively shrink the search area from the whole image till achieving precise localization of the target person, by fully exploiting information from the query and contextual cues in every recursive search step. We develop the Neural Person Search Machines (NPSM) to implement such recursive localization for person search. Benefiting from its neural search mechanism, NPSM is able to selectively shrink its focus from a loose region to a tighter one containing the target automatically. In this process, NPSM employs an internal primitive memory component to memorize the query representation which modulates the attention and augments its robustness to other distracting regions. Evaluations on two benchmark datasets, CUHK-SYSU Person Search dataset and PRW dataset, have demonstrated that our method can outperform current state-of-the-arts in both mAP and top-1 evaluation protocols. Hao Liu 0003, Jiashi Feng, Zequn Jie, Jayashree Karlekar, Bo Zhao 0032, Meibin Qi, Shuicheng Yan |
ICCV | 6 |
| 2017 | End-to-End Comparative Attention Networks for Person Re-IdentificationabstractPerson re-identification across disjoint camera views has been widely applied in video surveillance yet it is still a challenging problem. One of the major challenges lies in the lack of spatial and temporal cues, which makes it difficult to deal with large variations of lighting conditions, viewing angles, body poses, and occlusions. Recently, several deep-learning-based person re-identification approaches have been proposed and achieved remarkable performance. However, most of those approaches extract discriminative features from the whole frame at one glimpse without differentiating various parts of the persons to identify. It is essentially important to examine multiple highly discriminative local regions of the person images in details through multiple glimpses for dealing with the large appearance variance. In this paper, we propose a new soft attention-based model, i.e., the end-to-end comparative attention network (CAN), specifically tailored for the task of person re-identification. The end-to-end CAN learns to selectively focus on parts of pairs of person images after taking a few glimpses of them and adaptively comparing their appearance. The CAN model is able to learn which parts of images are relevant for discerning persons and automatically integrates information from different parts to determine whether a pair of images belongs to the same person. In other words, our proposed CAN model simulates the human perception process to verify whether two images are from the same person. Extensive experiments on four benchmark person re-identification data sets, including CUHK01, CHUHK03, Market-1501, and VIPeR, clearly demonstrate that our proposed end-to-end CAN for person re-identification outperforms well established baselines significantly and offer the new state-of-the-art performance. Hao Liu 0003, Jiashi Feng, Meibin Qi, Shuicheng Yan |
IEEE Trans. Image Process. | 3 |
| 2015 | Using binary particle swarm optimization to search for maximal successful coalition
Guofu Zhang, Renzhi Yang, Zhaopin Su, Meibin Qi |
Appl. Intell. | 6 |
| 2015 | Kernelized Relaxed Margin Components Analysis for Person Re-identificationabstractPerson re-identification across disjoint camera views plays a significant role in video surveillance. Several margin-based metric learning algorithms have recently been proposed to learn an optimal metric, with the goal that samples of the same person always belong to the same class while those from different classes are separated by a large margin. These approaches require no modification or extension in order to solve problems of multiple (as opposed to binary) classification. However, the formation of the margin in these methods is not scalable, and thus cannot adequately use inter-class information according to the relevant practical application. To address this issue, we propose a novel algorithm called Relaxed Margin Components Analysis (RMCA) to “relax” the margin constraint. Furthermore, we equip our RMCA with a kernel function to form a Kernelized RMCA (KRMCA) to learn non-linear distance metrics in order to further improve re-identification accuracy. Promising results from experiments on several public datasets demonstrate the effectiveness of our method. Hao Liu 0003, Meibin Qi |
IEEE Signal Process. Lett. | 2 |
| 2014 | Person Re-identification based on nonlinear ranking with difference vectors
Tianfeng Zhou, Meibin Qi, Shijie Hao, Yulong Jin |
Inf. Sci. | 2 |
| 2010 | Searching for overlapping coalitions in multiple virtual organizations
Guofu Zhang, Zhaopin Su, Meibin Qi |
Inf. Sci. | 4 |