EDBT 2026 Demo / reviewers in the wild / expert
Jingjing Wu 0001
dblp:27/2384-1
· DBLP profile ↗
27ranked-venue papers
9as first author
24since 2021 · last 2026
0000-0002-3818-4277ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Computer networks · 5 · 3 first-author · 5 since 2021Security and privacy · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generalized adversarial feature aggregation and KAN-enhanced network for semi-supervised visible-infrared person re-identification
Rui Sun 0004, Jicheng Shen, Guoxi Huang, Jingjing Wu 0001 |
Image Vis. Comput. | 5 |
| 2026 | Interview-Based Depression Detection Using LLM-Based Text Restatement and Emotion LexiconabstractDepression is a mental health disorder that significantly impacts modern society. Developing accurate depression detection models by leveraging discriminant features from multimedia or physiological data can aid medical professionals in making informed diagnoses. According to psychological studies, emotion is a critical indicator of depression. However, emotion has not been utilized as a central role in current research on assistive depression detection, usually serving as a supplementary information source or a guidance for integrating diverse data modalities. In contrast to existing studies, we investigate the feasibility of detecting depression by concentrating on emotion information. Specifically, focusing on modeling emotion feature representation during interviews, we propose an interview-based depression detection model via leveraging large language (LLM) based text restatement and emotion lexicon (IDD-LTE). In this model, we employ LLM to enhance text quality through restatement to address the potentially low quality of interview text data. Using an emotion lexicon, the open contents in restated texts are mapped to a fixed-size matrix representation that captures the interviewee's emotional state and mood swings during the conversation, serving as the fundamental representation for the following discriminant feature learning. The proposed IDD-LTE model is evaluated on four primary datasets for depression detection. The promising results confirm the feasibility and effectiveness of our model. Shijie Hao, Jingjing Wu 0001, Yanrong Guo, Richang Hong |
IEEE Trans. Affect. Comput. | 3 |
| 2026 | Subthreshold Depression Detection With Text-Guided Multimodal LearningabstractDepression, a widespread global mental health problem, affects millions of people annually, making early detection of subclinical depression crucial for timely intervention. Current automatic depression detection (ADD) methods, valuable for diagnosis, often neglect subthreshold populations and face difficulties in extracting diagnostic data from long sequences of multimodal information. These methods also inadequately leverage text modality, which is less noisy and information-rich compared with other modalities. Furthermore, existing datasets for depression research are often too small, limiting the generalizability of developed methods. To address these issues, this article proposes a new approach for detecting depression in subthreshold populations. For long-sequence samples in the field of depression, we construct an autoencoder that compresses along both temporal and feature dimensions, aiming to extract the most compact and effective features from the samples. To exploit the text modality’s advantages, we integrate the RoBERTa pretrained model with an attention mechanism for high-quality text encoding. We then develop a text-guided multimodal fusion (TGMF) module, using text encoding as an anchor for guiding audio and video modality encoding, ensuring multimodal alignment. Additionally, contrastive learning is applied to discern differences between classes, enhancing the model’s generalizability. Our method demonstrates superior performance in the tasks of detecting depression and identifying subthreshold populations on the E-DAIC and MMDA datasets. Yanrong Guo, Youwei Guo, Bingxin Yang, Jingjing Wu 0001, Shijie Hao, Richang Hong |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2026 | Learning Corruption-Invariant Components and Cross-Modal Correspondence for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised Visible-Infrared Person Re-Identification (US-VI-ReID) has great potential prospects because it does not require label information. However, corrupted pedestrian images collected due to corruption factors in real-world scenarios (e.g., noise, blur, and weather changes) largely limit the scalability of US-VI-ReID. In this paper, we explore the robustness of US-VI-ReID for the first time and propose a Multi-Granularity Spatial-Frequency Prototype Learning (MSPL) framework. The framework mainly consists of Multi-Channel Soft Augmentation (MSA), Robust Frequency Domain Feature Learning (RFL) module and Cross-modal Spatial-Frequency Prototype Matching (CSPM). Specifically, the MSA alleviates the sensitivity of model to color and abnormal samples through rich channel combinations and soft erasing. Subsequently, the RFL performs deep global filtering and amplitude attention compensated InstanceNorm to complete frequency and style modulation, concentrating on degradation-robust frequency content. Finally, the CSPM is designed to achieve multi-granularity prototype contrastive learning on cluster level and view level, then conduct cross-modal matching of multi-granularity spatial-frequency prototypes, thus establishing robust label association. With the above modules, our proposed framework can learn corruption-invariant feature components and generate robust cross-modal correspondence from unlabeled cross-modal images. Extensive experiments demonstrate that our MSPL outperforms other state-of-the-art methods by a large margin on the challenging SYSU-MM01-C and RegDB-C, while maintaining competitive on the SYSU-MM01 and RegDB. Rui Sun 0004, Guoxi Huang, Jingjing Wu 0001, Wei Jia 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2026 | Dynamic Correlation-Guided Disentanglement and Contrastive Learning for RGB-D Cross-Modal Re-IdentificationabstractPerson re-identification (Re-ID) across RGB and depth modalities offers complementary cues for robust pedestrian matching under challenging conditions. However, the significant discrepancy between RGB appearance features and depth structural features complicates cross-modal alignment. Existing methods either depend on static architectural designs or impose strong constraints to capture the common features of the two modalities, often suffering from branch imbalance or distorted identity features. In this work, we propose a novel framework, Dynamic Correlation-Guided Disentanglement and Contrastive Learning (DCG-DCL), for RGB-D cross-modal Re-ID. First, the Dynamic Correlation-guided Disentanglement (DCGD) dynamically decouples features with the guidance of inter-modal correlation, which explicitly enforces common-feature learning via a cross-correlation constraint and adaptively separates common and unique components without predefined assumptions. Second, a Common & Unique Contrastive Learning (CUCL) strategy fully leverages these decoupled features, which aligns RGB/depth features closer to their common representation and pushes them away from unique redundancies. This dual mechanism effectively narrows modality discrepancy and boosts robustness against modality-specific noise. Extensive experiments on multiple public benchmarks demonstrate that our method achieves state-of-the-art performance, with ablation studies validating the necessity of each component. Zhibo Lei, Jingjing Wu 0001, Yaxiong Wang, Yanrong Guo, Shijie Hao, Richang Hong |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2026 | Infrared Object Tracking via Complementary Dual-domain Interaction with Target-guided Frequency TransformationabstractInfrared Object Tracking (IOT) is challenging due to the low contrast of infrared images, which limits effective spatial feature extraction. Although recent works have explored frequency-domain information, their utilization remains insufficient, and fusion strategies either retain redundancy or fail to fully explore distinctive differences, thus limiting complementary enhancement. To overcome this, we propose a novel tracker that introduces a Target-guided Frequency Transformation Module (TFTM) and a Dual-domain Interactive Fusion Network (DIFN). The former extracts multi-frequency representations across scales and orientations, guided by an adaptive mask strategy to suppress background interference. The latter fuses the two domains with differentiated attention to achieve complementary enhancement. Extensive experiments show that our approach achieves superior performance over state-of-the-art trackers, highlighting the effectiveness of comprehensive frequency-domain integration in IOT. Pengyu Huang, Jingjing Wu 0001, Yanrong Guo, Richang Hong |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Discrete to Continuous: Generating Smooth Transition Poses from Sign Language ObservationsabstractGenerating continuous sign language videos from discrete segments is challenging due to the need for smooth transitions that preserve natural flow and meaning. Traditional approaches that simply concatenate isolated signs often result in abrupt transitions, disrupting video coherence. To address this, we propose a novel framework, Sign-D2C, that employs a conditional diffusion model to synthesize contextually smooth transition frames, enabling the seamless construction of continuous sign language sequences. Our approach transforms the unsupervised problem of transition frame generation into a supervised training task by simulating the absence of transition frames through random masking of segments in long-duration sign videos. The model learns to predict these masked frames by denoising Gaussian noise, conditioned on the surrounding sign observations, allowing it to handle complex, unstructured transitions. During inference, we apply a linearly interpolating padding strategy that initializes missing frames through interpolation between boundary frames, providing a stable foundation for iterative refinement by the diffusion model. Extensive experiments on the PHOENIX14T, USTCCSL100, and USTC-SLR500 datasets demonstrate the effectiveness of our method in producing continuous, natural sign language videos. Shengeng Tang, Lechao Cheng, Jingjing Wu 0001, Dan Guo 0001, Richang Hong |
CVPR | 4 |
| 2025 | SRConvNet: A Transformer-Style ConvNet for Lightweight Image Super-Resolution
Feng Li 0037, Runmin Cong, Jingjing Wu 0001, Huihui Bai 0001, Meng Wang 0001, Yao Zhao 0001 |
Int. J. Comput. Vis. | 3 |
| 2025 | Implicit Alignment-Based Cross-Modal Symbiotic Network for Text-to-Image Person Re-IdentificationabstractText-to-image person re-identification aims to utilize textual descriptions to retrieve specific person images from large image databases. The core challenge of this task lies in the significant feature differences between the abstract nature of text and the intuitiveness of images. Existing solutions primarily rely on explicit alignment of global or fine-grained local features, which lack flexibility and struggle to effectively capture and leverage subtle features and relationship information in multimodal data. Particularly, for different images of the same person, the emphasis in feature extraction should be adjusted according to the differences in text descriptions. To address these issues, this paper proposes a Cross-Modal Symbiotic Network (CMSN) based on implicit alignment. First, CMSN employs an Implicit Multi-scale Feature Integration (IMFI) module to implicitly extract and fuse multiscale features from images and text, thereby adaptively capturing the feature relationships between the two modalities. Second, a Combined Representation Learning (CRL) module is used to produce a combined representation of the text and image features, utilizing a Combined-Representation Identity Alignment (CRIA) loss to align and constrain the identity centers of the three feature vectors. Finally, we design a Semi-Positive Triplet (SPT) loss function, which defines semi-positive samples using other images and texts of the same identity, providing additional supervisory information to the model and further reducing modality heterogeneity. Extensive experiments on the CUHK-PEDES dataset demonstrate that CMSN achieves an impressive Rank-1 and mAP accuracy of 76.46% and 70.28%, respectively, significantly outperforming existing SOTA methods. Rui Sun 0004, Guoxi Huang, Jingjing Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Person Re-Identification With Arbitrary Modalities: A Multi-Modal Dataset and a Unified FrameworkabstractThis paper proposes a unified visual person re-identification (re-id) framework capable of handling various re-id tasks, including modal-fusion re-id, cross-modal re-id, and single-modal re-id, to accommodate diverse modal scenarios. We begin by constructing a Multi-modal Person Re-identification (MPR) dataset comprising RGB, infrared (IR), and depth modalities. Then, the unified re-id framework is established by integrating an Adaptive Modality Aggregation Module (AMAM) and Multi-modal Auto-aligned Learning (MAL). The former autonomously aggregates distinct modalities by thoroughly exploring their relationships. It not only benefits modal-fusion re-id by promoting the modal-fusion representations, but also enhances cross-modal re-id by performing modal consistency learning on the modal-fusion features to narrow modal gaps. The latter automatically aligns multiple modalities through contrastive learning constraints to lessen modal gaps for multiple cross-modal re-id tasks. So, these two modules respectively balance the tasks of distinct types and various tasks of the same type, which are beneficial to realize more re-id tasks with diverse modal scenarios. Moreover, we evaluate state-of-the-art (SOTA) multi-modal methods in terms of plentiful testing settings constructed on MPR dataset. The experiments demonstrate that the proposed unified method that only needs to be trained once outperforms existing methods that require multiple training processes with specific modalities. Besides, it can cope with more scenarios. Extensive ablation studies investigate the effects of the proposed modules on all re-id tasks. Our datasets and code will be publicly available soon: https://github.com/hfutwujingjing/A-Multi-Modal-Dataset-and-A-Unified-Framework. Jingjing Wu 0001, Zhun Zhong, Yanrong Guo, Shejiao Hu, Richang Hong |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Local Fine-Grained Visual TrackingabstractThis paper introduces a novel local fine-grained visual tracking task, aiming to precisely locate arbitrary local parts of objects. This task is motivated by our observation that in many realistic scenarios, the user demands to track a local part instead of a holistic object. However, the absence of an evaluation dataset and the distinctive characteristics of local fine-grained targets present extra challenges in conducting this research. To tackle these issues, first, this paper constructs a local fine-grained tracking (LFT) dataset to evaluate the tracking performance for local fine-grained targets. Second, this paper designs a cutting-edge solution to handle the challenges posed by properties of local objects, including ambiguity and high-proportion backgrounds. It consists of a hierarchical adaptive mask mechanism and foreground-background differentiated learning. The former adaptively searches for and masks ambiguity, which drives the network to concentrate on the local target instead of the holistic objects. The latter is constructed to distinguish foreground and background in an unsupervised manner, which is beneficial to mitigate the impacts of high-proportion backgrounds. Extensive analytic experiments are performed to verify the effectiveness of each submodule in the proposed fine-grained tracker. Jingjing Wu 0001, Richang Hong |
IEEE Trans. Multim. | 1 |
| 2025 | Gloss-driven Conditional Diffusion Models for Sign Language ProductionabstractSign Language Production (SLP) aims to convert text or audio sentences into sign language videos corresponding to their semantics, which is challenging due to the diversity and complexity of sign languages, and cross-modal semantic mapping issues. In this work, we propose a Gloss-driven Conditional Diffusion Model (GCDM) for SLP. The core of the GCDM is a diffusion model architecture, in which the sign gloss sequence is encoded by a Transformer-based encoder and input into the diffusion model as a semantic prior condition. In the process of sign pose generation, the textual semantic priors carried in the encoded gloss features are integrated into the embedded Gaussian noise via cross-attention. Subsequently, the model converts the fused features into sign language pose sequences through T-round denoising steps. During the training process, the model uses the ground-truth labels of sign poses as the starting point, generates Gaussian noise through T rounds of noise, and then performs T rounds of denoising to approximate the real sign language gestures. The entire process is constrained by the MAE loss function to ensure that the generated sign language gestures are as close as possible to the real labels. In the inference phase, the model directly randomly samples a set of Gaussian noise, generates multiple sign language gesture sequence hypotheses under the guidance of the gloss sequence, and outputs a high-confidence sign language gesture video by averaging multiple hypotheses. Experimental results on the Phoenix2014T dataset show that the proposed GCDM method achieves competitiveness in both quantitative performance and qualitative visualization. Shengeng Tang, Feng Xue 0002, Jingjing Wu 0001, Shuo Wang 0008, Richang Hong |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Improving Consistency of Proxy-Level Contrastive Learning for Unsupervised Person Re-IdentificationabstractRecently, contrastive learning-based unsupervised person re-identification (Re-ID) methods have garnered significant attention due to their effectiveness. These methods rely on predicted pseudo-labels to construct contrastive pairs, optimizing the network gradually. Some methods also utilize camera labels to explore intra-camera and inter-camera contrastive relations, achieving state-of-the-art results. However, these methods fail to address the issue of inconsistency in proxy-level contrastive learning, which arises from variations in the distribution of instances belonging to the same proxy. Specifically, they are sensitive to the distribution of instances in a mini-batch used for contrastive pair construction, and uncertainty or noise in the data distribution can lead to turbulence in the contrastive loss, degrading the effectiveness of contrastive learning. In this work, we first propose a dual-branch contrastive learning (DBCL) framework. The framework comprises a dual-branch structure with an identity discrimination branch and a camera view awareness branch. These branches are mutually trained to produce a jointly optimized model with both high person identification accuracy and cross-camera robustness. Moreover, to mitigate the proxy-level contrastive inconsistency issue in the camera view awareness branch, we design intra-camera and inter-camera consistent contrastive losses. Our DBCL has been extensively evaluated on several person Re-ID datasets and has demonstrated superior performance compared to state-of-the-art methods. Notably, on the challenging MSMT17 dataset with complex scenes, our method achieved an mAP of 45.3% and Rank-1 accuracy of 75.3%. Yimin Liu 0001, Meibin Qi, Yongle Zhang 0001, Qiang Wu 0001, Jingjing Wu 0001, Shuo Zhuang |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Intermediary-Generated Bridge Network for RGB-D Cross-Modal Re-IdentificationabstractRGB-D cross-modal person re-identification (re-id) targets at retrieving the person of interest across RGB and depth image modalities. To cope with the modal discrepancy, some existing methods generate an auxiliary mode with either inherent properties of input modes or extra deep networks. However, such useful intermediary role included in generated mode is often overlooked in these approaches, leading to insufficient exploitation of crucial bridge knowledge. By contrast, in this article, we propose a novel approach that constructs an intermediary mode through the constraints of self-supervised intermediary learning, which is freedom from modal prior knowledge and additional module parameters. We then design a bridge network to fully mine the intermediary role of generated modality through carrying out multi-modal integration and decomposition. For one thing, this network leverages a multi-modal transformer to integrate the information of three modes via fully exploiting their heterogeneous relations with the intermediary mode as the bridge. It conducts the identification consistency constraint to promote cross-modal associations. For another, it employs circle contrastive learning to decompose the cross-modal constraint process into several subprocedures, which provides the intermediate relay during pulling two original modalities closer. Experiments on two public datasets demonstrate that the proposed method exceeds the state-of-the-arts. The effectiveness of each component in this method is verified through numerous ablation studies. Additionally, we have demonstrated the generalization ability of the proposed method through experiments. Jingjing Wu 0001, Richang Hong, Shengeng Tang |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2024 | SYRER: Synergistic Relational Reasoning for RGB-D Cross-Modal Re-IdentificationabstractRGB-D cross-modal person re-identification is designed to match the people across the RGB and depth image modalities, where the large modality discrepancy makes this task intractable to tackle. To alleviate the negative effect brought by the discrepancy, this paper proposes a novel SYnergistic RElational Reasoning (SYRER) method, which targets at exploring the synergy between hetero-modalities for recognizing persons. We design a heterogeneous relationship contrast branch to establish intra-class and inter-class cross-modal relationships, which implements the cross-modal relation contrast learning to cope with imperceptible cross-modal inter-class differences and large cross-modal intra-class discrepancy. Additionally, in order to adequately represent the irregular depth images, we propose a point-wise depth extractor to extract non-uniform discriminative point features from depth images. Experimental results on two public datasets indicate the proposed SYRER surpasses the state-of-the-arts. And we also perform a series of analytic experiments to verify the effectiveness of each submodule of our SYRER. Hao Liu 0003, Jingjing Wu 0001, Feng Li 0037, Richang Hong |
IEEE Trans. Multim. | 2 |
| 2024 | Asymmetric Deformable Spatio-temporal Framework for Infrared Object TrackingabstractThe Infrared Object Tracking (IOT) task aims to locate objects in infrared sequences. Since color and texture information is unavailable in infrared modality, most existing infrared trackers merely rely on capturing spatial contexts from the image to enhance feature representation, where other complementary information is rarely deployed. To fill this gap, we in this article propose a novel Asymmetric Deformable Spatio-Temporal Framework (ADSF) to fully exploit collaborative shape and temporal clues in terms of the objects. Firstly, an asymmetric deformable cross-attention module is designed to extract shape information, which attends to the deformable correlations between distinct frames in an asymmetric manner. Secondly, a spatio-temporal tracking framework is coined to learn the temporal variance trend of the object during the training process and store the template information closest to the tracking frame when testing. Comprehensive experiments demonstrate that ADSF outperforms state-of-the-art methods on three public datasets. Extensive ablation experiments further confirm the effectiveness of each component in ADSF. Furthermore, we conduct generalization validation to demonstrate that the proposed method also achieves performance gains in RGB-based tracking scenarios. Jingjing Wu 0001, Xi Zhou 0004, Xiaohong Li 0002, Hao Liu 0003, Meibin Qi, Richang Hong |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | Saliency and Granularity: Discovering Temporal Coherence for Video-Based Person Re-IdentificationabstractVideo-based person re-identification (ReID) matches the same people across the video sequences with rich spatial and temporal information in complex scenes. It is highly challenging to capture discriminative information when occlusions and pose variations exist between frames. A key solution to this problem rests on extracting the temporal invariant features of video sequences. In this paper, we propose a novel method for discovering temporal coherence by designing a region-level saliency and granularity mining network (SGMN). Firstly, to address the varying noisy frame problem, we design a temporal spatial-relation module (TSRM) to locate frame-level salient regions, adaptively modeling the temporal relations on spatial dimension through a probe-buffer mechanism. It avoids the information redundancy between frames and captures the informative cues of each frame. Secondly, a temporal channel-relation module (TCRM) is proposed to further mine the small granularity information of each frame, which is complementary to TSRM by concentrating on discriminative small-scale regions. TCRM exploits a one-and-rest difference relation on channel dimension to enhance the granularity features, leading to stronger robustness against misalignments. Finally, we evaluate our SGMN with four representative video-based datasets, including iLIDS-VID, MARS, DukeMTMC-VideoReID, and LS-VID, and the results indicate the effectiveness of the proposed method. Cuiqun Chen, Mang Ye, Meibin Qi, Jingjing Wu 0001, Yimin Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Structure-Aware Positional Transformer for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is a cross-modality retrieval problem, which aims at matching the same pedestrian between the visible and infrared cameras. Due to the existence of pose variation, occlusion, and huge visual differences between the two modalities, previous studies mainly focus on learning image-level shared features. Since they usually learn a global representation or extract uniformly divided part features, these methods are sensitive to misalignments. In this paper, we propose a structure-aware positional transformer (SPOT) network to learn semantic-aware sharable modality features by utilizing the structural and positional information. It consists of two main components: attended structure representation (ASR) and transformer-based part interaction (TPI). Specifically, ASR models the modality-invariant structure feature for each modality and dynamically selects the discriminative appearance regions under the guidance of the structure information. TPI mines the part-level appearance and position relations with a transformer to learn discriminative part-level modality features. With a weighted combination of ASR and TPI, the proposed SPOT explores the rich contextual and structural information, effectively reducing cross-modality difference and enhancing the robustness against misalignments. Extensive experiments indicate that SPOT is superior to the state-of-the-art methods on two cross-modal datasets. Notably, the Rank-1/mAP value on the SYSU-MM01 dataset has improved by 8.43%/6.80%. Cuiqun Chen, Mang Ye, Meibin Qi, Jingjing Wu 0001, Chia-Wen Lin |
IEEE Trans. Image Process. | 4 |
| 2022 | Improving Feature Discrimination for Object Tracking by Structural-similarity-based Metric LearningabstractExisting approaches usually form the tracking task as an appearance matching procedure. However, the discrimination ability of appearance features is insufficient in these trackers, which is caused by their weak feature supervision constraints and inadequate exploitation of spatial contexts. To tackle this issue, this article proposes a novel appearance matching tracking (AMT) method to strengthen the feature restraints and capture discriminative spatial representations. Specifically, we first utilize a triplet structural loss function, which improves the learning capability of features by applying a structural similarity constraint with a triplet metric format on the features. It leverages feature statistics to capture the complex interactions of visual parts. Second, we put forward an adaptive matching module that exploits the dual spatial enhancement module to reinforce target feature discrimination. This not only boosts the representation ability of spatial context but also realizes spatially dynamic feature selection by attending to target deformation information. Moreover, this model introduces a simple but effective matching unit to intuitively evaluate the relative appearance differences between the target and the proposals. In addition, with the obtained discriminative features, AMT is capable of providing precise localization for the target. Therefore, the impact of spatial suppression imposed by window functions can be alleviated, allowing for effective tracking of high-speed moving objects. Extensive experiments prove that AMT outperforms state-of-the-art methods on six public datasets and demonstrate the effectiveness of each component in AMT. Jingjing Wu 0001, Meibin Qi, Cuiqun Chen, Yimin Liu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | An End-to-end Heterogeneous Restraint Network for RGB-D Cross-modal Person Re-identificationabstractThe RGB-D cross-modal person re-identification (re-id) task aims to identify the person of interest across the RGB and depth image modes. The tremendous discrepancy between these two modalities makes this task difficult to tackle. Few researchers pay attention to this task, and the deep networks of existing methods still cannot be trained in an end-to-end manner. Therefore, this article proposes an end-to-end module for RGB-D cross-modal person re-id. This network introduces a cross-modal relational branch to narrow the gaps between two heterogeneous images. It models the abundant correlations between any cross-modal sample pairs, which are constrained by heterogeneous interactive learning. The proposed network also exploits a dual-modal local branch, which aims to capture the common spatial contexts in two modalities. This branch adopts shared attentive pooling and mutual contextual graph networks to extract the spatial attention within each local region and the spatial relations between distinct local parts, respectively. Experimental results on two public benchmark datasets, that is, the BIWI and RobotPKU datasets, demonstrate that our method is superior to the state-of-the-art. In addition, we perform thorough experiments to prove the effectiveness of each component in the proposed method. Jingjing Wu 0001, Meibin Qi, Cuiqun Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | Towards accurate estimation for visual object tracking with multi-hierarchy feature aggregation
Jingjing Wu 0001, Meibin Qi, Xiaohong Li 0002 |
Neurocomputing | 1 |
| 2021 | Global-Local Graph Convolutional Network for cross-modality person re-identification
Xiaohong Li 0002, Cuiqun Chen, Meibin Qi, Jingjing Wu 0001 |
Neurocomputing | 5 |
| 2021 | Learning discriminative features with a dual-constrained guided network for video-based person re-identification
Cuiqun Chen, Meibin Qi, Guanghong Huang, Jingjing Wu 0001, Xiaohong Li 0002 |
Multim. Tools Appl. | 4 |
| 2021 | Mask-guided dual attention-aware network for visible-infrared person re-identification
Meibin Qi, Suzhi Wang, Guanghong Huang, Jingjing Wu 0001, Cuiqun Chen |
Multim. Tools Appl. | 5 |
| 2020 | A Cross-Modal Multi-granularity Attention Network for RGB-IR Person Re-identification
Meibin Qi, Jingjing Wu 0001, Cuiqun Chen |
Neurocomputing | 5 |
| 2020 | Person Attribute Recognition by Sequence Contextual Relation LearningabstractPerson attribute recognition aims to identify the attribute labels from the pedestrian images. Extracting contextual relation from the images and attributes, including the spatial-semantic relations, the spatial context and the semantic correlation, is beneficial to enhance the discrimination of the features for recognizing the attributes. Thus, this work proposes a sequence contextual relation learning (SCRL) method to capture these relations. It first embeds the images and attributes into sequences in two branches. Then SCRL flexibly learns the contextual relation from the sequences with the parallel attention model structure, which integrates the inter-attention and intra-attention models. The inter-attention module is utilized to extract the spatial-semantic relations, while the intra-attention is designed to gain the spatial context and the semantic correlation. Both attention modules are comprised of several parallel attention units and each unit can obtain the pairwise relations in one subspace. Therefore, they obtain the relations in multiple subspaces, which can improve the comprehensiveness of the relation learning. Additionally, for the sake of better extraction of spatial-semantic relations, this paper employs connectionist temporal classification (CTC) loss which is capable of driving the network to enforce monotonic alignment between the image and attribute. It can also accelerate the convergence of the network by the algorithm in it. Extensive experiments on five public datasets, i.e., Market-1501 attribute, Duke attribute, PETA, RAP and PA-100K datasets, demonstrate the effectiveness of the proposed method. Jingjing Wu 0001, Hao Liu 0003, Meibin Qi, Bo Ren 0002, Xiaohong Li 0002, Yashen Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Independent metric learning with aligned multi-part features for video-based person re-identification
Jingjing Wu 0001, Meibin Qi, Hao Liu 0003 |
Multim. Tools Appl. | 1 |