Yifan Jiao

dblp:214/4159 · DBLP profile ↗
← Back
19ranked-venue papers
11as first author
16since 2021 · last 2026
0000-0002-8923-6997ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 12 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Computer networks · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 LEPD-Net: A Lightweight and Efficient Network for Pedestrian Detection
abstract
The pedestrian detection is crucial in practical applications, such as autonomous driving and video surveillance. However, the existing research mainly focuses on improving detection accuracy, with relatively little attention paid to model complexity and operational efficiency. In scenarios with high real-time requirements, the practical deployment of pedestrian detectors still faces many difficulties. To this end, we propose a lightweight and efficient pedestrian detection network (LEPD-Net). First, we design a PoolFormer-based detection head (PDH) to reduce the model computation and inference time. Second, to compensate for the deficiency of PDH in global context modeling, we design a triple-branch joint attention module (TJAM). TJAM uses only a small number of parameters and strengthens the model's contextual representation by capturing spatial location dependencies and global semantic information between channels. Finally, after incorporating PDH and TJAM into the backbone network, a lightweight and efficient pedestrian detector is constructed. We benchmarked the model on mainstream pedestrian datasets Caltech and CityPersons. The results show that our model achieves the current state-of-the-art performance level. In addition, our model reduces inference time by 25% while maintaining accuracy.
Wenliang Ge, Shucheng Huang, Mingxing Li 0001, Yifan Jiao
IEEE Trans. Neural Networks Learn. Syst.4
2025 GSOT3D: Towards Generic 3D Single Object Tracking in the Wild
abstract
In this paper, we present a novel benchmark, GSOT3D, that aims at facilitating development of generic 3D single object tracking (SOT) in the wild. Specifically, GSOT3D offers 620 sequences with 123K frames, and covers a wide selection of 54 object categories. Each sequence is offered with multiple modalities, including the point cloud (PC), RGB image, and depth. This allows GSOT3D to support various 3D tracking tasks, such as single-modal 3D SOT on PC and multi-modal 3D SOT on RGB-PC or RGB-D, and thus greatly broadens research directions for 3D object tracking. To provide highquality per-frame 3D annotations, all sequences are labeled manually with multiple rounds of meticulous inspection and refinement. To our best knowledge, GSOT3D is the largest benchmark dedicated to various generic 3D object tracking tasks. To understand how existing 3D trackers perform and to provide comparisons for future research on GSOT3D, we assess eight representative point cloud-based tracking models. Our evaluation results exhibit that these models heavily degrade on GSOT3D, and more efforts are required for robust and generic 3D object tracking. Besides, to encourage future research, we present a simple yet effective generic 3D tracker, named PROT3D, that localizes the target object via a progressive spatial-temporal network and outperforms all current solutions by a large margin. By releasing GSOT3D, we expect to advance further 3D tracking in future research and applications. Our benchmark and model as well as the evaluation results will be publicly released at our webpage https://github.com/ailovejinx/GSOT3D.
Yifan Jiao, Junhua Ding 0001, Qing Yang 0003, Song Fu, Heng Fan 0001, Libo Zhang 0001
ICCV1
2025 Attention to Trajectory: Trajectory-Aware Open-Vocabulary Tracking
abstract
Open-Vocabulary Multi-Object Tracking (OV-MOT) aims to enable approaches to track objects without being limited to a predefined set of categories. Current OV-MOT methods typically rely primarily on instance-level detection and association, often overlooking trajectory information that is unique and essential for object tracking tasks. Utilizing trajectory information can enhance association stability and classification accuracy, especially in cases of occlusion and category ambiguity, thereby improving adaptability to novel classes. Thus motivated, in this paper we propose \textbf{TRACT}, an open-vocabulary tracker that leverages trajectory information to improve both object association and classification in OV-MOT. Specifically, we introduce a \textit{Trajectory Consistency Reinforcement} (\textbf{TCR}) strategy, that benefits tracking performance by improving target identity and category consistency. In addition, we present \textbf{TraCLIP}, a plug-and-play trajectory classification module. It integrates \textit{Trajectory Feature Aggregation} (\textbf{TFA}) and \textit{Trajectory Semantic Enrichment} (\textbf{TSE}) strategies to fully leverage trajectory information from visual and language perspectives for enhancing the classification results. Extensive experiments on OV-TAO show that our TRACT significantly improves tracking performance, highlighting trajectory information as a valuable asset for OV-MOT. Code will be released.
Yifan Jiao, Dan Meng 0001, Heng Fan 0001, Libo Zhang 0001
ICCV2
2025 PlanarTrack: A high-quality and challenging benchmark for large-scale planar object tracking
Yifan Jiao, Xiaoqiong Liu, Xiaohui Yuan 0001, Heng Fan 0001, Libo Zhang 0001
Comput. Vis. Image Underst.1
2025 Ms-VLPD: A multi-scale VLPD based method for pedestrian detection
Shucheng Huang, Senbao Zhang, Yifan Jiao
Expert Syst. Appl.3
2025 A multi-scale network with multi-view correlation for vehicle re-identification
Shucheng Huang, Yifan Jiao, Mingxing Li 0001
Multim. Syst.4
2025 SVSRD: Spatial Visual and Statistical Relation Distillation for Class-Incremental Semantic Segmentation
abstract
Class-incremental semantic segmentation (CISS) aims to incrementally learn novel classes while retaining the ability to segment old classes, and suffers catastrophic forgetting since the old-class labels are unavailable. Most existing methods typically impose strict constraints on the consistency between the extracted features or output logits of each pixel from old and current models in an attempt to prevent forgetting through knowledge distillation (KD), which 1) results in a significant transfer of redundant knowledge while limiting the restoration of old classes (rigidity) due to potentially overlooking essential knowledge extraction, and 2) imposes strong constraints at the pixel level making it challenging for the model to learn novel classes (plasticity). To solve the above limitations, we propose a novel Spatial Visual and Statistical Relation Distillation (SVSRD) by applying multi-scale visual and statistical position relation distillation for CISS, which enjoys several merits. First, we introduce a region-based similarity matrix and impose a consistency constraint between current and old models, which preserves the essential visual knowledge to enhance the rigidity. Second, we propose a novel statistical feature calculation algorithm to investigate the distribution of the data and further preserve the rules of statistics through statistical consistency, which also promotes the model on the novel-class learning for improving the plasticity. Finally, the aforementioned constraints are jointly applied in multiple scales to alleviate old-class forgetting and enhance novel-class learning. Extensive experiments on Pascal-VOC 2012 and ADE20 K demonstrate that the proposed approach performs favorably against the state-of-the-art CISS methods.
Yuyang Chang, Yifan Jiao, Bing-Kun Bao
IEEE Trans. Multim.2
2025 Replay-Based Incremental Object Detection With Local Response Exploration
abstract
Incremental object detection (IOD) aims to train an object detector on non-stationary data streams without forgetting previous knowledge. Prevalent replay-based methods keep a buffer composed of carefully selected instances towards this goal. However, due to the limited storage space and uniform feature distribution, existing methods are prone to overfit on replayed instances, leading to poor generalization on diverse test data. Additionally, the imbalance in data quantity makes the detector fail to distinguish old and new classes that are visually similar, introducing bias toward new classes. To enhance the diversity of stored instances and eliminate bias, we propose a Local Response Exploration (LRE) framework, which comprises three modules. First, Region-Entropy Instance Selector (REIS) introduces a novel metric to assess instance diversity based on the entropy of local responses. Second, Confusion-Guided Instance Replay (CGIR) replaces the previous random replay approach by replaying specific old class instances based on class similarity, ensuring that parameters for similar new and old classes are updated together, thereby mitigating bias and helping mining discriminative patterns. Third, Confusion-Aware Region Segregation (CARS) adaptively differentiates biased regions from other regions based on local responses, reducing bias toward new classes while preserving relationships between new and old classes. Extensive evaluations on Pascal-VOC and MS COCO datasets demonstrate that our approach outperforms State-of-the-Art methods in incremental object detection.
Yifan Jiao, Bing-Kun Bao
IEEE Trans. Multim.2
2025 Unified Text-Image Space Alignment with Cross-Modal Prompting in CLIP for UDA
abstract
Unsupervised Domain Adaptation (UDA) aims to transfer models trained on a labeled source domain to an unlabeled target domain. Due to the excellent generalization ability of Vision Language Models (VLMs) such as CLIP in downstream tasks, most recent methods apply CLIP to UDA tasks through learning domain-specific text prompts for source and target domains separately. However, these methods fail to dynamically adjust image features based on the characteristics of their respective domains, thereby limiting their alignment with domain-specific text prompts in CLIP’s joint space, which is a key factor in improving classification performance in the target domain. To bridge this gap, we propose a Unified Text-Image Space Alignment with Cross-Modal Prompting (UTISA) framework for UDA. First, we introduce a Cross-Modal Prompt Learning (CMP) module to generate domain-specific image prompts and layer-specific image prompts for the visual branch to encode domain-specific knowledge globally and locally. Second, under the guidance of image prompts, we introduce a Domain-Aware Multi-Layer Feature Fusion (DMF) module to construct multi-layer domain features for each domain and enhance the image features with these multi-layer domain features, which enables the image features to better reflect the characteristics of their respective domains, thereby promoting their alignment with domain-specific text prompts. Moreover, we introduce a Perturbation-Driven Regularization (PDR) mechanism for the target domain to enhance the robustness and generalization of the model. The experiments demonstrate that UTISA achieves the best performance on three mainstream UDA benchmarks, including 87.9% on Office-Home, 90.9% on VisDA-2017, and 62.4% on DomainNet.
Yifan Jiao, Chenglong Cai, Bing-Kun Bao
ACM Trans. Multim. Comput. Commun. Appl.1
2025 Bool Prompt with Decomposition and Enhancement: Zero-Shot VQA Based on PVLMs
abstract
Zero-Shot Visual Question Answering (ZSVQA) aims to answer questions about images without prior training on explicit image question pairs. Most existing methods usually apply Pre-trained Visual and Language Models (PVLMs) by designing prompts to convert questions into predefined input templates, which (1) ignores the text details and associations of the question when the question is complex, hindering comprehensive understanding, and (2) does not pay attention to local information of image, resulting in overlooking some details of the image that are important to the question when the image content is particularly complex or requires detailed observation. To address these challenges, we propose the Bool Prompt with Decomposition and Enhancement (BPDE) framework for ZSVQA. Specifically, we propose the Bool Sub-Questions Generating module to extract keywords from the original question and generate captions from the image, then use these keywords and captions to guide the transformation of original questions into simpler bool sub-questions, which focus on a specific logical point or piece of information, and guided from the captions can provide the model with local visual information, thereby enhancing the model’s understanding of complex questions and attention to local visual information. Additionally, an Adaptive Sub-Questions Selecting mechanism is designed to ensure non-redundant selection and that the meanings of the sub-questions can cover the original question. Extensive experiments on VQAv2 and AOKVQA demonstrate that the proposed approach performs favorably against the state-of-the-art methods.
Liyong Xu, Yifan Jiao, Bing-Kun Bao
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Source-Guided Target Feature Reconstruction for Cross-Domain Classification and Detection
abstract
Existing cross-domain classification and detection methods usually apply a consistency constraint between the target sample and its self-augmentation for unsupervised learning without considering the essential source knowledge. In this paper, we propose a Source-guided Target Feature Reconstruction (STFR) module for cross-domain visual tasks, which applies source visual words to reconstruct the target features. Since the reconstructed target features contain the source knowledge, they can be treated as a bridge to connect the source and target domains. Therefore, using them for consistency learning can enhance the target representation and reduce the domain bias. Technically, source visual words are selected and updated according to the source feature distribution, and applied to reconstruct the given target feature via a weighted combination strategy. After that, consistency constraints are built between the reconstructed and original target features for domain alignment. Furthermore, STFR is connected with the optimal transportation algorithm theoretically, which explains the rationality of the proposed module. Extensive experiments onnine benchmarksandtwo cross-domain visual tasksprove the effectiveness of the proposed STFR module,e.g., 1)cross-domain image classification: obtaining average accuracy of 91.0%, 73.9%, and 87.4% onOffice-31,Office-Home, andVisDA-2017, respectively; 2)cross-domain object detection: obtaining mAP of 44.50% onCityscapes→Foggy Cityscapes, AP on car of 78.10% onCityscapes→KITTI, MR-2of 8.63%, 12.27%, 22.10%, and 40.58% onCOCOPersons→Caltech,CityPersons→Caltech,COCOPersons→CityPersons, andCaltech→CityPersons, respectively.
Yifan Jiao, Hantao Yao, Bing-Kun Bao, Changsheng Xu
IEEE Trans. Image Process.1
2023 TPM: Two-Stage Prediction Mechanism for Universal Adversarial Patch Defense
Huaize Dong, Yifan Jiao, Bing-Kun Bao
ICIG (5)2
2023 Rescue decision via Earthquake Disaster Knowledge Graph reasoning
Yifan Jiao, Sisi You
Multim. Syst.1
2023 Dual Instance-Consistent Network for Cross-Domain Object Detection
abstract
Cross-domain object detection aims to transfer knowledge from a labeled dataset to an unlabeled dataset. Most existing methods apply a unified embedding model to generate the tightly coupled source and target descriptions for domain alignment, leading to the destroyed feature distribution of the target domain because the embedding model is mainly controlled by the source domain. To reduce the representation bias of the target domain, we apply two independent networks to extract two types of discriminative descriptions with mutual consistency, i.e., a novel Dual Instance-Consistent Network (DICN) is proposed for cross-domain object detection. Especially, Dual Instance-Consistent Module containing the instance mutual consistency between Primary Network and Auxiliary Network is applied to align two domains, where Primary and Auxiliary Networks are used to obtain the source-specific and target-specific information, respectively. The instance mutual consistency consists of two terms: feature consistency and detection consistency, which is applied to align the instance feature and the output of detection head, respectively. With the instance mutual consistency, optimizing the Primary (Auxiliary) Network only with source (target) images by fixing the Auxiliary (Primary) Network can generate the source(target)-specific description. Extensive experiments on several benchmarks demonstrate the effectiveness of the proposed DICN, e.g., obtaining mAP of 44.10% for Cityscapes$\rightarrow$Foggy Cityscapes, AP on car of 76.50% for Cityscapes$\rightarrow$KITTI, MR$^{-2}$of 8.87%, 12.66%, 22.27%, and 42.06% for COCOPersons$\rightarrow$Caltech, CityPersons$\rightarrow$Caltech, COCOPersons$\rightarrow$CityPersons, and Caltech$\rightarrow$CityPersons, respectively.
Yifan Jiao, Hantao Yao, Changsheng Xu
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 PEN: Pose-Embedding Network for Pedestrian Detection
abstract
In the past years, pedestrian detection has achieved significant progress via improving the visual description. However, the visual description is not robust to discover the occluded pedestrian, which is the bottleneck of the existing pedestrian methods. Targeting to overcome the shortcoming of visual description, we employ the human pose information, which is complementary to the visual description, to address the occlusion and false positive failure problems in pedestrian detection. The advantage of using human pose information is that the pose estimation model can localize the local part of the pedestrian once the pedestrian is occluded. By embedding the human pose information with the visual description, we propose a novel Pose-Embedding Network for pedestrian detection, which consists of two components: a Region Proposal Network, and a Pedestrian Recognization Network. The Region Proposal Network targets to generate lots of candidate proposals and corresponding confidence scores. Once obtaining the candidate proposals, the Pedestrian Recognization Network is proposed to distinguish pedestrian proposals by taking the visual information and pose information into consideration to refine the confidence scores and eliminate the false positives. Given the proposal image, the visual information is extracted with the Visual Feature Module. The Human Pose Module, which is proposed based on the pose estimation model, is used to predict the pose information. Further, the Classification Module is employed to fuse the visual and pose information and generates a pose-embedding pedestrian description. Extensive experiments on three challenging datasets, i.e., Caltech, CityPersons, and COCOPersons, show that the proposed approach achieves a significant improvement upon the state-of-the-art methods.
Yifan Jiao, Hantao Yao, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.1
2021 SAN: Selective Alignment Network for Cross-Domain Pedestrian Detection
abstract
Cross-domain pedestrian detection, which has been attracting much attention, assumes that the training and test images are drawn from different data distributions. Existing methods focus on aligning the descriptions of whole candidate instances between source and target domains. Since there exists a giant visual difference among the candidate instances, aligning whole candidate instances between two domains cannot overcome the inter-instance difference. Compared with aligning the whole candidate instances, we consider that aligning each type of instances separately is a more reasonable manner. Therefore, we propose a novel Selective Alignment Network for cross-domain pedestrian detection, which consists of three components: a Base Detector, an Image-Level Adaptation Network, and an Instance-Level Adaptation Network. The Image-Level Adaptation Network and Instance-Level Adaptation Network can be regarded as the global-level and local-level alignments, respectively. Similar to the Faster R-CNN, the Base Detector, which is composed of a Feature module, an RPN module and a Detection module, is used to infer a robust pedestrian detector with the annotated source data. Once obtaining the image description extracted by the Feature module, the Image-Level Adaptation Network is proposed to align the image description with an adversarial domain classifier. Given the candidate proposals generated by the RPN module, the Instance-Level Adaptation Network firstly clusters the source candidate proposals into several groups according to their visual features, and thus generates the pseudo label for each candidate proposal. After generating the pseudo labels, we align the source and target domains by maximizing and minimizing the discrepancy between the prediction of two classifiers iteratively. Extensive evaluations on several benchmarks demonstrate the effectiveness of the proposed approach for cross-domain pedestrian detection.
Yifan Jiao, Hantao Yao, Changsheng Xu
IEEE Trans. Image Process.1
2019 Video Highlight Detection via Region-Based Deep Ranking Model
abstract
The video highlight detection task is to localize key elements (moments of user’s major or special interest) in a video. Most of the existing highlight detection approaches extract features from the video segment as a whole without considering the difference of local features spatially. In spatial extent, not all regions are worth watching because some of them only contain the background of the environment without human or other moving objects, especially when there is lots of clutter in the background. To deal with this issue, we propose a novel region-based model which can automatically localize the key elements in a video without any extra supervised annotations. Specifically, the proposed model produces position-sensitive score maps for local regions in the spatial dimension of the video segment, and then aggregates all position-wise scores with position-pooling operation. The regions with higher response values will be extracted as key elements. Thus more effective features of the video segment are obtained to predict the highlight score. The proposed position-sensitive scheme can be easily integrated into an end-to-end fully convolutional network which aims to update parameters via stochastic gradient descent method in the backward propagation to improve the robustness of the model. Extensive experimental results on the YouTube and SumMe datasets demonstrate that the proposed approach achieves significant improvement over state-of-the-art methods.
Yifan Jiao, Tianzhu Zhang 0001, Shucheng Huang, Bin Liu 0014, Changsheng Xu
Int. J. Pattern Recognit. Artif. Intell.1
2018 Three-Dimensional Attention-Based Deep Ranking Model for Video Highlight Detection
abstract
The video highlight detection task is to localize key elements (moments of user's major or special interest) in a video. Most of existing highlight detection approaches extract features from the video segment as a whole without considering the difference of local features both temporally and spatially. Due to the complexity of video content, this kind of mixed features will impact the final highlight prediction. In temporal extent, not all frames are worth watching because some of them only contain the background of the environment without human or other moving objects. In spatial extent, it is similar that not all regions in each frame are highlights especially when there are lots of clutters in the background. To solve the above problem, we propose a novel three-dimensional (3-D) (spatial+temporal) attention model that can automatically localize the key elements in a video without any extra supervised annotations. Specifically, the proposed attention model produces attention weights of local regions along both the spatial and temporal dimensions of the video segment. The regions of key elements in the video will be strengthened with large weights. Thus, the more effective feature of the video segment is obtained to predict the highlight score. The proposed 3-D attention scheme can be easily integrated into a conventional end-to-end deep ranking model that aims to learn a deep neural network to compute the highlight score of each video segment. Extensive experimental results on the YouTube and SumMe datasets demonstrate that the proposed approach achieves significant improvement over state-of-the-art methods. With the proposed 3-D attention model, video highlights can be accurately retrieved in spatial and temporal dimensions without human supervision in several domains, such as gymnastics, parkour, skating, skiing, surfing, and dog activities, on the public datasets.
Yifan Jiao, Zhetao Li, Shucheng Huang, Xiaoshan Yang, Bin Liu 0014, Tianzhu Zhang 0001
IEEE Trans. Multim.1
2017 Video Highlight Detection via Deep Ranking Modeling
Yifan Jiao, Xiaoshan Yang, Tianzhu Zhang 0001, Shucheng Huang, Changsheng Xu
PSIVT1