Junkai Yan

dblp:227/6500 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
12since 2021 · last 2025
0009-0009-6531-0070ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
abstract
Recent open-vocabulary detectors achieve promising performance with abundant region-level annotated data. In this work, we show that an open-vocabulary detector co-training with a large language model by generating image-level detailed captions for each image can further improve performance. To achieve the goal, we first collect a dataset, GroundingCap-1M, wherein each image is accompanied by associated grounding labels and an image-level detailed caption. With this dataset, we finetune an open-vocabulary detector with training objectives including a standard grounding loss and a caption generation loss. We take advantage of a large language model to generate both region-level short captions for each region of interest and image-level long captions for the whole image. Under the supervision of the large language model, the resulting detector, LLMDet, outperforms the baseline by a clear margin, enjoying superior open-vocabulary ability. Further, we show that the improved LLMDet can in turn build a stronger large multi-modal model, achieving mutual benefits. The code, model, and dataset are available at https://github.com/iSEE-Laboratory/LLMDet.
Shenghao Fu, Qize Yang, Qijie Mo, Junkai Yan, Xihan Wei, Jingke Meng, Xiaohua Xie, Wei-Shi Zheng 0001
CVPR4
2025 A Versatile Framework for Multi-Scene Person Re-Identification
abstract
Person Re-identification (ReID) has been extensively developed for a decade in order to learn the association of images of the same person across non-overlapping camera views. To overcome significant variations between images across camera views, mountains of variants of ReID models were developed for solving a number of challenges, such as resolution change, clothing change, occlusion, modality change, and so on. Despite the impressive performance of many ReID variants, these variants typically function distinctly and cannot be applied to other challenges. To our best knowledge, there is no versatile ReID model that can handle various ReID challenges at the same time. This work contributes to the first attempt at learning a versatile ReID model to solve such a problem. Our main idea is to form a two-stage prompt-based twin modeling framework called VersReID. Our VersReID firstly leverages the scene label to train a ReID Bank that contains abundant knowledge for handling various scenes, where several groups of scene-specific prompts are used to encode different scene-specific knowledge. In the second stage, we distill a V-Branch model with versatile prompts from the ReID Bank for adaptively solving the ReID of different scenes, eliminating the demand for scene labels during the inference stage. To facilitate training VersReID, we further introduce the multi-scene properties into self-supervised learning of ReID via a multi-scene prioris data augmentation (MPDA) strategy. Through extensive experiments, we demonstrate the success of learning an effective and versatile ReID model for handling ReID tasks under multi-scene conditions without manual assignment of scene labels in the inference stage, including general, low-resolution, clothing change, occlusion, and cross-modality scenes.
Wei-Shi Zheng 0001, Junkai Yan, Yi-Xing Peng
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 A Hierarchical Semantic Distillation Framework for Open-Vocabulary Object Detection
abstract
Open-vocabulary object detection (OVD) aims to detect objects beyond the training annotations, where detectors are usually aligned to a pre-trained vision-language model, e.g., CLIP, to inherit its generalizable recognition ability so that detectors can recognize new or novel objects. However, previous works directly align the feature space with CLIP and fail to learn the semantic knowledge effectively. In this work, we propose a hierarchical semantic distillation framework named HD-OVD to construct a comprehensive distillation process, which exploits generalizable knowledge from the CLIP model in three aspects. In the first hierarchy of HD-OVD, the detector learns fine-grainedinstance-wise semanticsfrom the CLIP image encoder by modeling relations among single objects in the visual space. Besides, we introduce text space novel-class-aware classification to help the detector assimilate the highly generalizableclass-wise semanticsfrom the CLIP text encoder, representing the second hierarchy. Lastly, abundantimage-wise semanticscontaining multi-object and their contexts are also distilled by an image-wise contrastive distillation. Benefiting from the elaborated semantic distillation in triple hierarchies, our HD-OVD inherits generalizable recognition ability from CLIP in instance, class, and image levels. Thus, we boost the novel AP on the OV-COCO dataset to 46.4% with a ResNet50 backbone, which outperforms others by a clear margin. We also conduct extensive ablation studies to analyze how each component works.
Shenghao Fu, Junkai Yan, Qize Yang, Xihan Wei, Xiaohua Xie, Wei-Shi Zheng 0001
IEEE Trans. Multim.2
2024 Bridge Past and Future: Overcoming Information Asymmetry in Incremental Object Detection
Qijie Mo, Yipeng Gao, Shenghao Fu, Junkai Yan, Ancong Wu, Wei-Shi Zheng 0001
ECCV (16)4
2024 DreamView: Injecting View-Specific Text Guidance Into Text-to-3D Generation
Junkai Yan, Yipeng Gao, Qize Yang, Xihan Wei, Xuansong Xie, Ancong Wu, Wei-Shi Zheng 0001
ECCV (25)1
2024 PTMA: Pre-trained Model Adaptation for Transfer Learning
Xiao Li 0074, Junkai Yan, Jianjian Jiang, Wei-Shi Zheng 0001
KSEM (1)2
2024 Loc4Plan: Locating Before Planning for Outdoor Vision and Language Navigation
abstract
Vision and Language Navigation (VLN) is a challenging task that requires agents to understand instructions and navigate to the destination in a visual environment. One of the key challenges in outdoor VLN is keeping track of which part of the instruction was completed. To alleviate this problem, previous works mainly focus on grounding the natural language to the visual input, but neglecting the crucial role of the agent's spatial position information in the grounding process. In this work, we first explore the substantial effect of spatial position locating on the grounding of outdoor VLN, drawing inspiration from human navigation. In real-world navigation scenarios, before planning a path to the destination, humans typically need to figure out their current location. This observation underscores the pivotal role of spatial localization in the navigation process. In this work, we introduce a novel framework, Locating before Planning (Loc4Plan), designed to incorporate spatial perception for action planning in outdoor VLN tasks. The main idea behind Loc4Plan is to perform the spatial localization before planning a decision action based on corresponding guidance, which comprises a block-aware spatial locating (BAL) module and a spatial-aware action planning (SAP) module. Specifically, to help the agent perceive its spatial location in the environment, we propose to learn a position predictor that measures how far the agent is from the next intersection for reflecting its position, which is achieved by the BAL module. After this locating process, we propose the PSA module to associate visual observations After the locating process, we propose the SAP module to incorporate spatial information to ground the corresponding guidance and enhance the precision of action planning. Extensive experiments on the Touchdown and map2seq datasets show that the proposed Loc4Plan outperforms the SOTA methods.
Huilin Tian, Jingke Meng, Wei-Shi Zheng 0001, Yuan-Ming Li, Junkai Yan, Yunong Zhang
ACM Multimedia5
2024 Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models
abstract
Recent vision foundation models can extract universal representations and show impressive abilities in various tasks. However, their application on object detection is largely overlooked, especially without fine-tuning them. In this work, we show that frozen foundation models can be a versatile feature enhancer, even though they are not pre-trained for object detection. Specifically, we explore directly transferring the high-level image understanding of foundation models to detectors in the following two ways. First, the class token in foundation models provides an in-depth understanding of the complex scene, which facilitates decoding object queries in the detector's decoder by providing a compact context. Additionally, the patch tokens in foundation models can enrich the features in the detector's encoder by providing semantic details. Utilizing frozen foundation models as plug-and-play modules rather than the commonly used backbone can significantly enhance the detector's performance while preventing the problems caused by the architecture discrepancy between the detector's backbone and the foundation model. With such a novel paradigm, we boost the SOTA query-based detector DINO from 49.0% AP to 51.9% AP (+2.9% AP) and further to 53.8% AP (+4.8% AP) by integrating one or two foundation models respectively, on the COCO validation set after training for 12 epochs with R50 as the detector's backbone. Code will be available.
Shenghao Fu, Junkai Yan, Qize Yang, Xihan Wei, Xiaohua Xie, Wei-Shi Zheng 0001
NeurIPS2
2023 AsyFOD: An Asymmetric Adaptation Paradigm for Few-Shot Domain Adaptive Object Detection
abstract
In this work, we study few-shot domain adaptive object detection (FSDAOD), where only a few target labeled images are available for training in addition to sufficient source labeled images. Critically, in FSDAOD, the data scarcity in the target domain leads to an extreme data imbalance between the source and target domains, which potentially causes over-adaptation in traditional feature alignment. To address the data imbalance problem, we propose an asymmetric adaptation paradigm, namely AsyFOD, which leverages the source and target instances from different perspectives. Specifically, by using target distribution estimation, the AsyFOD first identifies the target-similar source instances, which serves to augment the limited target instances. Then, we conduct asynchronous alignment between target-dissimilar source instances and augmented target instances, which is simple yet effective for alleviating the over-adaptation. Extensive experiments demonstrate that the proposed AsyFOD outperforms all state-of-the-art methods on four FSDAOD benchmarks with various environmental variances, e.g., 3.1% mAP improvement on Cityscapes-to-FoggyCityscapes and 2.9% mAP increase on Sim10k-to-Cityscapes. The code is available at https://github.com/Hlings/AsyFPD.
Yipeng Gao, Kun-Yu Lin, Junkai Yan, Yaowei Wang 0001, Wei-Shi Zheng 0001
CVPR3
2023 ASAG: Building Strong One-Decoder-Layer Sparse Detectors via Adaptive Sparse Anchor Generation
abstract
Recent sparse detectors with multiple, e.g. six, decoder layers achieve promising performance but much inference time due to complex heads. Previous works have explored using dense priors as initialization and built one-decoder-layer detectors. Although they gain remarkable acceleration, their performance still lags behind their six-decoder-layer counterparts by a large margin. In this work, we aim to bridge this performance gap while retaining fast speed. We find that the architecture discrepancy between dense and sparse detectors leads to feature conflict, hampering the performance of one-decoder-layer detectors. Thus we propose Adaptive Sparse Anchor Generator (ASAG) which predicts dynamic anchors on patches rather than grids in a sparse way so that it alleviates the feature conflict problem. For each image, ASAG dynamically selects which feature maps and which locations to predict, forming a fully adaptive way to generate image-specific anchors. Further, a simple and effective Query Weighting method eases the training instability from adaptiveness. Extensive experiments show that our method outperforms dense-initialized ones and achieves a better speed-accuracy trade-off. The code is available at https://github.com/iSEE-Laboratory/ASAG.
Shenghao Fu, Junkai Yan, Yipeng Gao, Xiaohua Xie, Wei-Shi Zheng 0001
ICCV2
2023 Self-supervised Cross-stage Regional Contrastive Learning for Object Detection
abstract
Cross-stage object similarity is a vital property of generic supervised object detectors, which maintains similar feature responses to the same object across feature maps of different intermediate stages of the backbone network. Since an object can be predicted by multiple stages, this similarity is beneficial for accurate object classification and localization. Inspired by this property, we introduce Cross-stage regional Contrastive Learning (CrossCL) to learn the cross-stage object similarity during the model pre-training. Since labels are unavailable in self-supervised learning, we treat the regions sharing the same position in different stages as the same object and constrain them to have similar feature responses across stages to achieve cross-stage object similarity. The learned feature representations of CrossCL share a similar property with supervised detectors, thus showing strong transfer capability to object detection tasks. Besides, we also provide in-depth discussions, ablation studies, and visualizations to understand better how CrossCL works. Code is available at https://github.com/yanjk3/CrossCL.
Junkai Yan, Lingxiao Yang, Yipeng Gao, Wei-Shi Zheng 0001
ICME1
2022 Space-correlated Contrastive Representation Learning with Multiple Instances
abstract
Self-supervised contrastive learning methods have shown promising transferability in pretraining by maximizing the mutual information between two cropped regions as views from the same image. In order to effectively extract mutual information between views, the cropped regions need to be the same instance as prior hypothesis. However, the data collected in general scenes usually have multiple instances, so the two cropped regions probably contain different instances which will mislead the contrastive learning process. In this paper, we make the first attempt to exploit the spatial position relationships of the two cropped regions in self-supervised contrastive learning with images that include multiple instances. Then, we propose an effective method called Space-correlated Contrastive Learning (SpaceCL). Specifically, given two randomly cropped regions as contrastive pairs from the same image, we implement self-supervised contrastive learning by optimizing a space correspondence contrastive similarity loss. As a result, our method achieves state-of-the-art performance and remarkably outperforms other counterparts when pretrained on the COCO dataset of which images contain multiple instances. Experiments show our method outperforms ReSim with 2.6%AP on PASCAL VOC object detection, 0.8%APbband 0.6%APmkon COCO object detection and instance segmentation, 1.3%APmkon Cityscapes instance segmentation.
Danming Song, Yipeng Gao, Junkai Yan, Wei Sun 0007, Wei-Shi Zheng 0001
ICPR3