Jinpeng Mi

dblp:195/9013 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-0506-9707ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 12 since 2021Systems, architecture and hardware · 6 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Zero-Shot Visual Grounding via Cascade Semantic Prototype Learning
abstract
Compared to standard visual grounding (VG), zero-shot visual grounding (ZSVG) requires robust localization of novel image–query pairs unseen during training, creating a clear gap from supervised settings. Existing two-stage approaches rely too heavily on external knowledge from pre-trained vision-language models and underexplore in-context semantics, while current one-stage models fail to leverage region-level cues to enhance discriminative representation learning. In this paper, inspired by the information processing procedure of human beings, we propose a novel one-stage network, dubbed Cascade Semantic Prototype Learning (CaSePro), to tackle these limitations. Specifically, CaSePro builds two types of prototypes that capture prototypical multimodal representations for global contexts and local semantics, respectively. With global context prototypes, CaSePro explores context-related cues to enhance the coarsely fused global multimodal features, enabling the model to learn more comprehensive representations. With local semantic prototypes, CaSePro searches for region-level multimodal correspondences to augment the prototypical global patterns, thereby strengthening the cross-modal relationship between the query and visual representations. Moreover, to enhance the generalization ability of the proposed approach on sample-imbalanced benchmark datasets, CaSePro further incorporates an Adaptive Focal Loss that dynamically reweights hard and easy samples by accounting for the model’s grounding uncertainty. Extensive experiments on ten benchmarks demonstrate that CaSePro outperforms all existing one-stage approaches and achieves performance comparable to state-of-the-art two-stage methods.
Jinpeng Mi, Jinbo Yang, Xian Wei, Jianwei Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Relaxed Rotational Equivariance via G-Biases in Vision
abstract
Group Equivariant Convolution (GConv) can capture rotational equivariance from original data. It assumes uniform and strict rotational equivariance across all features as the transformations under the specific group. However, the presentation or distribution of real-world data rarely conforms to strict rotational equivariance, commonly referred to as Rotational Symmetry-Breaking (RSB) in the system or dataset, making GConv unable to adapt effectively to this phenomenon. Motivated by this, we propose a simple but highly effective method to address this problem, which utilizes a set of learnable biases called G-Biases under the group order to break strict group constraints and then achieve a Relaxed Rotational Equivariant Convolution (RREConv). To validate the efficiency of RREConv, we conduct extensive ablation experiments on the discrete rotational group Cn. Experiments demonstrate that the proposed RREConv-based methods achieve excellent performance compared to existing GConv-based methods in both classification and 2D object detection tasks on the natural image datasets.
Licheng Sun, Jian Yang 0034, Shing-Ho J. Lin, Jinpeng Mi, Xian Wei
AAAI8
2025 3D Dense Captioning via Prototypical Momentum Distillation
abstract
3D dense captioning aims to describe the crucial regions in 3D visual scenes in the form of natural language. Recent prevailing approaches achieve promising results by leveraging complicated structures incorporated with large-scale models, which necessitate abundant parameters and pose challenges regarding its practical applications. Besides, with limited training data, 3D dense captioners are often susceptible to overfitting, directly degrading caption generation performance. Drawing inspiration from the recent advancements in knowledge distillation, we propose a novel approach termed Prototypical Momentum Distillation (PMD) to prompt the model to generate more detailed captions. PMD incorporates Momentum Distillation (MD) with an Uncertainty-aware Prototype-anchored Clustering (UPC) strategy to transfer knowledge by considering the uncertainty of the teacher knowledge. Specifically, we employ the original captioner as the student model and maintain an Exponential Moving Average (EMA) copy of the captioner as the teacher model to impart knowledge as the auxiliary supervision of the student. To abate the misleading caused by uncertain knowledge, we present an Uncertainty-aware Prototype-anchored Clustering (UPC) strategy to cluster the distilled knowledge according to its confidence. We then transfer the rearranged knowledge from the teacher to guide the training route of the student. We conduct extensive experiments and ablation studies on two widely used benchmark datasets, ScanRefer and Nr3D. Experimental results demonstrate that PMD outperforms all state-of-the-art approaches on the benchmarks with MLE training, highlighting its effectiveness.
Jinpeng Mi, Shaofei Jin, Xian Wei, Jianwei Zhang 0001
ICRA1
2025 KiteRunner: Language-Driven Cooperative Local-Global Navigation Policy with UAV Mapping in Outdoor Environments
abstract
Autonomous navigation in open-world outdoor environments faces challenges in integrating dynamic conditions, long-distance spatial reasoning, and semantic understanding. Traditional methods struggle to balance local planning, global planning, and semantic task execution, while existing large language models (LLMs) enhance semantic comprehension but lack spatial reasoning capabilities. Although diffusion models excel in local optimization, they fall short in large-scale long-distance navigation. To address these gaps, this paper proposes KiteRunner, a language-driven cooperative local-global navigation strategy that combines UAV orthophoto-based global planning with diffusion model-driven local path generation for long-distance navigation in open-world scenarios. Our method innovatively leverages real-time UAV orthophotography to construct a global probability map, providing traversability guidance for the local planner, while integrating large models like CLIP and GPT to interpret natural language instructions. Experiments demonstrate that KiteRunner achieves 5.6% and 12.8% improvements in path efficiency over state-of-the-art methods in structured and unstructured environments, respectively, with significant reductions in human interventions and execution time.
Shibo Huang, Chenfan Shi, Jian Yang 0034, Jinpeng Mi, Ke Li 0005, Miao Ding, Peidong Liang, Xiong You, Xian Wei
IROS5
2025 NeuroLoc: Encoding Navigation Cells for 6-DOF Camera Localization
abstract
Recently, camera localization has been widely adopted in autonomous robotic navigation due to its efficiency and convenience. However, autonomous navigation in unknown environments often suffers from scene ambiguity, environmental disturbances, and dynamic object transformation in camera localization. To address this problem, inspired by the brain cognitive navigation mechanism (such as grid cells, place cells, and head direction cells), we propose a novel neurobiological camera location method, namely NeuroLoc. Firstly, we designed a Hebbian learning module driven by place cells to save and replay historical information, aiming to restore the details of historical representations and solve the issue of scene fuzziness. Secondly, we utilized the head direction cell-inspired internal direction learning as multi-head attention embedding to help restore the true orientation in similar scenes. Finally, we added a 3D grid center prediction in the pose regression module to reduce the final wrong prediction. We evaluate the proposed NeuroLoc on commonly used benchmark indoor and outdoor datasets. The experimental results show that our NeuroLoc can enhance the robustness in complex environments and improve the performance of pose regression by using only a single image.
Jian Yang 0034, Fenli Jia, Muyu Wang, Jinpeng Mi, Jilin Hu, Peidong Liang, Ke Li 0005, Xiong You, Xian Wei
IROS6
2025 Learning economically for Chinese word segmentation: tuning pretrained model via active learning and N-gram preference
Zhiyuan Ma 0001, Jiwei Qin, Song Tang 0001, Jinpeng Mi
Neural Comput. Appl.4
2025 SAM-Net: Semantic-assisted multimodal network for action recognition in RGB-D videos
Jinpeng Mi, Mao Ye 0001, Qingdu Li, Jianwei Zhang 0001
Pattern Recognit.3
2025 FuseFormer: A Manifold Metric Fusing Attention for Pedestrian Trajectory Prediction
abstract
Accurate pedestrian trajectory prediction is critical for ensuring the safety of autonomous vehicles and advancing higher levels of driving automation. However, the complex interpersonal interactions and highly dynamic trajectory patterns in real-world scenarios pose significant challenges to achieving precise predictions. Recently, Transformers have shown remarkable success in pedestrian trajectory prediction, primarily due to their effective modeling of temporal and spatial dependencies via Multi-Head Self-Attention (MHA) mechanisms. Despite these advancements, existing self-attention methods often rely on Euclidean distance-based metrics and dot-product operations, which are inadequate for capturing interaction-induced trajectory curvatures. To address this limitation, we propose a novel hybrid Transformer architecture, FuseFormer, that incorporates Geodesic Self-Attention (GSA) mechanisms. GSA utilizes geodesic distances to characterize interaction features effectively, complementing MHA, which excels in capturing local features and maintaining temporal correlations. FuseFormer employs a gating network to adaptively combine GSA and MHA embeddings, leveraging their complementary strengths. Additionally, FuseFormer integrates a Transformer-based Neural Ordinary Differential Equation (ODE) decoder to model trajectory temporal dynamics. This design enables the generation of future trajectories that align closely with motion trends while adapting the network depth to input sequence lengths. Experimental results demonstrate that FuseFormer achieves state-of-the-art performance across widely used pedestrian trajectory prediction datasets, including ETH/UCY, SDD, and NBA. These results underscore the model’s effectiveness and generalization capability in capturing complex interaction patterns and handling diverse scenarios.
Kohsin Ko, Jian Yang 0034, Ke Li 0005, Xiong You, Jinpeng Mi, Mingsong Chen 0001, Xian Wei
IEEE Trans. Intell. Transp. Syst.7
2024 Temporal cues enhanced multimodal learning for action recognition in RGB-D videos
Zhiyuan Ma 0001, Jinpeng Mi, Yan Gan, Mao Ye 0001, Jianwei Zhang 0001
Neurocomputing5
2024 Zero-shot visual grounding via coarse-to-fine representation learning
Jinpeng Mi, Shaofei Jin, Zhiqian Chen, Xian Wei, Jianwei Zhang 0001
Neurocomputing1
2024 Adaptive knowledge distillation and integration for weakly supervised referring expression comprehension
Jinpeng Mi, Stefan Wermter, Jianwei Zhang 0001
Knowl. Based Syst.1
2023 Weakly Supervised Referring Expression Grounding via Target-Guided Knowledge Distillation
abstract
Weakly supervised referring expression grounding aims to train a model without the manual labels between image regions and referring expressions during the training phase. Current predominant models often adopt deep structures to reconstruct the region-expression correspondence. A crucial deficiency of the existing approaches lies in that these models neglect to exploit potential valuable information to further improve their grounding performance. To address this issue, we leverage knowledge distillation as a unique scheme to excavate and transfer helpful information for acquiring a better model. Specifically, we propose a target-guided knowledge distillation framework that accounts for region-expression pairs reconstruction and matching. We reactivate the target-related prediction information learned by a pre-trained teacher model and transfer the target-related prediction knowledge from the teacher to guide the training process and boost the performance of the student model. We conduct extensive experiments on three benchmark datasets, i.e., RefCOCO, RefCOCO+, and RefCOCOg. Without bells and whistles, our approach achieves state-of-the-art results on several splits of benchmark datasets. The implementation codes and trained models are available at: https://github.com/dami23/WREG_KD.
Jinpeng Mi, Song Tang 0001, Zhiyuan Ma 0001, Qingdu Li, Jianwei Zhang 0001
ICRA1
2023 Weakly Supervised Referring Expression Grounding via Dynamic Self-Knowledge Distillation
abstract
Weakly supervised referring expression grounding (WREG) is an attractive and challenging task for grounding target regions in images by understanding given referring expressions. WREG learns to ground target objects without the manual annotations between image regions and referring expressions during the model training phase. Different from the predominant grounding pattern of existing models, which locates target objects by reconstructing the region-expression correspondence, we investigate WREG from a novel perspective and enrich the prevailing pattern with self-knowledge distillation. Specifically, we propose a target-guided self-knowledge distillation approach that adopts the target prediction knowledge learned from the previous training iterations as the teacher to guide the subsequent training procedure. In order to avoid the misleading caused by the teacher knowledge with low prediction confidence, we present an uncertaintyaware knowledge refinement strategy to adaptively rectify the teacher knowledge by learning dynamic threshold values based on the model prediction uncertainty. To validate the proposed approach, we implement extensive experiments on three benchmark datasets, i.e., Ref Coco, RefCOCO+, and RefCOCOg. Our approach achieves new state-of-the-art results on several splits of the benchmark datasets, showcasing the advantage of the proposed framework for WREG. The implementation codes and trained models are available at: https://github.com/dami23IWREG.sar_KD.
Jinpeng Mi, Zhiqian Chen, Jianwei Zhang 0001
IROS1
2023 Cross-domain video action recognition via adaptive gradual learning
Zhenwei Bao, Jinpeng Mi, Yan Gan, Mao Ye 0001, Jianwei Zhang 0001
Neurocomputing3
2019 Visual Domain Adaptation Exploiting Confidence-Samples
abstract
Domain adaptation methods are used to address a problem, in which train scenario (source domain) and test scenario (target domain) are different. The existing methods mainly perform adaptation via reducing domain discrepancy from the view of a probability distribution. However, the idea of probability distribution matching always leads to a complex optimization process. Thereby these methods are difficult to apply in some scenario like online application or fast perception in dynamic environments. In this paper, we propose a new and simple domain adaptation method that utilizes confidence- samples to facilitate the classifier training on the target domain. Here, the confidence-samples are a subset of the target samples, and they have very credibly predicted labels. In order to detect the samples, a Category Similarity Collaborative Representation (CSCR) is first developed, by which the raw labels of all target samples are predicted using the smallest projection error according to the law of category. After this, the confidence score of the raw predicted labels is evaluated by the energy context information of CSCR. Finally, the target samples with a high confidence score are selected. Because of the linearity of CSCR, our method avoids complex optimization for matching the probability distribution. Empirical studies on a standard dataset demonstrate the advantages of our method.
Song Tang 0001, Yunfeng Ji, Jianzhi Lyu, Jinpeng Mi, Qingdu Li, Jianwei Zhang 0001
IROS4