Yunda Sun

dblp:19/6167 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Decision-Invariant Sim-to-Real Vision-and-Language Navigation with Pseudo-Panoramic Observations
abstract
Following natural language instructions to complete navigation tasks is a crucial capability for real-world embodied robots. In vision-and-language navigation (VLN), agents typically assume access to complete, on-the-fly environmental observations and rely on them for decision making. However, most sim-to-real VLN approaches approximate the privileged complete panoramic sensing in simulation with monocular sensors, leading to significant semantic loss and incomplete perception. In this work, we present DIP2, a sim-to-real VLN framework that reduces the discrepancy between assumed and realizable observations by introducing a unified sensing and mapping representation shared across simulated and real-world domains. Specifically, dense pseudo-panoramic observations are synthesized from a self-assembled surround-view camera system to recover rich semantics comparable to simulation, while a structure-only metric map unifies simulated RGB-D inputs with real-world LiDAR observations. Furthermore, to promote sim-to-real decision invariance, DIP2 employs a learning-free local navigable waypoint prediction strategy via radial expansion, applied consistently in both domains. The predicted waypoints and dense semantic observations are integrated into a global topological map, allowing existing agents trained in simulation to be directly deployed in the real world. Extensive experiments in simulated and real environments demonstrate that DIP2 substantially improves sim-to-real navigation performance. Source code will be published at https://github.com/zheng19845/DIP2.
Yuanyu Zheng, Xumin Shen, Yunda Sun, Ying Shen 0005, Lin Zhang 0014
ICMR3
2026 Taming Generative Synthetic Data for X-Ray Prohibited Item Detection
abstract
Training prohibited item detection models requires a large amount of X-ray security images, but collecting and annotating these images is time-consuming and laborious. To address data insufficiency, X-ray security image synthesis methods composite images to scale up datasets. However, previous methods primarily follow a two-stage pipeline, where they implement labor-intensive foreground extraction in the first stage and then composite images in the second stage. Such a pipeline introduces inevitable extra labor cost and is not efficient. In this paper, we propose a one-stage X-ray security image synthesis pipeline (Xsyn) based on text-to-image generation, which incorporates two effective strategies to improve the usability of synthetic images. The Cross-Attention Refinement (CAR) strategy leverages the cross-attention map from the diffusion model to refine the bounding box annotation. The Background Occlusion Modeling (BOM) strategy explicitly models background occlusion in the latent space to enhance imaging complexity. To the best of our knowledge, compared with previous methods, Xsyn is the first to achieve high-quality X-ray security image synthesis without extra labor cost. Experiments demonstrate that our method outperforms all previous methods with 1.2% mAP improvement, and the synthetic images generated by our method are beneficial to improve prohibited item detection performance across various X-ray security datasets and detectors. Code is available at https://github.com/pILLOW-1/Xsyn/.
Jialong Sun, Hongguang Zhu, Weizhe Liu, Yunda Sun, Renshuai Tao, Yunchao Wei
IEEE Trans. Inf. Forensics Secur.4
2026 CaneSpeaker: An LLM-Assisted Speaker for Generating Human-Like Navigation Instructions
abstract
Navigation instruction generation aims to address data scarcity in Vision-and-Language Navigation (VLN) by generating navigation instructions for unannotated routes from data sources like simulators or online data. However, existing methods usually suffer from high reliance on panoramic views, poor cross-task generalization ability, and limited availability of training data. To address these challenges, we propose a novel speaker, CaneSpeaker, to generate human-like instructions from front-facing images for a variety of VLN tasks. First, to mitigate the limited amount of speaker training data, we propose an Large Language Model (LLM)-based instruction augmentation method, LLM-IA, that utilizes an off-the-shelf LLM to create augmented instructions for training by distilling and reformulating existing instructions. This method allows us to collect an instruction-augmented dataset with human-level accuracy for speaker training, namely Rx2R. Second, to eliminate the dependency on panoramic views, we propose a novel Vision-Language Model (VLM)-based speaker architecture, VL-Sp. By leveraging the advanced reasoning capabilities of a pre-trained VLM, CaneSpeaker can effectively generate high-quality instructions directly from front-facing images without relying on panoramic views. Also, the prompt-based characteristic of the VLM allows us to devise a unified input representation to enable the processing of multiple VLN tasks, thus further addressing the problem of data scarcity by combining multiple datasets from different VLN tasks. Finally, we utilize CaneSpeaker to synthesize a large-scale augmented dataset, CANE, from unannotated routes in the Matterport3D Simulator. Comprehensive experiments demonstrate that CaneSpeaker generates precise instructions with diverse expressions across various VLN tasks, and the VLN agent trained on our datasets obviously outperforms its counterparts. The source codes and datasets are available at https://github.com/zheng19845/CaneSpeaker .
Yuanyu Zheng, Lin Zhang 0014, Yunda Sun, Ying Shen 0005, Shengjie Zhao 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2025 MaGo-I2P: Image-to-Point Cloud Registration with Mamba and Geometry Recovery
abstract
Estimating the relative poses between images and point clouds is a fundamental problem in multi-sensor fusion, with extensive applications in tasks such as robot localization and navigation. However, existing methods fall short in registration accuracy and efficiency due to the modality gaps and resource-consuming backbones. To address these issues, we propose the first Mamba-based I2P registration framework called MaGo-I2P. On the one hand, MaGo-I2P recovers the geometric structure of images through depth estimation, thereby constructing an implicit 3D representation of the image scene to alleviate the modality gap between images and point clouds, facilitating cross-modal feature extraction. On the other hand, unlike transformer-based backbones applied in existing methods, a Mamba-based backbone with linear time complexity is utilized in our MaGo-I2P. Such a backbone allows our method to possess both context-aware capability and fast inference speed. In addition, by adopting a coarse-to-fine matching strategy, MaGo-I2P eliminates outlier matches by progressively narrowing the matching region, establishing more accurate 2D-3D correspondences. Experiments on KITTI Odometry and Oxford Robotcar datasets suggest that our method achieves state-of-the-art registration accuracy while maintaining high-efficiency. Meanwhile, we also demonstrate the application potential of MaGo-I2P in LiDAR-camera calibration through qualitative experiments. The source code will be released at https://cslinzhang.github.io/MaGo-I2P.
Yunda Sun, Lin Zhang 0014
ICMR1
2024 HILP: hardware-in-loop pruning of convolutional neural networks towards inference acceleration
Dong Li 0040, Qianqian Ye, Xiaoyue Guo, Yunda Sun, Li Zhang 0050
Neural Comput. Appl.4
2024 I2P Registration by Learning the Underlying Alignment Feature Space from Pixel-to-Point Similarities
abstract
Estimating the relative pose between a camera and a LiDAR holds paramount importance in facilitating complex task execution within multi-agent systems. Nonetheless, current methodologies encounter two primary limitations. First, amid the cross-modal feature extraction, they typically employ separate modal branches to extract cross-modal features from images and point clouds. This approach results in the feature spaces of images and point clouds being misaligned, thereby reducing the robustness of establishing correspondences. Second, due to the scale differences between images and point clouds, one-to-many pixel-point correspondences are inevitably encountered, which will mislead the pose optimization. To address these challenges, we propose a framework named I mage-to- P oint cloud registration by learning the underlying alignment feature space from P ixel-to- P oint SIM imilarities (I2P \({}_{\mathbf{ppsim}}\) ) . Central to \(\text{I2P}_{\text{ppsim}}\) is a Shared Feature Alignment Module (SFAM). It is designed under on a coarse-to-fine architecture and uses a weight-sharing network to construct an alignment feature space. Benefiting from SFAM, \(\text{I2P}_{\text{ppsim}}\) can effectively identify the co-view regions between images and point clouds and establish high-reliability 2D-3D correspondences. Moreover, to mitigate the one-to-many correspondence issue, we introduce a similarity maximization strategy termed point-max. This strategy effectively filters out outliers, thereby establishing accurate 2D-3D correspondences. To evaluate the efficacy of our framework, we conduct extensive experiments on KITTI Odometry and Oxford Robotcar. The results corroborate the effectiveness of our framework in improving image-to-point cloud registration. To make our results reproducible, the source codes have been released at https://cslinzhang.github.io/I2P
Yunda Sun, Lin Zhang 0014, Zhong Wang 0009, Yang Chen 0037, Shengjie Zhao 0001, Yicong Zhou
ACM Trans. Multim. Comput. Commun. Appl.1
2022 QuadNet: Quadruplet loss for multi-view learning in baggage re-identification
Hao Yang 0010, Xiuxiu Chu, Li Zhang 0050, Yunda Sun, Dong Li 0040, Stephen J. Maybank
Pattern Recognit.4
2022 Feedback Graph Convolutional Network for Skeleton-Based Action Recognition
abstract
Skeleton-based action recognition has attracted considerable attention since the skeleton data is more robust to the dynamic circumstances and complicated backgrounds than other modalities. Recently, many researchers have used the Graph Convolutional Network (GCN) to model spatial-temporal features of skeleton sequences by an end-to-end optimization. However, conventional GCNs are feedforward networks for which it is impossible for the shallower layers to access semantic information in the high-level layers. In this paper, we propose a novel network, named Feedback Graph Convolutional Network (FGCN). This is the first work that introduces a feedback mechanism into GCNs for action recognition. Compared with conventional GCNs, FGCN has the following advantages: (1) A multi-stage temporal sampling strategy is designed to extract spatial-temporal features for action recognition in a coarse to fine process; (2) A Feedback Graph Convolutional Block (FGCB) is proposed to introduce dense feedback connections into the GCNs. It transmits the high-level semantic features to the shallower layers and conveys temporal information stage by stage to model video level spatial-temporal features for action recognition; (3) The FGCN model provides predictions on-the-fly. In the early stages, its predictions are relatively coarse. These coarse predictions are treated as priors to guide the feature learning in later stages, to obtain more accurate predictions. Extensive experiments on three datasets, NTU-RGB+D, NTU-RGB+D120 and Northwestern-UCLA, demonstrate that the proposed FGCN is effective for action recognition. It achieves the state-of-the-art performance on all three datasets.
Hao Yang 0010, Dan Yan, Li Zhang 0050, Yunda Sun, Dong Li 0040, Stephen J. Maybank
IEEE Trans. Image Process.4
2020 Towards Optimal Filter Pruning with Balanced Performance and Pruning Speed
Yunda Sun
ACCV (4)4
2020 STA-CNN: Convolutional Spatial-Temporal Attention Learning for Action Recognition
abstract
Convolutional Neural Networks have achieved excellent successes for object recognition in still images. However, the improvement of Convolutional Neural Networks over the traditional methods for recognizing actions in videos is not so significant, because the raw videos usually have much more redundant or irrelevant information than still images. In this paper, we propose a Spatial-Temporal Attentive Convolutional Neural Network (STA-CNN) which selects the discriminative temporal segments and focuses on the informative spatial regions automatically. The STA-CNN model incorporates a Temporal Attention Mechanism and a Spatial Attention Mechanism into a unified convolutional network to recognize actions in videos. The novel Temporal Attention Mechanism automatically mines the discriminative temporal segments from long and noisy videos. The Spatial Attention Mechanism firstly exploits the instantaneous motion information in optical flow features to locate the motion salient regions and it is then trained by an auxiliary classification loss with a Global Average Pooling layer to focus on the discriminative non-motion regions in the video frame. The STA-CNN model achieves the state-of-the-art performance on two of the most challenging datasets, UCF-101 (95.8%) and HMDB-51 (71.5%).
Hao Yang 0010, Chunfeng Yuan, Li Zhang 0050, Yunda Sun, Weiming Hu 0004, Stephen J. Maybank
IEEE Trans. Image Process.4
2019 MVB: A Large-Scale Dataset for Baggage Re-Identification and Merged Siamese Networks
Zhulin Zhang, Dong Li 0040, Yunda Sun, Li Zhang 0050
PRCV (3)4
2006 Regression-Based Human Motion Capture From Voxel Data
abstract
A regression based method is proposed to recover human body pose from 3D voxel data. In order to do this we need to convert the voxel data into a feature vector. This is done using a Bayesian approach based on Mixture of Probabilistic PCA that transforms a collection of 3D shape context descriptors, extracted from the voxels, to a compact feature vector. For the regression, the newly-proposed Multi-Variate Relevance Vector Machine is explored to learn a single mapping from this feature vector to a low-dimensional representation of full body pose. We demonstrate the effectiveness and robustness of our method with experiments on both synthetic data and real sequences. 1
Yunda Sun, Matthieu Bray, Arasanathan Thayananthan, B. Yuan, Philip Torr 0001
BMVC1