Shijie Sun 0001

dblp:42/4605-1 · also Shi-Jie Sun 0001, ShiJie Sun 0001 · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
19since 2021 · last 2026
0000-0003-4043-8448ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From perception to cognition: Unifying multi-object 3D visual grounding and dense captioning in monocular images
Keyu Guo, Yongle Huang, Hongkai Wei, Shijie Sun 0001, Mingtao Feng, Huansheng Song
Expert Syst. Appl.4
2026 ArgusNet: Understanding 3D scenes more like humans
Keyu Guo, Hongkai Wei, Yongle Huang, Shijie Sun 0001, Mingtao Feng, Huansheng Song, Jianxin Li 0001
Neurocomputing5
2026 MedHyCLIP: Hyperbolic CLIP adaptation for universal medical anomaly detection
Keyu Guo, Hongkai Wei, Yongle Huang, Shijie Sun 0001, Yueming Shi, Huansheng Song, Nicola Strisciuglio
Pattern Recognit.6
2026 Learning coherent matrixized representation in latent space for volumetric 4D generation
Qitong Yang, Mingtao Feng, Shijie Sun 0001, Weisheng Dong, Yaonan Wang 0001, Mian M. Ajmal
Pattern Recognit.4
2026 Monocular Multi-Object 3D Visual Language Tracking
abstract
Visual Language Tracking (VLT) enables machines to perform tracking in real world through human-like language descriptions. However, existing VLT methods are limited to 2D spatial tracking or single-object 3D tracking and do not support multi-object 3D tracking within monocular video. This limitation arises because advancements in 3D multi-object tracking have predominantly relied on sensor-based data (e.g., point clouds, depth sensors) that lacks corresponding language descriptions. Moreover, natural language descriptions in existing VLT literature often suffer from redundancy, impeding the efficient and precise localization of multiple objects. We present the first technique to extend VLT to multi-object 3D tracking using monocular video. We introduce a comprehensive framework that includes (i) a Monocular Multi-object 3D Visual Language Tracking (MoMo-3DVLT) task, (ii) a large-scale dataset, MoMo-3DRoVLT, tailored for this task, and (iii) a custom neural model. Our dataset, generated with the aid of Large Language Models (LLMs) and manual verification, contains 8,216 video sequences annotated with both 2D and 3D bounding boxes, with each sequence accompanied by three freely generated, human-level textual descriptions. We propose MoMo-3DVLTracker, the first neural model specifically designed for MoMo-3DVLT. This model integrates a multimodal feature extractor, a visual language encoder-decoder, and modules for detection and tracking, setting a strong baseline for MoMo-3DVLT. Beyond existing paradigms, it introduces a task-specific structural coupling that integrates a differentiable linked-memory mechanism with depth-guided and language-conditioned reasoning for robust monocular 3D multi-object tracking. Experimental results demonstrate that our approach outperforms existing methods on the MoMo-3DRoVLT dataset. Our dataset and code are available at https://github.com/hongkai-wei/MoMo-3DVLT.
Hongkai Wei, Haixiang Hu, Shijie Sun 0001, Mingtao Feng, Keyu Guo, Yongle Huang, Naveed Akhtar
IEEE Trans. Image Process.4
2025 Beyond Human Perception: Understanding Multi-Object World from Monocular View
abstract
Language and binocular vision play a crucial role in human understanding of the world. Advancements in artificial intelligence have also made it possible for machines to develop 3D perception capabilities essential for high-level scene understanding. However, only monocular cameras are often available in practice due to cost and space constraints. Enabling machines to achieve accurate 3D understanding from a monocular view is practical but presents significant challenges. We introduce MonoMulti-3DVG, a novel task aimed at achieving multi-object 3D Visual Grounding (3DVG) based on monocular RGB images, allowing machines to better understand and interact with the 3D world. To this end, we construct a large-scale benchmark dataset, MonoMulti3D-ROPE, and propose a model, CyclopsNet that integrates a State-Prompt Visual Encoder (SPVE) module with a Denoising Alignment Fusion (DAF) module to achieve robust multi-modal semantic alignment and fusion. This leads to more stable and robust multi-modal joint representations for downstream tasks. Experimental results show that our method significantly outperforms existing techniques on the MonoMulti3D-ROPE dataset. Our dataset and code are available at https://github.com/JasonHuang516/MonoMulti-3DVG
Keyu Guo, Yongle Huang, Shijie Sun 0001, Mingtao Feng, Huansheng Song, Jianxin Li 0001, Naveed Akhtar, Ajmal Mian
CVPR3
2025 Mono3DVLT: Monocular-Video-Based 3D Visual Language Tracking
abstract
Visual-Language Tracking (VLT) is emerging as a promising paradigm to bridge the human-machine performance gap. For single objects, VLT broadens the problem scope to text-driven video comprehension. Yet, this direction is still confined to 2D spatial extents, currently lacking the ability to deal with 3D tracking in the confines of monocular video. Unfortunately, advances in 3D tracking mainly rely on expensive sensor inputs, e.g., point clouds, depth measurements, radar. Absence of language counterpart for the outputs of these mildly democratized sensors in the literature also hinders VLT expansion to 3D tracking. Addressing that, we make the first attempt towards extending VLT to 3D tracking based on monocular video. We present a comprehensive framework, introducing (i) the Monocular-Video-based 3D Visual Language Tracking (Mono3DVLT) task, (ii) a large-scale dataset for the task, called Mono3DVLT-V2X, and (iii) a customized neural model for the task. Our dataset is carefully curated, leveraging a Large Langauge Model (LLM) followed by human verification, composing natural language descriptions for 79,158 video sequences aiming at single object tracking, providing 2D and 3D bounding box annotations. Our neural model, termed Mono3DVLT-MT, is the first targeted approach for the Mono3DVLT task. Comprising the pipeline of multi-modal feature extractor, visual-language encoder, tracking decoder and a tracking head, our model sets a strong baseline for the task on Mono3DVLT-V2X. Experimental results show that our method significantly outperforms existing techniques on the Mono3DVLT-V2X dataset. Our dataset and code are available in https://github.com/hongkai-wei/Mono3DVLT.
Hongkai Wei, Shijie Sun 0001, Mingtao Feng, Hongli Hu, Huansheng Song, Naveed Akhtar, Ajmal Mian
CVPR3
2025 MTTM: Memory-Augmented with Mamba for 3D Medical Images Analysis
abstract
The rapid advancement of artificial intelligence has propelled the healthcare industry into a new era of diagnostic precision. A pivotal component of this evolution is the accurate classification of 3D medical images, which necessitates extracting robust feature representations capable of effectively modeling long-range dependencies within the data. This paper introduces the Mamba Token Turing Machine (MTTM), a novel architecture that integrates the efficiency of Mamba with the memory mechanisms of the Token Turing Machine (TTM), effectively addressing limitations of Transformers in long-range dependency modeling. MTTM’s Memory-Augmented Processing Unit (MAPU) employs four blending methods, achieving state-of-the-art accuracy and efficiency on the MedMNIST v2 dataset, thereby advancing diagnostic precision in 3D medical image analysis. The code is available at https://github.com/hongkai-wei/MTTM.
Hongkai Wei, Shijie Sun 0001, Huansheng Song, Keyu Guo, Yongfeng Bu
ICASSP3
2025 MSALNet: Capturing Contextual Relationships for Monocular 3D Visual Grounding
abstract
Monocular 3D Visual Grounding (Mono3DVG) aims to predict the 3D localization of objects in monocular RGB images based on natural language descriptions. This task has broad applications in areas such as autonomous driving, human-computer interaction and robotic manipulation, making it both a significant and challenging task. To tackle this challenge, we propose a novel network architecture, the Monocular Selective Attention Learning Network (MSALNet). This network enhances the understanding and localization of objects by introducing an Adaptive Learning Module (ALM) and a Vision-Text Interaction Encoder. Specifically, the ALM further learns the features extracted from the 3D scene and textual descriptions, capturing the contextual relationships within the features. This enables the model to better understand the meaning of both the scene and the text. Meanwhile, the Vision-Text Interaction Encoder facilitates refined cross-modal interaction and fusion, promoting alignment between visual and textual information and providing more discriminative feature representations. Experimental results demonstrate that our method achieves competitive performance on the Mono3DRefer dataset.
Keyu Guo, Yongle Huang, Yongfeng Bu, Hongkai Wei, Shijie Sun 0001
IJCNN5
2025 MGSGM: Multi-Granularity Selective Graph Mamba for Image-Text Retrieval
abstract
Image-Text Retrieval (ITR) serves as a fundamental task in multimodal information retrieval, aiming to bridge the semantic gap between visual and textual modalities. The core challenge resides in overcoming intrinsic modality discrepancies while effectively capturing cross-modal semantic relationships. Fine-grained retrieval approaches can model associations between salient visual regions and their linguistic counterparts, but are prone to interference from irrelevant visual embeddings and non-predicative textual components, leading to misalignments and degraded embedding quality for cross-modal retrieval. To address these challenges, we propose a Multi-Granularity Selective Graph Mamba (MGSGM) framework for discriminative embeddings. It uses the Selective Relationship Reasoning Module (SRRM) to suppress intra-modal redundancy, followed by graph refinement for medium-grained feature extraction. The Cross-modal Mamba Alignment (CMA) establishes selective semantic bridges between inter-modal entities. A hierarchical fusion mechanism ultimately integrates multi-granularity representations into discriminative embeddings. Extensive experiments on Flickr30K and MS-COCO datasets demonstrate the superiority of our method.
Yongle Huang, Yongfeng Bu, Keyu Guo, Shijie Sun 0001
ICMR6
2025 A novel pipeline for tunnel multi-object tracking integrating cross-modality and motion model
Yongfeng Bu, Juan Zhao 0007, Shijie Sun 0001, Haoxiang Liang, Huansheng Song, Xinzhou Ma
Expert Syst. Appl.4
2025 YCFA-Net: A unified framework for vehicle detection and fire anomaly recognition in tunnel scenarios
Lichen Liu, Huansheng Song, Shijie Sun 0001, Zhaoyang Zhang 0007, Zhaoquan Gu, Bangyang Wei, Hanke Luo
Expert Syst. Appl.4
2025 Bootstrapping vision-language transformer for monocular 3D visual grounding
abstract
Abstract In the task of 3D visual grounding using monocular RGB images, it is a challenging problem to perceive visual features and accurately predict the localization of 3D objects based on given geometric and appearances descriptions. Traditional text‐guided attention‐based methods have achieved better results than baselines, but it is argued that there is still potential for improvement in the area of multi‐modal fusion. Thus, Mono3DVG‐TRv2, an end‐to‐end transformer‐based architecture that employs a visual‐text multi‐modal encoder for the alignment and fusion of multi‐modal features, incorporating an enhanced transformer module proven in 2D detection, is introduced. The depth features predicted by the multi‐modal features and the visual‐text features are associated with the learnable queries in the decoder, facilitating more efficient and effective acquisition of geometric information in intricate scenes. Following a comprehensive comparison and ablation study on the Mono3DRefer dataset, this method achieves state‐of‐the‐art performance, markedly surpassing the prior approach. The code will be released at https://github.com/Jade-Ray/Mono3DVGv2 .
Shijie Sun 0001, Huansheng Song, Mingtao Feng, Chengzhong Wu
IET Image Process.2
2025 Causally-guided graph Mamba for detecting socially abnormal vehicle trajectories
Yongfeng Bu, Haoxiang Liang, Huansheng Song, Shijie Sun 0001, Zhaoyang Zhang 0007
Neurocomputing4
2025 SFAN: Selective Filter and Alignment Network for Cross-Modal Retrieval
abstract
Bridging the gap between visual and textual modalities effectively has consistently been a key challenge in cross-modal retrieval. Fine-grained matching approaches improve performance by precisely aligning salient region features in visual modality with word embeddings in textual modality. However, how to effectively and efficiently filter out irrelevant features (e.g., irrelevant background regions and nonmeaningful prepositions) in multimodality remains a significant challenge. Furthermore, capturing key cross-modal relationships while minimizing misalignment interference is crucial for effective cross-modal retrieval. In this work, we propose a novel approach called the selective filter and alignment network (SFAN) to tackle these challenges. First, we propose modality-specific selective filter modules (SFMs) to selectively and implicitly filter out redundant information within each modality. We then propose the state-space models (SSMs)-based selective alignment module (SAM) to selectively capture key correspondences and reduce the disturbance of irrelevant associations. Finally, we utilize a fusion operation to combine these embeddings from both SFM and SAM to derive the final embeddings for similarity computation. Extensive experiments on the Flickr30k, MS-COCO, and MSR-VTT datasets reveal that our proposed SFAN can effectively learn robust patterns, significantly outperforming the state-of-the-art (SOTA) cross-modal retrieval methods by a wide margin.
Yongle Huang, Shijie Sun 0001, Ningning Cui, Jianxin Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Yolo-3DMM for Simultaneous Multiple Object Detection and Tracking in Traffic Scenarios
abstract
Video-based multiple object tracking (MOT) is a fundamental task in intelligent transportation with applications ranging from automated traffic surveillance to autonomous driving. MOT methods commonly follow a tracking-by-detection paradigm, tracking objects by associating their detections across video frames. However, insofar, these methods have not used the entire vehicle trajectory motion characteristics to perform tracking, which converts the vehicle localization problem into a motion parameter estimation problem. Moreover, MOT methods mainly rely on off-the-shelf detectors. An independently trained detector is sub-optimal for the tracking-by-detection paradigm and adversely affects the overall system performance. In this article, we address these issues by proposing a novel MOT method for moving vehicles in traffic scenarios. Our tracker treats the vehicle tracks as unified 3D spatio-temporal trajectory instances and leverages the power of deep learning to extract vehicle motion from the 3D instances. We propose a new simultaneous detection and tracking network, called YOLO-3D Motion Model Network (Yolo-3DMM) that employs spatio-temporal features of traffic videos for simultaneous vehicle detection and tracking in an end-to-end manner. We adopt a variety of different vehicle tracking datasets to evaluate our method. Moreover, we also propose a tunnel MOT dataset from real highway tunnel surveillance in Guangdong, China to expand the experimental scenarios. To establish the efficacy of our method, we evaluate it on 100 different roadside traffic scenarios. Our method shows excellent performance on UA-DETRAC and Omni-MOT datasets. It achieves a PR-MOTA score of 29.40% on UA-DETRAC and gets a 69.7% MOTA score on the Omni-MOT dataset.
Lichen Liu, Huansheng Song, Shijie Sun 0001, Xian-Feng Han, Naveed Akhtar, Ajmal Mian
IEEE Trans. Intell. Transp. Syst.4
2022 Data association in multiple object tracking: A survey of recent techniques
Lionel Rakai, Huansheng Song, Shijie Sun 0001, Yanni Yang 0004
Expert Syst. Appl.3
2021 Novel methods for noisy 3D point cloud based object recognition
Xian-Feng Han, Xin-Yu Yan, Shijie Sun 0001
Multim. Tools Appl.3
2021 Deep Affinity Network for Multiple Object Tracking
abstract
Multiple Object Tracking (MOT) plays an important role in solving many fundamental problems in video analysis and computer vision. Most MOT methods employ two steps: Object Detection and Data Association. The first step detects objects of interest in every frame of a video, and the second establishes correspondence between the detected objects in different frames to obtain their tracks. Object detection has made tremendous progress in the last few years due to deep learning. However, data association for tracking still relies on hand crafted constraints such as appearance, motion, spatial proximity, grouping etc. to compute affinities between the objects in different frames. In this paper, we harness the power of deep learning for data association in tracking by jointly modeling object appearances and their affinities between different frames in an end-to-end fashion. The proposed Deep Affinity Network (DAN) learns compact, yet comprehensive features of pre-detected objects at several levels of abstraction, and performs exhaustive pairing permutations of those features in any two frames to infer object affinities. DAN also accounts for multiple objects appearing and disappearing between video frames. We exploit the resulting efficient affinity computations to associate objects in the current frame deep into the previous frames for reliable on-line tracking. Our technique is evaluated on popular multiple object tracking challenges MOT15, MOT17 and UA-DETRAC. Comprehensive benchmarking under twelve evaluation metrics demonstrates that our approach is among the best performing techniques on the leader board for these challenges. The open source implementation of our work is available at https://github.com/shijieS/SST.git.
Shijie Sun 0001, Naveed Akhtar, Huansheng Song, Ajmal Mian, Mubarak Shah
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Simultaneous Detection and Tracking with Motion Modelling for Multiple Object Tracking
Shijie Sun 0001, Naveed Akhtar, Huansheng Song, Ajmal Mian, Mubarak Shah
ECCV (24)1
2019 Benchmark Data and Method for Real-Time People Counting in Cluttered Scenes Using Depth Sensors
abstract
Vision-based automatic counting of people has widespread applications in intelligent transportation systems, security, and logistics. However, there is currently no large-scale public dataset for benchmarking approaches on this problem. This paper fills this gap by introducing the first real-world RGB-D people counting dataset (PCDS) containing over 4500 videos recorded at the entrance doors of buses in normal and cluttered conditions. It also proposes an efficient method for counting people in real-world cluttered scenes related to public transportations using depth videos. The proposed method computes a point cloud from the depth video frame and re-projects it onto the ground plane to normalize the depth information. The resulting depth image is analyzed for identifying potential human heads. The human head proposals are meticulously refined using a 3D human model. The proposals in each frame of the continuous video stream are tracked to trace their trajectories. The trajectories are again refined to ascertain reliable counting. People are eventually counted by accumulating the head trajectories leaving the scene. To enable effective head and trajectory identification, we also propose two different compound features. A thorough evaluation on PCDS demonstrates that our technique is able to count people in cluttered scenes with high accuracy at 45 fps on a 1.7-GHz processor, and hence it can be deployed for effective real-time people counting for intelligent transportation systems.
Shijie Sun 0001, Naveed Akhtar, Huansheng Song, ChaoYang Zhang, Jianxin Li 0001, Ajmal Mian
IEEE Trans. Intell. Transp. Syst.1