VLDB 2026 Research / reviewers in the wild / expert
Huansheng Song
dblp:86/11308 · also HuanSheng Song
· DBLP profile ↗
28ranked-venue papers
1as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From perception to cognition: Unifying multi-object 3D visual grounding and dense captioning in monocular images
Keyu Guo, Yongle Huang, Hongkai Wei, Shijie Sun 0001, Mingtao Feng, Huansheng Song |
Expert Syst. Appl. | 9 |
| 2026 | ArgusNet: Understanding 3D scenes more like humans
Keyu Guo, Hongkai Wei, Yongle Huang, Shijie Sun 0001, Mingtao Feng, Huansheng Song, Jianxin Li 0001 |
Neurocomputing | 7 |
| 2026 | MedHyCLIP: Hyperbolic CLIP adaptation for universal medical anomaly detection
Keyu Guo, Hongkai Wei, Yongle Huang, Shijie Sun 0001, Yueming Shi, Huansheng Song, Nicola Strisciuglio |
Pattern Recognit. | 8 |
| 2025 | Beyond Human Perception: Understanding Multi-Object World from Monocular ViewabstractLanguage and binocular vision play a crucial role in human understanding of the world. Advancements in artificial intelligence have also made it possible for machines to develop 3D perception capabilities essential for high-level scene understanding. However, only monocular cameras are often available in practice due to cost and space constraints. Enabling machines to achieve accurate 3D understanding from a monocular view is practical but presents significant challenges. We introduce MonoMulti-3DVG, a novel task aimed at achieving multi-object 3D Visual Grounding (3DVG) based on monocular RGB images, allowing machines to better understand and interact with the 3D world. To this end, we construct a large-scale benchmark dataset, MonoMulti3D-ROPE, and propose a model, CyclopsNet that integrates a State-Prompt Visual Encoder (SPVE) module with a Denoising Alignment Fusion (DAF) module to achieve robust multi-modal semantic alignment and fusion. This leads to more stable and robust multi-modal joint representations for downstream tasks. Experimental results show that our method significantly outperforms existing techniques on the MonoMulti3D-ROPE dataset. Our dataset and code are available at https://github.com/JasonHuang516/MonoMulti-3DVG Keyu Guo, Yongle Huang, Shijie Sun 0001, Mingtao Feng, Huansheng Song, Jianxin Li 0001, Naveed Akhtar, Ajmal Mian |
CVPR | 7 |
| 2025 | Mono3DVLT: Monocular-Video-Based 3D Visual Language TrackingabstractVisual-Language Tracking (VLT) is emerging as a promising paradigm to bridge the human-machine performance gap. For single objects, VLT broadens the problem scope to text-driven video comprehension. Yet, this direction is still confined to 2D spatial extents, currently lacking the ability to deal with 3D tracking in the confines of monocular video. Unfortunately, advances in 3D tracking mainly rely on expensive sensor inputs, e.g., point clouds, depth measurements, radar. Absence of language counterpart for the outputs of these mildly democratized sensors in the literature also hinders VLT expansion to 3D tracking. Addressing that, we make the first attempt towards extending VLT to 3D tracking based on monocular video. We present a comprehensive framework, introducing (i) the Monocular-Video-based 3D Visual Language Tracking (Mono3DVLT) task, (ii) a large-scale dataset for the task, called Mono3DVLT-V2X, and (iii) a customized neural model for the task. Our dataset is carefully curated, leveraging a Large Langauge Model (LLM) followed by human verification, composing natural language descriptions for 79,158 video sequences aiming at single object tracking, providing 2D and 3D bounding box annotations. Our neural model, termed Mono3DVLT-MT, is the first targeted approach for the Mono3DVLT task. Comprising the pipeline of multi-modal feature extractor, visual-language encoder, tracking decoder and a tracking head, our model sets a strong baseline for the task on Mono3DVLT-V2X. Experimental results show that our method significantly outperforms existing techniques on the Mono3DVLT-V2X dataset. Our dataset and code are available in https://github.com/hongkai-wei/Mono3DVLT. Hongkai Wei, Shijie Sun 0001, Mingtao Feng, Hongli Hu, Huansheng Song, Naveed Akhtar, Ajmal Mian |
CVPR | 9 |
| 2025 | MTTM: Memory-Augmented with Mamba for 3D Medical Images AnalysisabstractThe rapid advancement of artificial intelligence has propelled the healthcare industry into a new era of diagnostic precision. A pivotal component of this evolution is the accurate classification of 3D medical images, which necessitates extracting robust feature representations capable of effectively modeling long-range dependencies within the data. This paper introduces the Mamba Token Turing Machine (MTTM), a novel architecture that integrates the efficiency of Mamba with the memory mechanisms of the Token Turing Machine (TTM), effectively addressing limitations of Transformers in long-range dependency modeling. MTTM’s Memory-Augmented Processing Unit (MAPU) employs four blending methods, achieving state-of-the-art accuracy and efficiency on the MedMNIST v2 dataset, thereby advancing diagnostic precision in 3D medical image analysis. The code is available at https://github.com/hongkai-wei/MTTM. Hongkai Wei, Shijie Sun 0001, Huansheng Song, Keyu Guo, Yongfeng Bu |
ICASSP | 4 |
| 2025 | Highway spillage detection using an improved STPM anomaly detection network from a surveillance perspective
Haoxiang Liang, Huansheng Song, Shaoyang Zhang, Yongfeng Bu |
Appl. Intell. | 2 |
| 2025 | A novel pipeline for tunnel multi-object tracking integrating cross-modality and motion model
Yongfeng Bu, Juan Zhao 0007, Shijie Sun 0001, Haoxiang Liang, Huansheng Song, Xinzhou Ma |
Expert Syst. Appl. | 7 |
| 2025 | YCFA-Net: A unified framework for vehicle detection and fire anomaly recognition in tunnel scenarios
Lichen Liu, Huansheng Song, Shijie Sun 0001, Zhaoyang Zhang 0007, Zhaoquan Gu, Bangyang Wei, Hanke Luo |
Expert Syst. Appl. | 3 |
| 2025 | Bootstrapping vision-language transformer for monocular 3D visual groundingabstractAbstract In the task of 3D visual grounding using monocular RGB images, it is a challenging problem to perceive visual features and accurately predict the localization of 3D objects based on given geometric and appearances descriptions. Traditional text‐guided attention‐based methods have achieved better results than baselines, but it is argued that there is still potential for improvement in the area of multi‐modal fusion. Thus, Mono3DVG‐TRv2, an end‐to‐end transformer‐based architecture that employs a visual‐text multi‐modal encoder for the alignment and fusion of multi‐modal features, incorporating an enhanced transformer module proven in 2D detection, is introduced. The depth features predicted by the multi‐modal features and the visual‐text features are associated with the learnable queries in the decoder, facilitating more efficient and effective acquisition of geometric information in intricate scenes. Following a comprehensive comparison and ablation study on the Mono3DRefer dataset, this method achieves state‐of‐the‐art performance, markedly surpassing the prior approach. The code will be released at https://github.com/Jade-Ray/Mono3DVGv2 . Shijie Sun 0001, Huansheng Song, Mingtao Feng, Chengzhong Wu |
IET Image Process. | 4 |
| 2025 | Causally-guided graph Mamba for detecting socially abnormal vehicle trajectories
Yongfeng Bu, Haoxiang Liang, Huansheng Song, Shijie Sun 0001, Zhaoyang Zhang 0007 |
Neurocomputing | 3 |
| 2025 | VRAR: Video-Radar Automatic Registration Method Based on Trajectory Spatiotemporal Features and Bidirectional MappingabstractAutomating video and radar spatial registration without sensor layout constraints is crucial for enhancing the flexibility of perception systems. However, this remains challenging due to the lack of effective approaches for constructing and utilizing matching information between heterogeneous sensors. Existing methods rely on human intervention or prior knowledge, making it difficult to achieve true automation. Consequently, establishing a registration model that automatically extracts matching information from heterogeneous sensor data remains a key challenge. To address these issues, we propose a novel Video-Radar Automatic Registration (VRAR) method based on vehicle trajectory spatiotemporal feature encoding and a bidirectional mapping network. We first establish a unified representation for heterogeneous sensor data by encoding spatiotemporal features of vehicle trajectories. Based on this, we automatically extract a large number of high-quality matching points from synchronized trajectory pairs using a frame synchronization strategy. Subsequently, we utilize the proposed Video-Radar Bidirectional Mapping Network to process these matching points. This network learns the bidirectional mapping between the two sensor modalities, extending the alignment from discrete local observation points to the entire observable space. Experimental results demonstrate that the VRAR method exhibits significant performance advantages in various traffic scenarios, verifying its effectiveness and generalizability. This capability of automated and adaptive registration highlights the method’s potential for broader applications in heterogeneous sensor integration. Kong Li, Xuan Wang 0021, Huansheng Song |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | A Nonoverlapping Sampling Approach With Peak Data Utilization for Hyperspectral ClassificationabstractData-driven hyperspectral image classification has gained significant attention across various applications, leading to the development of numerous novel neural networks for effective classification. However, existing spatial-spectral works often overlook the critical aspect of how training and testing sets are split, resulting in data leakage between these sets and suboptimal performance in real-world scenarios. To address this issue, this paper proposes a non-overlapping sampling approach named PDUnS, drawing inspiration from the theory of extremal polyominoes to achieve peak data utilization, that is, to maximize the number of testing patches. Specifically, PDUnS begins by determining the number of training patchesnand a random seed point in each connected componentRof each class. Subsequently, among all constructed polyominoes that cover the seed point and intersect at leastnpixels withR, our PDUnS method searches for those with the minimum perimeterp(n). Finally, we remove pixels in the polyominoes that are not inR, and then delete pixels with a degree of 2 and farthest from the boundary ofRuntilnpixels are retained to construct training patches. Experimental results and in-depth analysis on the Indian Pines dataset reveal that PDUnS demonstrates an average 12.38% increase in the number of testing samples to the comparing method. This affirms the superior effectiveness of PDUnS in maximizing data utilization. The code will be available at https://github.com/Yanzi-S/PDUnS. Yanzi Shi, Yaping Yin, Huansheng Song |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2024 | Yolo-3DMM for Simultaneous Multiple Object Detection and Tracking in Traffic ScenariosabstractVideo-based multiple object tracking (MOT) is a fundamental task in intelligent transportation with applications ranging from automated traffic surveillance to autonomous driving. MOT methods commonly follow a tracking-by-detection paradigm, tracking objects by associating their detections across video frames. However, insofar, these methods have not used the entire vehicle trajectory motion characteristics to perform tracking, which converts the vehicle localization problem into a motion parameter estimation problem. Moreover, MOT methods mainly rely on off-the-shelf detectors. An independently trained detector is sub-optimal for the tracking-by-detection paradigm and adversely affects the overall system performance. In this article, we address these issues by proposing a novel MOT method for moving vehicles in traffic scenarios. Our tracker treats the vehicle tracks as unified 3D spatio-temporal trajectory instances and leverages the power of deep learning to extract vehicle motion from the 3D instances. We propose a new simultaneous detection and tracking network, called YOLO-3D Motion Model Network (Yolo-3DMM) that employs spatio-temporal features of traffic videos for simultaneous vehicle detection and tracking in an end-to-end manner. We adopt a variety of different vehicle tracking datasets to evaluate our method. Moreover, we also propose a tunnel MOT dataset from real highway tunnel surveillance in Guangdong, China to expand the experimental scenarios. To establish the efficacy of our method, we evaluate it on 100 different roadside traffic scenarios. Our method shows excellent performance on UA-DETRAC and Omni-MOT datasets. It achieves a PR-MOTA score of 29.40% on UA-DETRAC and gets a 69.7% MOTA score on the Omni-MOT dataset. Lichen Liu, Huansheng Song, Shijie Sun 0001, Xian-Feng Han, Naveed Akhtar, Ajmal Mian |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Transfer Learning With Nonlinear Spectral Synthesis for Hyperspectral Target DetectionabstractSpectral distortion severely limits detection performance in hyperspectral imagery, while feature learning with neural networks could provide sufficient capacity to enhance spectral consistency. This paper designs an end-to-end hyperspectral target detection (HTD) network based on transfer learning and nonlinear spectral synthesis (TLNSS). We first utilize bilinear mixture model (BMM) to synthesize nonlinear target and background spectra for training sample augmentation, which could better characterize ground objects in complex environments. Due to the mutual constraints between the quantity and diversity of the synthesized spectra, transfer learning is introduced to further address data insufficiency. Specifically, we propose an asymmetric autoencoder with a particularly designed multi-level loss to maximally distinguish the reconstruction residuals of background and target, where the multi-scale feature extraction sub-network is trained with abundant reference data, and the simple restoration sub-network is updated with the simulated spectra. To effectively reconstruct the input as expected, the features extracted from different blocks are complementarily integrated through residual attention. Lastly, we accumulate reconstruction residuals across all levels for final detection. The experimental results and ablation analysis of single-data detection on three hyperspectral images verify the superiority and effectiveness of the proposed method, and further cross-data detection consolidates the satisfactory tolerance of TLNSS to spectral variation. Yanzi Shi, Yaping Yin, Huansheng Song, Yunsong Li 0001, Paolo Gamba |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Data association in multiple object tracking: A survey of recent techniques
Lionel Rakai, Huansheng Song, Shijie Sun 0001, Yanni Yang 0004 |
Expert Syst. Appl. | 2 |
| 2022 | Parallelized Nonlinear Target Detection for Asbestos Identification in Large-Scale Remote Sensing DataabstractDue to the side effects of asbestos on human health and environments, many countries have banned the use of asbestos-containing materials, but there are still illegal products with asbestos in daily life. In order to investigate the distributions of asbestos to facilitate its removal, this paper studies the feasibility of asbestos identification with HyperSpectral (HS) and panchromatic (PAN) data, taking images captured by the PRISMA and ZY1E 2D satellites over Pavia, Italy as examples. In this work, a pansharpening method with guided filter was used to improve HS image quality in terms of spectral fidelity and spatial details. Then, the possible location of asbestos could be obtained by a nonlinear target detector named BSTD. Considering high computational cost for large-scale remote sensing data processing, we further develop BSTD to its parallelized version (denoted as PBSTD). Given the groundtruth of asbestos over Pavia by the Regional Environmental Protection Agency-ARPA Lombardia, our PBSTD and several popular methods are evaluated from both qualitative and quantitative perspectives, showing that most algorithms could correctly detect large-size asbestos roofs, and the nonlinear PBSTD and MSDinter perform better in small-size asbestos identification than other linear detectors. However, the detection accuracy on small-size asbestos is insufficient in practical applications, which indicates that there are still issues to achieve accurate small-size asbestos identification using coarse-spatial-resolution spaceborne remote sensing. Yanzi Shi, Jiahui Qu, Yunsong Li 0001, Huansheng Song, Anna Vizziello, Paolo Gamba |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2021 | Multi-camera traffic scene mosaic based on camera calibrationabstractAbstract Recently, the traffic application based on vision in the single traffic‐monitoring scene has been widely studied and developed. However, cross‐regional research is still in its infancy. In order to help solve the application of cross‐regional traffic surveillance scenarios, this paper proposes a more reliable and accurate road scene mosaic method under multi‐camera surveillance. The mosaic road panorama contains physical information, which can be used to achieve a cross‐regional measurement. It also lays the foundation for vehicle spatial location, vehicle speed and traffic incident detection across regions. First, the mapping relationship between the three‐dimensional sub‐world coordinates and their corresponding two‐dimensional image coordinates is established by camera calibration. Second, the projection transformation relationship between two cameras is established by two sub‐world coordinate systems and their common information. Finally, we use the proposed inverse projection idea and translation vector relationship to complete the mosaic of two traffic‐monitoring road scenes. The experimental results show that the camera calibration accuracy can reach more than 97% in a single scene. The measurement accuracy of the mosaic block is over 95%. The experimental results show that the proposed method has a higher accuracy, which has great value in related theoretical research and practical applications. Feifan Wu, Huansheng Song, Wei Wang 0026 |
IET Comput. Vis. | 2 |
| 2021 | Deep Affinity Network for Multiple Object TrackingabstractMultiple Object Tracking (MOT) plays an important role in solving many fundamental problems in video analysis and computer vision. Most MOT methods employ two steps: Object Detection and Data Association. The first step detects objects of interest in every frame of a video, and the second establishes correspondence between the detected objects in different frames to obtain their tracks. Object detection has made tremendous progress in the last few years due to deep learning. However, data association for tracking still relies on hand crafted constraints such as appearance, motion, spatial proximity, grouping etc. to compute affinities between the objects in different frames. In this paper, we harness the power of deep learning for data association in tracking by jointly modeling object appearances and their affinities between different frames in an end-to-end fashion. The proposed Deep Affinity Network (DAN) learns compact, yet comprehensive features of pre-detected objects at several levels of abstraction, and performs exhaustive pairing permutations of those features in any two frames to infer object affinities. DAN also accounts for multiple objects appearing and disappearing between video frames. We exploit the resulting efficient affinity computations to associate objects in the current frame deep into the previous frames for reliable on-line tracking. Our technique is evaluated on popular multiple object tracking challenges MOT15, MOT17 and UA-DETRAC. Comprehensive benchmarking under twelve evaluation metrics demonstrates that our approach is among the best performing techniques on the leader board for these challenges. The open source implementation of our work is available at https://github.com/shijieS/SST.git. Shijie Sun 0001, Naveed Akhtar, Huansheng Song, Ajmal Mian, Mubarak Shah |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Simultaneous Detection and Tracking with Motion Modelling for Multiple Object Tracking
Shijie Sun 0001, Naveed Akhtar, Huansheng Song, Ajmal Mian, Mubarak Shah |
ECCV (24) | 4 |
| 2020 | A robust meaningful image encryption scheme based on block compressive sensing and SVD embedding
Liya Zhu, Huansheng Song, Maode Yan |
Signal Process. | 2 |
| 2020 | A high accurate vehicle speed estimation method
Sheng-Nan Lu, Yu-Ping Wang 0002, Huansheng Song |
Soft Comput. | 3 |
| 2019 | Benchmark Data and Method for Real-Time People Counting in Cluttered Scenes Using Depth SensorsabstractVision-based automatic counting of people has widespread applications in intelligent transportation systems, security, and logistics. However, there is currently no large-scale public dataset for benchmarking approaches on this problem. This paper fills this gap by introducing the first real-world RGB-D people counting dataset (PCDS) containing over 4500 videos recorded at the entrance doors of buses in normal and cluttered conditions. It also proposes an efficient method for counting people in real-world cluttered scenes related to public transportations using depth videos. The proposed method computes a point cloud from the depth video frame and re-projects it onto the ground plane to normalize the depth information. The resulting depth image is analyzed for identifying potential human heads. The human head proposals are meticulously refined using a 3D human model. The proposals in each frame of the continuous video stream are tracked to trace their trajectories. The trajectories are again refined to ascertain reliable counting. People are eventually counted by accumulating the head trajectories leaving the scene. To enable effective head and trajectory identification, we also propose two different compound features. A thorough evaluation on PCDS demonstrates that our technique is able to count people in cluttered scenes with high accuracy at 45 fps on a 1.7-GHz processor, and hence it can be deployed for effective real-time people counting for intelligent transportation systems. Shijie Sun 0001, Naveed Akhtar, Huansheng Song, ChaoYang Zhang, Jianxin Li 0001, Ajmal Mian |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2018 | Early ramp warning using vehicle behavior analysis
Zefa Wei, Xuan Wang 0021, Pannong Li, Huansheng Song |
Soft Comput. | 6 |
| 2018 | Vehicle trajectory clustering based on 3D information via a coarse-to-fine strategy
Huansheng Song, Xuan Wang 0021, Zhaoyang Zhang 0001 |
Soft Comput. | 1 |
| 2016 | Digital weighted autocorrelation receiver using channel characteristic sequences for transmitted reference UWB communication systemsabstractWeighted autocorrelation receivers have been proposed in the literature to suppress noise or interference for transmitted reference ultra-wideband communication systems. Usually weight optimization is performed for partitioned integration bins. To improve the optimization effect, this paper proposes a digital weighted autocorrelation receiver, in which the sampled auto-correlated signal is first rearranged following a channel characteristic vector that sorts the strengths of the received channel samples, and then it is divided into multiple partitions. Finally, the integration bins corresponding to these partitions are weighted according to the minimum mean square error principle. Results show that compared to the digital versions of existing weighted autocorrelation receivers, the proposed digital scheme can achieve significantly better bit error rate performance with limited penalty in terms of implementation complexity. Zhonghua Liang, Xiaodai Dong, Huansheng Song |
WCNC | 4 |
| 2015 | Robust Framework of Single-Frame Face Superresolution Across Head Pose, Facial Expression, and Illumination VariationsabstractThis paper presents a robust framework to solve the face hallucination problem across multiple factors, i.e., different expressions, head poses, and illuminations. It proposes a redundant transformation with diagonal loading for modeling the mappings among different new face factors, and a local reconstruction with geometry and position constraints for incorporating image details in the new factor spaces. Our proposed redundant and sparse strategies are discussed, and the experiments indicate that it is not necessary to adopt sparse representation in the proposed framework. The experimental results demonstrate that the proposed framework offers robustness when dealing with the inputs that have different expressions, head poses, and illuminations compared with the state-of-the-art methods, can generate high-resolution face images with better image qualities than the hierarchical tensor-based method, and improves the state of the art from single one output to multiple outputs with new factors. Huansheng Song, Xueming Qian |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2012 | Sparse representation and position prior based face hallucination upon classified over-complete dictionaries
Hiêp Quang Luong, Wilfried Philips, Huansheng Song |
Signal Process. | 4 |