VLDB 2026 Research / reviewers in the wild / expert
Junge Shen
dblp:126/9199
· DBLP profile ↗
21ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0002-6563-9206ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Breakthrough in fine-grained video anomaly detection on highway: New benchmark and model
Chenlin Meng, Chi Zhang 0119, Zhaoyong Mao, Junge Shen, Zhiyong Cheng 0008 |
Expert Syst. Appl. | 5 |
| 2026 | Enhancing weakly supervised video anomaly detection via prototype-driven pseudo-labeling
Guanglin Liu, Junge Shen, Chi Zhang 0119, Zhaoyong Mao |
Vis. Comput. | 2 |
| 2025 | Boosting Weakly Supervised Video Anomaly Detection with Generative Description
Chenlin Meng, Zhaoyong Mao, Chi Zhang 0119, Junge Shen |
PRCV (5) | 5 |
| 2025 | Dynamic Anchor: Density Map Guided Small Object Detector for Tiny Persons
Xingzhou Xu, Zhaoyong Mao, Qinhao Tu, Junge Shen |
Comput. Vis. Image Underst. | 5 |
| 2025 | Pseudo-Label Guided Object Detection in Sparsely Annotated Underwater Optical ImagesabstractObject detection in underwater optical imagery plays a crucial role in various fields related to underwater exploration. However, manual annotation of such images often results in incomplete ground truth due to its severe degradation. In this study, we address the issue of incomplete supervision signals in degraded underwater images by reframing it as a sparse annotation challenge. Specifically, we present a novel method for object detection in sparsely annotated underwater scenarios. Our approach involves an effective pseudo-label generation network designed to produce labels for the degraded foreground lacking annotations. To mitigate potential background noise resulting from the discrepancy between the fixed confidence threshold and its dynamic distribution, we introduce a novel dynamic adaptive confidence threshold method. Additionally, a novel adaptive geometric prior-based noise reduction strategy is designed to eliminate noisy pseudo-labels with low-quality localization. We validate and analyze our approach through experiments on publicly available underwater optical image datasets. The results demonstrate that our approach achieves significant performance improvements across various sparsity conditions. Compared with existing state-of-the-art models, our proposed approach delivers significantly superior Average Precision (AP) performance while maintaining fast inference speeds. The code is available at https://github.com/chenyyyxxx/Underwater-aisi. Gangqi Chen, Zhaoyong Mao, Junge Shen, Zhiyong Cheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | SL-Seg: A CNN-Transformer Fusion Network for Road Surface and Lane Segmentation in Complex ScenariosabstractRoad image segmentation plays a pivotal role in traffic video surveillance for environmental perception. Precise segmentation of roads and lanes is essential for effective traffic monitoring and management. However, unlike the perspective encountered in autonomous driving, the surveillance perspective poses unique challenges due to its wider scope and susceptibility to complex environments. This complexity makes the segmentation task in road surveillance videos particularly demanding. To overcome these challenges, we introduce an end-to-end semantic segmentation network that leverages a CNN-Transformer architecture. Firstly, a spatial pyramid attention-style convolution (SP-AttnConv) module, built upon the Transformer is introduced, to ensure accurate segmentation across long distances while preserving fine boundary information. This module enhances local information and fosters a “global-local” feature fusion framework. Secondly, to tackle the issue of scale imbalance during segmentation, a lightweight multi-scale (LMS) module is introduced to capture multi-scale feature. Additionally, an occlusion relief branch (ORB) module is integrated into the decoder, specifically addressing occlusions caused by irrelevant objects. Recognizing the need for a dedicated benchmark dataset for road surface and lane segmentation, surface-lane (SL) for complex scenarios is built in our paper to promote the development of traffic surveillance system. Comparative experiments demonstrate that our method achieves the best overall performance on the SL dataset. Chenlin Meng, Qinhao Tu, Zhaoyong Mao, Junge Shen |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Enhancing Weakly Supervised Anomaly Detection in Surveillance Videos: The CLIP-Augmented Bimodal Memory Enhanced NetworkabstractAiming at the challenges of surveillance video anomaly detection(SVAD),especially the diversity and openness of its event types, we propose CLIP-Augmented Bimodal Memory Enhanced Network for weakly-supervised surveillance video anomaly detection. Specifically, we design a video feature extraction module based on CLIP feature, which significantly improves the ability to capture the semantic content of surveillance videos. Given the problem of semantic diversity of abnormal events, we further design a Bimodal Memory Unit(BMMU), which is used to enhance the model for all types of abnormal events by means of two kinds of memory module, storing the visual features and the textual descriptive features, in order to enhance the model's ability to remember and distinguish various types of anomalous features. Extensive experiments demonstrate that our approach achieves state-of-the-art performance on the UCF-Crime and XD-Violence benchmark datasets. Yinglong Wu, Zhaoyong Mao, Chenyang Yu, Guanglin Liu, Junge Shen |
ICARCV | 5 |
| 2024 | Advanced Object Detection in Multibeam Forward-Looking Sonar Images Using Linear Cross-Attention TechniquesabstractSonar images object detection plays a crucial role in marine resource exploration and defense. Existing methods encounter challenges arising from non-rigid, deformed and low-resolution objects in sonar images, making them difficult to build a stable feature representation. In this paper, we present a novel object detection method based on linear cross-attention, aiming to construct a more robust feature representation tailored for sonar objects. Specifically, we introduce a novel feature fusion network designed to efficiently extract global object context. It helps to construct a more robust feature representation, effectively improving the model’s ability to detect non-rigidly deformed and low-resolution objects. Moreover, we propose a linear attention mechanism to constitute the linear cross-attention module, leading to a significant reduction in computational load. Extensive experiments conducted on the public available dataset demonstrate the superiority of our approach. Our method surpasses a 1.4 mAP over the baseline, while requiring comparable parameters. Gangqi Chen, Zhaoyong Mao, Junge Shen |
ICIP | 3 |
| 2024 | StreamTrack: real-time meta-detector for streaming perception in full-speed domain driving scenarios
Weizhen Ge, Zhaoyong Mao, Jing Ren 0005, Junge Shen |
Appl. Intell. | 5 |
| 2024 | CoocNet: a novel approach to multi-label text classification with improved label co-occurrence modeling
Junge Shen, Zhaoyong Mao |
Appl. Intell. | 2 |
| 2024 | A Cooperative Training Framework for Underwater Object Detection on a Clearer ViewabstractUnderwater optical image object detection plays a crucial role in fields such as ocean exploration. However, constructing a comprehensive annotated dataset for training is challenging, especially when dealing with severely degraded underwater imagery. The sparsity of annotations can significantly reduce the performance of object detection algorithms. Existing methods designed for sparsely annotated object detection (SAOD) in terrestrial scenarios are not optimal for underwater conditions. To address these challenges, we propose a novel underwater cooperative training framework (CTF). Specifically, we propose a novel conjugate data generation module (CDGM) to tackle the issue of noise accumulation inherent in the existing data generation module, thereby greatly enhancing pseudo label generation. Furthermore, to mitigate the impacts of noisy pseudo labels, we present a pseudo label calibration strategy (PLCS) that manipulates the foreground confidence trend toward a low entropy distribution, effectively eliminating noisy pseudo labels. Finally, we propose a novel decoupled detection module to alleviate interference between position information and foreground confidence, further reducing noisy pseudo labels. Compared with methods tailored for terrestrial conditions with sparse annotations, our approach demonstrates superior performance in underwater scenarios. We conducted extensive experiments on various underwater datasets, including URPC2018, DUO, etc. The results show that our method outperforms the existing state-of-the-art by 2.0 mean average precision (mAP) on the URPC2018 and 0.9 mAP on the DUO datasets, while achieving state-of-the-art performance. Code will be released athttps://github.com/bobchenlut/coorporate-learing. Gangqi Chen, Zhaoyong Mao, Qinhao Tu, Junge Shen |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Multi-Scale Semantic Map Distillation for Lightweight Pavement Crack DetectionabstractTimely and accurate pavement crack detection plays a crucial role in urban transportation and pavement management. Recently, deep learning-based pavement crack detection approaches have witnessed significant advancements. However, high-performance networks heavily rely on complex network structures and massive model parameters, making their application in practical scenarios challenging. To achieve efficient yet accurate pavement crack detection, this paper presents a novel Lightweight Pavement Crack Detection based on Multi-Scale Semantic Map Distillation (LPCD-MSMD). Specifically, we improve the U-Net by designing a novel Cascaded U-Net (C-UNet) structure and incorporating additional branches to capture more subtle features. The proposed C-UNet is optimized by reducing the number of convolutional layers and compressing the channel dimensions, benefiting to a more lightweight model while maintaining comparable performance. Additionally, we utilize the original C-UNet as the teacher network and the simplified network as the student network for a lightweight pavement crack detection. Through the proposed multi-scale semantic map distillation strategy, the student network can acquire multi-scale output knowledge from the cascaded output of the teacher network, benefiting significantly improved performance of the CU-Net(s). Thorough experimental results demonstrate that the elaborately designed student network possesses a remarkably small parameter size of only 0.54 MB. Moreover, the performance evaluated on the Crack500 dataset and the GAPS384 dataset indicate the superiority of proposed method over several typical deep learning-based approaches including SegNet, FCN, BiSeNet, and U-Net. In particular, our method outperforms the second-best method by 0.09 and 0.124 in terms of F1 score evaluated on the Crack500 dataset and the GAPS384 dataset, respectively. Zhaoyong Mao, Junge Shen |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Be an Excellent Student: Review, Preview, and CorrectionabstractIn the letter, we propose a novel yet effective knowledge distillation scheme which mimics an all-round learning process of an excellent student from the teacher, i.e, knowledge review, knowledge preview, and knowledge correction, to acquire more informative and complementary knowledge. In the newly proposed method, to better leverage comprehensive feature knowledge from the teacher model, we propose Knowledge Review and Knowledge Preview Distillation to amalgamate multi-level features from different intermediate layers in both forward and backward pathways and fully distill them through hierarchical context loss, which greatly improves the student's feature learning efficiency. Moreover, we further present a Response Correction Mechanism to reinforce the prediction of student, which can more fully excavate the student's own knowledge, effectively alleviating the negative influence caused by the knowledge gap between the teacher and the student. We verify the effectiveness of our method with various networks on the CIFAR-100 datasets and the proposed method achieves competitive results compared with other state-of-the-art competitors. The code will be available athttps://github.com/kbzhang0505/RPC. Qizhi Cao, Kaibing Zhang, Xin He 0029, Junge Shen |
IEEE Signal Process. Lett. | 4 |
| 2022 | Assessing learning engagement based on facial expression recognition in MOOC's scenario
Junge Shen, Haopeng Yang, Zhiyong Cheng 0001 |
Multim. Syst. | 1 |
| 2022 | Remote Sensing Scene Classification Based on Attention-Enabled Progressively SearchingabstractRemote sensing image scene classification plays a significant role in remote sensing image analysis. Aiming at the problems of large transformation and scale variation of background and key objects in remote sensing images, we propose a neural architecture search (NAS) method based on attention search space. The network adaptively searches convolution, pooling, and attention operations in the appropriate layers. To ensure the stability of the searching process, a multistage network progressive fusion search method is proposed, which discards useless operations in stages, reduces the burden of search algorithm, and improves the search efficiency. Finally, paying attention to the association information between objects and scenes, a bottom-up multiscale fusion network connection strategy is proposed to fully reuse the semantics of multiscale feature maps in each stage. The experimental results show that the proposed method performs better than the manual method and the current neural network architecture search method. Junge Shen, Bin Cao 0006, Ruxin Wang 0002, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Image phylogeny tree construction based on local inheritance relationship correction
Junge Shen |
Multim. Tools Appl. | 2 |
| 2016 | Attraction recommendation: Towards personalized tourism via collective intelligence
Junge Shen, Cheng Deng 0002, Xinbo Gao 0001 |
Neurocomputing | 1 |
| 2016 | Landmark Reranking for Smart Travel Guide Systems by Combining and Analyzing Diverse MediaabstractAdvanced networking technologies and massive online social media have stimulated a booming growth of travel heterogeneous information in recent years. By employing such information, smart travel guide systems, such as landmark ranking systems, have been proposed to offer diverse online travel services. It is essential for a landmark ranking system to structure, analyze, and search the travel heterogeneous information to produce human-expected results. Therefore, currently the most fundamental yet challenging problems can be concluded: 1) how to fuse heterogeneous tourism information and 2) how to model landmark ranking. In this paper, a novel landmark search system is introduced based on a newly designed heterogeneous information fusion scheme and a query-dependent landmark ranking strategy. Different from the existing travel guide systems, the proposed system can effectively combine the heterogeneous information from multimodality media into a landmark reranking list via a user's query. Experimental results conducted on a large travel information collection illustrate the advantages of the proposed system in terms of both effectiveness and efficiency. Junge Shen, Jialie Shen 0001, Tao Mei 0001, Xinbo Gao 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2014 | The Evolution of Research on Multimedia Travel Guide Search and Recommender Systems
Junge Shen, Zhiyong Cheng 0001, Jialie Shen 0001, Tao Mei 0001, Xinbo Gao 0001 |
MMM (2) | 1 |
| 2013 | Image search reranking with multi-latent topical graphabstractImage search reranking has attracted extensive attention. However, existing image reranking approaches deal with different features independently while ignoring the latent topics among them. It is important to mine multi-latent topic from the features to solve the image search reranking problem. In this paper, we propose a new image reranking model, named reranking with multi-latent topical graph (RMTG), which not only exploits the explicit information of local and global features, but also mines multi-latent topic from these features. We evaluate RMTG over the MSRA-MM dataset and show that RMTG outperforms several existing reranking methods. Junge Shen, Tao Mei 0001, Qi Tian 0001, Xinbo Gao 0001 |
ISCAS | 1 |
| 2013 | Video archaeology: understanding video manipulation history
Junge Shen, Tao Mei 0001, Xinbo Gao 0001 |
Multim. Tools Appl. | 1 |