Mengke Song

dblp:172/2658 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0001-9618-0656ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Computer networks · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Adapting Visual Trackers to Dynamic View Transitions With Shift-View Prompt Tuning
abstract
Visual tracking is essential across numerous video analysis applications, surveillance systems, entertainment, and autonomous applications. However, most conventional state-of-the- art visual trackers are designed for constant-view scenarios with fixed camera viewpoints, and they only achieve satisfactory performance under stable visual features scenarios. In reality, visual tracking often encounters shift-view scenarios (e.g., sports broadcasting, ground-aerial surveillance), where cameras’ dynamic view transitions between ground and aerial views. These shifts lead to large variations in target scale and environmental complexity, resulting in inconsistent visual features that ultimately degrade the robustness of conventional visual trackers. Although developing a dedicated tracker for such shift-view scenarios is possible, it requires expensive temporal and computational costs. To address this challenge, we propose Shift-view Prompt Tuning, a cost-efficient method that enables conventional trackers to handle dynamic view transitions. We use sample pairs from different view datasets as prompts to guide the tracker’s adaptation. By embedding distinctive visual information from these prompts into training samples, we help the tracker learn about dynamic view transitions without requiring it to be relearned from scratch. This approach seamlessly transforms any constant-view trackers into shift-view trackers. Our extensive experiments on 14 datasets with 3 different view types show that our approach significantly enhances tracking performance. This advancement extends the application scope of current trackers and offers a robust solution for multimedia content production, sports analytics, and security monitoring in video analysis systems.
Chenglizhao Chen, Shaofeng Liang, Luming Li, Mengke Song, Xu Yu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Domain Adaptation for Cross-View Localization via Multi-Teacher Knowledge Distillation
abstract
Cross-view fine-grained localization estimates a ground camera’s pixel-level coordinates in aerial images by analyzing visual correspondences between views. Recent studies have made significant progress in this task, but when the models trained in a source area are directly applied to a new target area, their localization performance often suffers significant degradation due to the domain gap between the two areas. Moreover, obtaining accurate ground truth (GT) for the target area to retrain the models is prohibitively expensive. To adapt the localization model to the target area, this article proposes a weakly supervised learning approach based on multi-teacher knowledge distillation. This approach utilizes multiple pre-trained teacher models to make predictions for the target area and employs a learning-free cross-view instance matching and view alignment (CVMA) module to evaluate the quality of predicted coordinates from geometric, semantic, and visual perspectives. Based on the evaluation results, the best prediction is selected as pseudo-GT, and potential anomalous training samples are filtered out. The CVMA module also functions as a learning-free fine-grained localization method, achieving performance comparable to some learning-based methods. Our approach is validated on the VIGOR benchmark using three state-of-the-art models, and experimental results show that our method significantly improves the localization performance of models in the target area.
Chenglizhao Chen, Qianxi Yuan, Shaofeng Liang, Mengke Song, Xinyu Liu 0029
ACM Trans. Multim. Comput. Commun. Appl.4
2026 Channel-Wise Contribution Assessment for RGB-D Salient Object Detection
abstract
In RGB-D salient object detection (SOD), a common approach to improving accuracy is by using a dual-stream architecture to combine RGB and depth data. However, the effectiveness of depth information varies depending on the scene. In scenarios where depth maps provide a limited contribution, integrating them with RGB can be challenging and sometimes even detrimental to performance rather than enhancing it. Conventional RGB-D SOD methods often lack precision in assessing depth map quality, neglecting to account for the distinct contributions of various regions within the map, often resulting in a suboptimal fusion of regions where depth information is minimally beneficial or irrelevant. To address these issues, this article presents a novel channel-wise contribution assessment method that precisely evaluates the contributions of both the RGB and depth channels. By employing a controlled perturbation process to challenge the saliency detection model with specific, manageable disturbances, we are able to measure how much RGB and depth information each contributes to the final saliency map. Based on this analysis, we have developed a novel routing-style fusion of modality that dynamically adjusts the integration of the two modalities. This approach significantly lessens the negative impact of regions where depth data have a low, no, or even detrimental contribution, leading to a more effective and balanced fusion of RGB and depth information. Extensive experiments on multiple benchmark datasets demonstrate that the proposed method consistently achieves competitive performance and improves the robustness of RGB-D salient object detection across diverse and challenging scenarios.
Chenglizhao Chen, Mengke Song, Xinyu Liu 0029, Wenfeng Song
ACM Trans. Multim. Comput. Commun. Appl.3
2025 SharpEdge: High-quality data-driven monocular depth estimation for enhanced boundary precision
Mengke Song, Luming Li, Xu Yu 0001, Chenglizhao Chen
Eng. Appl. Artif. Intell.1
2025 Unveiling Context-Related Anomalies: Knowledge Graph Empowered Decoupling of Scene and Action for Human-Related Video Anomaly Detection
abstract
Video anomaly detection methods are mainly classified into two categories based on their primary feature types: appearance-based and action-based. Appearance-based methods rely on low-level visual features like color, texture, and shape, learning patterns specific to training scenes. While effective in familiar settings, they struggle with unknown or altered scenes due to poor generalization and limited understanding of action-scene relationships. In contrast, action-based methods focus on detecting action anomalies but often overlook contextual scene associations, leading to misjudgments (e.g., running on a street being deemed normal without considering scene context). To overcome these limitations, we propose a novel decoupling-based anomaly detection architecture (DecoAD). Its core lies in the decoupling and interweaving of scenes and actions, enabling explicit modeling of their complex relationships. By reconstructing these interactions using knowledge graphs, DecoAD achieves a deeper understanding of behaviors and contexts. This design ensures strong performance in both known and unknown scenarios, significantly enhancing generalization. To evaluate its effectiveness in dynamic scenes and its ability to handle scene-related anomalies, we introduce UFSR, the first video anomaly detection dataset featuring dynamic scenes and scene-related anomalies. DecoAD supports fully-supervised, weakly-supervised, and unsupervised settings, improving AUC on UBnormal by 1.1%, 3.1%, and 2.1% in fully-supervised, weakly-supervised, and unsupervised settings, and on UFSR by 1.2% and 8.2% in weakly-supervised and unsupervised settings. The source code and datasets are available at:https://github.com/liuxy3366/DecoAD.
Chenglizhao Chen, Xinyu Liu 0029, Mengke Song, Luming Li, Shaojiang Yuan, Xu Yu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 UNI-IQA: A Unified Approach for Mutual Promotion of Natural and Screen Content Image Quality Assessment
abstract
To date, the image quality assessment (IQA) research field has mainly focused on natural images (NIs)-based IQA and screen content images (SCIs)-based IQA. Usually, these two research branches are quite independent due to the large differences between NIs and SCIs, where NIs, captured by cameras directly, contain pictorial information solely, yet, SCIs, synthesized or GPU-rendered, have pictures and textures. Moreover, the distortion types are also different, and subjective scores of different datasets assigned by participants are usually not well aligned. So, due to the above-mentioned “domain shifts” and “dataset misalignments”, our research community has widely believed that it could be very difficult to achieve joint mutual promotions between NIs- and SCIs-based IQA. In this paper, we argue that despite the “differences”, there still are some “common characteristics” — our human visual system perceives the “pictures” in both SCIs and NIs almost the same way. Thus, we can still achieve mutual performance promotion if we can appropriately use the “common characteristics” between SCIs and NIs. Our key idea is to devise a “content-aware” data switch, which, from the perspective of input’s contents (i.e., pictures or textures), aims at letting the model automatically enhance the commonness and compress the discrepancies between the two tasks. Notice that none of the existing fusion schemes can reach this goal since they are actually content-unaware, degenerating the “mutual interactions” into “mutual interferences”. This paper is the first attempt to achieve full end-to-end “mutual interactions” between NIs- and SCIs-based IQA. Using the proposed switch, we are also the first to achieve solid mutual promotions for the two tasks, reaching new SOTA results.
Mengke Song, Chenglizhao Chen, Wenfeng Song, Yuming Fang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Pushing the Boundaries of Salient Object Detection: A Denoising-Driven Approach
abstract
Salient Object Detection (SOD) aims to identify the most attention-grabbing regions in an image and focuses on distinguishing salient objects from their backgrounds. Current SOD methods primarily use a discriminative approach, which works well for clear images but struggles in complex scenes with similar colors and textures between objects and backgrounds. To address these limitations, we introduce the diffusion-based salient object detection model (DiffSOD), which leverages a noise-to-image denoising process within a diffusion framework, enhancing saliency detection in both RGB and RGB-D images. Unlike conventional fusion-based SOD methods that directly merge RGB and depth information, we treat RGB and depth as distinct conditions, i.e., the appearance condition and the structure condition, respectively. These conditions serve as controls within the diffusion UNet architecture, guiding the denoising process. To facilitate this guidance, we employ two specialized control adapters: the appearance control adapter and the structure control adapter. Moreover, conventional denoising UNet models may struggle when handling low-quality depth maps, potentially introducing detrimental cues into the denoising process. To mitigate the impact of low-quality depth maps, we introduce a quality-aware filter. This filter selectively processes only high-quality depth data, ensuring that the denoising process is based on reliable information. Comparative evaluations on benchmark datasets have shown that DiffSOD substantially surpasses existing RGB and RGB-D saliency detection methods, improving average performance by 1.5% and 1.2% respectively, thus setting a new benchmark for diffusion-based dense prediction models in visual saliency detection.
Mengke Song, Luming Li, Xu Yu 0001, Chenglizhao Chen
IEEE Trans. Image Process.1
2025 Adapting Generic RGB-D Salient Object Detection for Specific Traffic Scenarios
abstract
Existing RGB-D salient object detection (SOD) models are primarily trained on general-purpose datasets, which may lead to domain shift issues when applied directly to new, specific scenes, such as stereo traffic datasets. Though “large-scale datasets (COME15K and ReDweb-S)” have been released, they only partially address the domain shift problem. From the perspective of data augmentation, this paper presents a novel solution, which follows a weakly-supervised way to adapt generic RGB-D SOD models for specific scenarios, with a focus on traffic scene imagery. Our key idea is to equip plain videos (specific scenarios, i.e., traffic scenes) with newly estimated saliency informative depth maps and pseudo-SOD GTs, enabling them to support the retraining of existing RGB-D SOD models for meeting the requirements of these specific scenes. To achieve this, we offer a fresh perspective on how depth information can be leveraged in the SOD task and introduce a new paradigm for extracting intrinsic information from optical flows derived from videos to refine RGB-D SOD models. Our method achieves a 1.2% improvement in F-measure on RGB-D datasets and a 27% enhancement on real-world street view datasets compared to baseline models. These results demonstrate the effectiveness of our approach in enhancing model adaptability for traffic scene imagery, even with limited target domain data. Codes, datasets, and results are available at https://github.com/MengkeSong/AGSS.
Chenglizhao Chen, Mengke Song, Chong Peng 0001
IEEE Trans. Intell. Transp. Syst.2
2024 Rethinking Object Saliency Ranking: A Novel Whole-Flow Processing Paradigm
abstract
Existing salient object detection methods are capable of predicting binary maps that highlight visually salient regions. However, these methods are limited in their ability to differentiate the relative importance of multiple objects and the relationships among them, which can lead to errors and reduced accuracy in downstream tasks that depend on the relative importance of multiple objects. To conquer, this paper proposes a new paradigm for saliency ranking, which aims to completely focus on ranking salient objects by their "importance order". While previous works have shown promising performance, they still face ill-posed problems. First, the saliency ranking ground truth (GT) orders generation methods are unreasonable since determining the correct ranking order is not well-defined, resulting in false alarms. Second, training a ranking model remains challenging because most saliency ranking methods follow the multi-task paradigm, leading to conflicts and trade-offs among different tasks. Third, existing regression-based saliency ranking methods are complex for saliency ranking models due to their reliance on instance mask-based saliency ranking orders. These methods require a significant amount of data to perform accurately and can be challenging to implement effectively. To solve these problems, this paper conducts an in-depth analysis of the causes and proposes a whole-flow processing paradigm of saliency ranking task from the perspective of "GT data generation", "network structure design" and "training protocol". The proposed approach outperforms existing state-of-the-art methods on the widely-used SALICON set, as demonstrated by extensive experiments with fair and reasonable comparisons. The saliency ranking task is still in its infancy, and our proposed unified framework can serve as a fundamental strategy to guide future work. The code and data will be available at https://github.com/MengkeSong/Saliency-Ranking-Paradigm.
Mengke Song, Dunquan Wu, Wenfeng Song, Chenglizhao Chen
IEEE Trans. Image Process.1
2023 A Comprehensive Survey on Video Saliency Detection With Auditory Information: The Audio-Visual Consistency Perceptual is the Key!
abstract
Video saliency detection (VSD) aims at fast locating the most attractive objects/things/patterns in a given video clip. Existing VSD-related works have mainly relied on the visual system but paid less attention to the audio aspect. In contrast, our audio system is the most vital complementary part of our visual system. Also, audio-visual saliency detection (AVSD), one of the most representative research topics for mimicking human perceptual mechanisms, is currently in its infancy, and none of the existing survey papers have touched on it, especially from the perspective of saliency detection. Thus, the ultimate goal of this paper is to provide an extensive review to bridge the gap between audio-visual fusion and saliency detection. In addition, as another highlight of this review, we have provided a deep insight into key factors that could directly determine AVSD deep models’ performances. We claim that the audio-visual consistency degree (AVC) — a long-overlooked issue, can directly influence the effectiveness of using audio to benefit its visual counterpart when performing saliency detection. Moreover, to make the AVC issue more practical and valuable for future followers, we have newly equipped almost all existing publicly available AVSD datasets with additional frame-wise AVC labels. Based on these upgraded datasets, we have conducted extensive quantitative evaluations to ground our claim on the importance of AVC in the AVSD task. In a word, our ideas and new sets serve as a convenient platform with preliminaries and guidelines, all of which can potentially facilitate future works in further promoting state-of-the-art (SOTA) performance.
Chenglizhao Chen, Mengke Song, Wenfeng Song, Li Guo 0016, Muwei Jian
IEEE Trans. Circuits Syst. Video Technol.2
2022 Improving RGB-D Salient Object Detection via Modality-Aware Decoder
abstract
Most existing RGB-D salient object detection (SOD) methods are primarily focusing on cross-modal and cross-level saliency fusion, which has been proved to be efficient and effective. However, these methods still have a critical limitation, i.e., their fusion patterns - typically the combination of selective characteristics and its variations, are too highly dependent on the network's non-linear adaptability. In such methods, the balances between RGB and D (Depth) are formulated individually considering the intermediate feature slices, but the relation at the modality level may not be learned properly. The optimal RGB-D combinations differ depending on the RGB-D scenarios, and the exact complementary status is frequently determined by multiple modality-level factors, such as D quality, the complexity of the RGB scene, and degree of harmony between them. Therefore, given the existing approaches, it may be difficult for them to achieve further performance breakthroughs, as their methodologies belong to some methods that are somewhat less modality sensitive. To conquer this problem, this paper presents the Modality-aware Decoder (MaD). The critical technical innovations include a series of feature embedding, modality reasoning, and feature back-projecting and collecting strategies, all of which upgrade the widely-used multi-scale and multi-level decoding process to be modality-aware. Our MaD achieves competitive performance over other state-of-the-art (SOTA) models without using any fancy tricks in the decoder's design. Codes and results will be publicly available at https://github.com/MengkeSong/MaD.
Mengke Song, Wenfeng Song, Guowei Yang 0002, Chenglizhao Chen
IEEE Trans. Image Process.1