VLDB 2026 Research / reviewers in the wild / expert
Guangxiao Ma
dblp:183/2777
· DBLP profile ↗
13ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-8968-9304ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ConsistencyTrack: A robust multi-object tracker with a generation strategy of consistency model
Lifan Jiang, Zhihui Wang 0003, Siqi Yin, Guangxiao Ma, Peng Zhang 0057, Boxi Wu 0001 |
Pattern Recognit. | 4 |
| 2025 | Comprehensive Feature Processing Based on Attention Mechanism for Co-Salient Object DetectionabstractCo-salient object detection (CoSOD) aims to detect common salient objects across multiple related images. However, existing methods often struggle with limited attention coverage, missing some co-salient objects. To address this, we propose a two-stage feature processing module (FPM) comprising comprehensive feature extraction module (CFE) and feature enhancement module (FEM). CFE extracts comprehensive cosalient features while reducing background noise, and FEM enhances feature representation and adjusts attention weights for full object coverage. Additionally, we introduce an adversarial learning module (ALM) to improve prediction quality by reducing noise in the co-salient regions. Extensive experiments on three benchmark datasets—CoCA, CoSOD3k, and CoSal2015—demonstrate that our model significantly outperforms state-of-the-art methods. The source code is available at https://github.com/yaobaimiao/CFPAM. Guohua Lv, Mao Yuan, Zengbin Zhang, Zhengyang Zhang, Zhenhui Ding, Guangxiao Ma |
ICASSP | 6 |
| 2025 | Self-supervised Co-salient Object Detection via Unified Multi-granularity Feature Learning
Mao Yuan, Guohua Lv, Guangxiao Ma, Zhengyang Zhang |
PRCV (16) | 3 |
| 2025 | Saliency-Free and Aesthetic-Aware Panoramic Video NavigationabstractMost of the existing panoramic video navigation approaches are saliency-driven, whereby off-the-shelf saliency detection tools are directly employed to aid the navigation approaches in localizing video content that should be incorporated into the navigation path. In view of the dilemma faced by our research community, we rethink if the "saliency clues" are really appropriate to serve the panoramic video navigation task. According to our in-depth investigation, we argue that using "saliency clues" cannot generate a satisfying navigation path, failing to well represent the given panoramic video, and the views in the navigation path are also low aesthetics. In this paper, we present a brand-new navigation paradigm. Although our model is still trained on eye-fixations, our methodology can additionally enable the trained model to perceive the "meaningful" degree of the given panoramic video content. Outwardly, the proposed new approach is saliency-free, but inwardly, it is developed from saliency but biasing more to be "meaningful-driven"; thus, it can generate a navigation path with more appropriate content coverage. Besides, this paper is the first attempt to devise an unsupervised learning scheme to ensure all localized meaningful views in the navigation path have high aesthetics. Thus, the navigation path generated by our approach can also bring users an enjoyable watching experience. As a new topic in its infancy, we have devised a series of quantitative evaluation schemes, including objective verifications and subjective user studies. All these innovative attempts would have great potential to inspire and promote this research field in the near future. Chenglizhao Chen, Guangxiao Ma, Wenfeng Song, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | IDFusion: An Infrared and Visible Image Fusion Network for Illuminating DarknessabstractThe purpose of infrared and visible image fusion is to combine the background information of the visible images and the thermal target information of the infrared images. Current fusion methods often neglect the challenges of low-light conditions. In the night scene, the existing methods fail to capture texture information from the visible images that is hidden in the darkness, producing suboptimal fusion results that can hinder subsequent visual applications. Therefore, we propose a method for infrared and visible image fusion under night scenes, termed as IDFusion. Specifically, our method is divided into three parts, first, a dense auto-encoder is designed to obtain more useful features from the source images. Then we design a brightness enhancement network that removes visible degraded illumination maps to obtain brightness-enhanced features. Finally, a texture fusion network is designed so that the infrared features and enhanced visible features avoid texture loss during the fusion process. Experimental results show that our network can obtain better fusion results, outperforming state-of-the-art methods in terms of subjective visual effects and quantitative metrics. Guohua Lv, Xiyan Wang, Zhonghe Wei, Jinyong Cheng, Guangxiao Ma, Hanju Bao |
CSCWD | 5 |
| 2024 | Rafmnet: Reinforced Attention Fusion and Multiscale Network For Noisy Infrared and Visible Image FusionabstractThe purpose of infrared and visible image fusion is to combine the advantages of different types of images to produce more robust and informative images. However, if the source images are noisy, existing fusion methods may not produce clear results. To address this issue, we propose a novel method for infrared and visible image fusion with noise reduction. This method enhances the visual perception of fused images by integrating features of different scales extracted by the denoising network into the fusion network. By using deformable convolutional denoising networks, noise in images can be removed and features can be enhanced. Then, a set of reinforced attention fusion modules (RAFM) are designed to fuse the features extracted by the denoising network. Experimental results demonstrate the effectiveness of our proposed method, which outperforms existing state-of-the-art methods in terms of fusion accuracy and visual perception. Guohua Lv, Xiyan Wang, Yongbiao Gao, Yi Zhai 0003, Guixin Zhao, Guangxiao Ma |
ICIP | 6 |
| 2024 | GLEGNet: Infrared and Visible Image Fusion via Global-Local Feature Extraction and Edge-Gradient Preservation
Guohua Lv, Wenkuo Song, Zhonghe Wei, Aimei Dong, Jinyong Cheng, Guangxiao Ma |
ICONIP (7) | 6 |
| 2024 | CEDP-YOLO: UAV Object Detection Based on Context Enhancement and Dynamic Perception
Zhenhui Ding, Zengbin Zhang, Mao Yuan, Guangxiao Ma, Guohua Lv |
PRCV (3) | 4 |
| 2024 | Scd-yolo: a novel object detection method for efficient road crack detection
Kuiye Ding, Zhenhui Ding, Zengbin Zhang, Mao Yuan, Guangxiao Ma, Guohua Lv |
Multim. Syst. | 5 |
| 2021 | Rethinking Image Salient Object Detection: Object-Level Semantic Saliency Reranking First, Pixelwise Saliency Refinement LaterabstractHuman attention is an interactive activity between our visual system and our brain, using both low-level visual stimulus and high-level semantic information. Previous image salient object detection (SOD) studies conduct their saliency predictions via a multitask methodology in which pixelwise saliency regression and segmentation-like saliency refinement are conducted simultaneously. However, this multitask methodology has one critical limitation: the semantic information embedded in feature backbones might be degenerated during the training process. Our visual attention is determined mainly by semantic information, which is evidenced by our tendency to pay more attention to semantically salient regions even if these regions are not the most perceptually salient at first glance. This fact clearly contradicts the widely used multitask methodology mentioned above. To address this issue, this paper divides the SOD problem into two sequential steps. First, we devise a lightweight, weakly supervised deep network to coarsely locate the semantically salient regions. Next, as a postprocessing refinement, we selectively fuse multiple off-the-shelf deep models on the semantically salient regions identified by the previous step to formulate a pixelwise saliency map. Compared with the state-of-the-art (SOTA) models that focus on learning the pixelwise saliency in single images using only perceptual clues, our method aims at investigating the object-level semantic ranks between multiple images, of which the methodology is more consistent with the human attention mechanism. Our method is simple yet effective, and it is the first attempt to consider salient object detection as mainly an object-level semantic reranking problem. Guangxiao Ma, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Salient Object Detection via Multiple Instance Joint Re-LearningabstractIn recent years deep neural networks have been widely applied to visual saliency detection tasks with remarkable detection performance improvements. As for the salient object detection in single image, the automatically computed convolutional features frequently demonstrate high discriminative power to distinguish salient foregrounds from its non-salient surroundings in most cases. Yet, the obstinate feature conflicts still persist, which naturally gives rise to the learning ambiguity, arriving at massive failure detections. To solve such problem, we propose to jointly re-learn common consistency of inter-image saliency and then use it to boost the detection performance. Its core rationale is to utilize the easy-to-detect cases to re-boost much harder ones. Compared with the conventional methods, which focus on their problem domain within the single image scope, our method attempts to utilize those beyond-scope information to facilitate the current salient object detection. To validate our new approach, we have conducted a comprehensive quantitative comparisons between our approach and 13 state-of-the-art methods over 5 publicly available benchmarks, and all the results suggest the advantage of our approach in terms of accuracy, reliability, and versatility. Guangxiao Ma, Chenglizhao Chen, Shuai Li 0001, Chong Peng 0001, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | Stage-wise Salient Object Detection in 360° Omnidirectional Image via Object-level Semantical Saliency RankingabstractThe 2D image based salient object detection (SOD) has been extensively explored, while the 360° omnidirectional image based SOD has received less research attention and there exist three major bottlenecks that are limiting its performance. Firstly, the currently available training data is insufficient for the training of 360° SOD deep model. Secondly, the visual distortions in 360° omnidirectional images usually result in large feature gap between 360° images and 2D images; consequently, the widely used stage-wise training-a widely-used solution to alleviate the training data shortage problem, becomes infeasible when conducing SOD in 360° omnidirectional images. Thirdly, the existing 360° SOD approach has followed a multi-task methodology that performs salient object localization and segmentation-like saliency refinement at the same time, being faced with extremely large problem domain, making the training data shortage dilemma even worse. To tackle all these issues, this paper divides the 360° SOD into a multi-staqe task, the key rationale of which is to decompose the original complex problem domain into sequential easy sub problems that only demand for small-scale training data. Meanwhile, we learn how to rank the "object-level semantical saliency", aiming to locate salient viewpoints and objects accurately. Specifically, to alleviate the training data shortage problem, we have released a novel dataset named 360-SSOD, containing 1,105 360° omnidirectional images with manually annotated object-level saliency ground truth, whose semantical distribution is more balanced than that of the existing dataset. Also, we have compared the proposed method with 13 SOTA methods, and all quantitative results have demonstrated the performance superiority. Guangxiao Ma, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2017 | Multi-faceted Visual Analysis on Tropical Cyclone
Cui Xie, Guangxiao Ma, Xiaotian Gao, Junyu Dong |
CDVE | 3 |