VLDB 2026 Research / reviewers in the wild / expert
Xiaole Ma
dblp:211/5042
· DBLP profile ↗
16ranked-venue papers
3as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SemDNet: Semantic-guided despeckling network for SAR images
Fuyu Bo, Yi Jin 0001, Xiaole Ma, Yi-Gang Cen, Shaohai Hu, Yidong Li |
Expert Syst. Appl. | 3 |
| 2025 | CroCaps: A CLIP-assisted cross-domain video captioner
Wanru Xu, Yenan Xu, Zhenjiang Miao, Yi-Gang Cen, Xiaole Ma |
Expert Syst. Appl. | 6 |
| 2025 | ACFNet: An adaptive cross-fusion network for infrared and visible image fusion
Xiaoxuan Chen, Shuwen Xu 0002, Shaohai Hu, Xiaole Ma |
Pattern Recognit. | 4 |
| 2025 | UIFD: A Unified Interactive Network for Image Fusion and Traffic Object Detection Under Low-Light ConditionsabstractUnder low-light conditions, perceptual fusion techniques, such as image fusion, can mitigate the inherent limitations of source images, thereby enhancing the safety performance of automated driving. However, these methods typically perform fusion and detection tasks separately, making it difficult to improve detection precision, as the fused images often lack the rich target information necessary for effective low-light traffic detection. In this paper, an end-to-end unified interactive network is proposed, which eliminates the problem of structural inconsistency between fusion and detection tasks by constructing interaction relationships. In this work, a dual-branch coupled feature extraction module is proposed. Different from other dual-branch methods, this module couples weak features from different modalities to enhance features. In addition, an interactive fusion module is proposed to achieve mutual enhancement between fusion and detection tasks while adaptively weighting the fusion of infrared and visible features. Extensive experimental results on traffic scene datasets, such as LLVIP and FMB, demonstrate that the proposed unified network not only achieves excellent fusion results but also significantly improves traffic object detection precision in low-light environments. Xiaoxuan Chen, Shuwen Xu 0002, Shaohai Hu, Xiaole Ma |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Cross-Subject EEG Emotion Recognition Based on Interconnected Dynamic Domain AdaptationabstractElectroencephalogram (EEG) is widely utilized in emotion recognition owing to its unique advantages. To achieve more optimal cross-subject emotion recognition, a cross subject emotion recognition method based on interconnection dynamic domain adaptation (IDDA) is proposed. In IDDA, dynamic graph convolution (DGC) is employed to dynamically learn the intrinsic relationships between different EEG channels and to extract domain invariant features. And dynamic domain adaptation (DDA) is employed to align the source domain and target domain, at the same time the emotional sub-domains is aligned, achieving more optimal cross subject emotion recognition. To select suitable subjects as the source domain, a multi-source selection algorithm is incorporated before dynamic adaptive computation reducing migration noise and achieving interconnection between DGC and DDA. IDDA enhances the emotion discrimination ability of domain invariant features, thereby improving the accuracy of cross-subject EEG emotion recognition. This method achieves classification results of 85.75% and 72.36% in cross subject experiments on SEED and SEED-IV. Yanling An, Shaohai Hu, Shuaiqi Liu 0001, Zeyao Wang, Xiaole Ma |
ICASSP | 6 |
| 2024 | Probabilistic Distillation Transformer: Modelling Uncertainties for Visual Abductive ReasoningabstractVisual abduction reasoning aims to find the most plausible explanation for incomplete observations, and suffers from inherent uncertainties and ambiguities, which mainly stem from the latent causal relations, incomplete observations, and the reasoning itself. To address this, we propose a probabilistic model named Uncertainty-Guided Probabilistic Distillation Transformer (UPD-Trans) to model uncertainties for Visual Abductive Reasoning. In order to better discover the correct cause-effect chain, we model all the potential causal relations into a unified reasoning framework, thus both the direct relations and latent relations are considered. In order to reduce the effect of the stochasticity and uncertainty for reasoning: 1) we extend the deterministic Transformer to a probabilistic Transformer by considering those uncertain factors as Gaussian random variables and explicitly modeling their distribution; 2) we introduce a distillation mechanism between the posterior branch with complete observations and the prior branch with incomplete observations to transfer posterior knowledge. Evaluation results on the benchmark datasets, consistently demonstrate the commendable performance of our UPD-Trans, with significant improvements after latent relation modeling and uncertainty modeling. Wanru Xu, Zhenjiang Miao, Yi-Gang Cen, Xiaole Ma |
ACM Multimedia | 6 |
| 2024 | MGFA : A multi-scale global feature autoencoder to fuse infrared and visible images
Xiaoxuan Chen, Shuwen Xu 0002, Shaohai Hu, Xiaole Ma |
Signal Process. Image Commun. | 4 |
| 2024 | SAR Image Speckle Reduction Based on Nuclear Norm Minus Frobenius Norm RegularizationabstractSynthetic aperture radar (SAR) is a powerful imaging system with all-day and all-weather capabilities, making it suitable for a wide range of applications. However, SAR images often suffer from coherent speckle noise, which degrades image quality and hampers subsequent analysis and interpretation. Recently, methods based on the Fisher-Tippett (FT) distribution and nonlocal low-rank (NLR) techniques have shown great potential in SAR despeckling. Building upon these methods, this article proposes a novel SAR image despeckling method named SAR nuclear norm minus Frobenius norm (SAR-NNFN). This method effectively restores clean images using singular value shrinkage and allows for adaptive shrinkage without the need for additional weighting parameters. SAR-NNFN utilizes NNFN to achieve rank relaxation, resulting in a more robust low-rank solution for speckle reduction. The proposed model comprises two components: a data fidelity term that captures the statistical characteristics of SAR images using the FT distribution in the logarithmic domain, and an NNFN regularization term that enhances low-rank approximations. The optimization problem associated with SAR-NNFN is solved using the alternating direction method of multipliers (ADMM) algorithm. Extensive experiments conducted on both simulated and real SAR images demonstrate that SAR-NNFN can not only adequately suppress speckle noise but also preserve fine textures. Fuyu Bo, Xiaole Ma, Yi-Gang Cen, Shaohai Hu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Image fusion based on discrete Chebyshev moments
Xiaoxuan Chen, Shuwen Xu 0002, Shaohai Hu, Xiaole Ma |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | MRI Image Fusion Based on Optimized Dictionary Learning and Binary Map Refining in Gradient Domain
Qiu Hu, Shaohai Hu, Xiaole Ma, Fengzhen Zhang |
Multim. Tools Appl. | 3 |
| 2022 | Dynamics of Blink and Non-Blink Cyclicity for Affective Assessment: A Case Study for Stress IdentificationabstractPrevious studies have shown that eye activities, including blinks, can indicate the psychological state of an individual. However, almost all previous studies analyzing blinks merely concentrated on traditional descriptive statistics, which are unable to reflect their dynamic processes. Furthermore, the states of non-blink (opening the eyes) and blink alternate with each other, forming a physiological cycle. If we only investigate blinks alone, it may be inadequate to describe how blinking works. Therefore, we attempted to recognize the affective state (“relaxation” vs. “stress”) of an individual through the dynamics of blink and non-blink cyclicity (BNBC), as one example, to illustrate this method. First, the “Stroop Test” was employed for emotion elicitation. Then, features were extracted from a categorical time series (0: non-blink; 1: blink), which was recorded by the eye-tracking system. Finally, the areas under the receiver operating characteristic curve (AUC) values were obtained via eight commonly used classifiers. The results show that, compared with the traditional approaches for blink analysis, BNBC exhibits more compelling proficiency to detect stress. In summation, BNBC can be considered a new type of psychophysiological measure, which could be widely applied in psychology, medicine, and engineering. Peng Ren 0002, Armando Barreto, Xiaole Ma, Shengnan Liu, Ying Wang 0061, Yeyun Dong, Dezhong Yao 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2021 | Multi-focus image fusion based on multi-scale sparse representation
Xiaole Ma, Shaohai Hu |
J. Vis. Commun. Image Represent. | 1 |
| 2021 | SAR Speckle Removal Using Hybrid Frequency ModulationsabstractSynthetic aperture radar (SAR) images often interfere with speckle artifacts that have a great impact on subsequent processing and analysis operations. To remove speckle artifacts, this article introduces a hybrid denoising approach by using a convolutional neural network (CNN) and consistent cycle spinning (CCS) in the nonsubsample shearlet transform (NSST) domain. First, we apply NSST to a noisy SAR image to gain low- and high-frequency coefficients. Second, we adopt a learned deep CNN model to eliminate the speckle noise in the low-frequency coefficients, which retains more contour information. Third, we employ CCS to enhance the high-frequency coefficients, which preserves more details of the original SAR image. Finally, we obtain the denoised image by using inverse NSST applied to the denoised coefficients. Compared with state-of-the-art algorithms, the results of the experiment indicate that our method not only achieves better speckle removal performance but also maintains more detailed information retention. Shuaiqi Liu 0001, Lele Gao, Miaohui Wang, Xiaole Ma, Yudong Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2019 | IoT Structured Long-Term Wearable Social Sensing for Mental WellbeingabstractLong-term wellbeing monitoring is an underlying theme for evaluating health status by collecting physiological signs through behavioral traits. In alignment with Internet of Things (IoT), nonintrusive and trustworthy wearable social sensing technology holds a potential way for researchers to find and establish the interrelationships between unobtrusive social cues and physical mental health. This paper implements an IoT structured wearable social sensing platform with the integration of privacy audio feature, behavior monitoring, and environment sensing in a naturalistic environment. Particularly, four privacy protected audio-wellbeing features are embedded into the platform to automatically evaluate speech information without preserving raw audio data. Four weeks of long-term monitoring experimental studies have been conducted. A series of well-being questionnaires in conjunction with a group of students are engaged to objectively investigate the relationships between physical and mental health by utilizing the feature fusion strategy from speech, behavioral activities, and ambient factors. Sihao Yang, Bin Gao 0003, Long Jiang, Jikun Jin, Zhao Gao, Xiaole Ma, Wai Lok Woo |
IEEE Internet Things J. | 6 |
| 2019 | Multi-focus image fusion based on joint sparse representation and optimum theory
Xiaole Ma, Shaohai Hu, Shuaiqi Liu 0001, Shuwen Xu 0002 |
Signal Process. Image Commun. | 1 |
| 2018 | SAR image edge detection via sparse representation
Xiaole Ma, Shuaiqi Liu 0001, Shaohai Hu, Peng Geng, Jie Zhao 0008 |
Soft Comput. | 1 |