Xiaohu Sun

dblp:169/7244 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2025 SACM: spatial attributes of complex movements in multi-object tracking
Hongjun Li 0003, Jiaxin Li 0003, Xiaohu Sun
Multim. Tools Appl.3
2024 Clustering Multimodal Ensemble Learning for Predicting Gastric Cancer Neoadjuvant Chemotherapy Efficacy
abstract
Accurately predicting the response to neoadjuvant chemotherapy (NCT) is crucial for gastric cancer treatment planning. Multimodal diagnostic methods enhance prediction accuracy by integrating images and clinical features, but they often overlook intra-class differences, causing feature entanglement, and single models struggle to capture dataset diversity. To address these issues, we propose Clustering Multimodal Ensemble Learning (CMEL) framework, which aligns and fuses image features from different domains with clinical features and performs decision fusion using an ensemble model. Specifically, we design the clustering process using Siamese network and memory bank to partition images in feature space and introduce contrastive clustering loss to enhance clustering properties and reduce feature entanglement across domains. We design Hierarchical Feature Alignment Fusion (HFAF) module to generate multi-level image features at different depths and then separately align and fuse them with clinical features, providing diverse features to cover variations in the dataset. We design Dynamic Multi-Classifier Decision Fusion (DMDF) module to predict on fused features using a gating mechanism supervised by pseudo-labels for dynamic weights, enhancing prediction across different domains. On our collected the Gastric Cancer Chemotherapy Response (GCCR) dataset, CMEL achieves an AUC of 76.09%, accuracy of 72.00%, PPV of 86.02%. These results demonstrate the feasibility of joint image and clinical data in predicting NCT response and provide valuable insights for gastric cancer treatment. The code is available at https://github.com/WHX0259/CMEL.
Jianning Chi, Huixuan Wu, Yujin Shi, Zelan Li, Xiaohu Sun, Zitian Zhang, Yuehua Gong
BIBM5
2024 MTM-net: a multidimensional two-stage memory-guided network for vedio abnormal detection
Hongjun Li 0003, Xiaohu Sun
Multim. Tools Appl.3
2024 Appearance-motion heterogeneous networks for video anomaly detection
Hongjun Li 0003, Xiaohu Sun
Multim. Tools Appl.2
2024 Channel based approach via faster dual prediction network for video anomaly detection
Hongjun Li 0003, Xulin Shen, Xiaohu Sun
Multim. Tools Appl.3
2024 Inversion of Aerosol Optical Depth: Incorporating Multimodel Approach
abstract
Atmospheric aerosols originate from diverse sources and exert a notable influence on the radiation budget, atmospheric environment, and human health. However, current aerosol inversion models still have limitations in dealing with multiple types of variables and intricate scenarios, for which a two-stage hybrid model named convolutional neural network-random forest (CNNRF) is proposed in this study. Convolution is employed to extract continuous spectral signals. This study takes into account the synergistic impact of spatiotemporal, meteorological, and surface information. Ensemble learning is then applied to adeptly handle diverse-independent input variables. In this article, eight regions were selected globally for modeling and testing based on different scenario types. Additionally, an aerosol hotspot region (India) was chosen for independent experiments. Accuracy was validated at the site scale using a 10-fold cross-validation (10-CV) approach and cross-comparison with the MCD19A2 product, and convolutional neural network (CNN) and random forest (RF) model results. The sample-based CV of the CNNRF model demonstrates high and consistent accuracy, with a Pearson correlation coefficient ($R$) value of 0.958, mean absolute error (MAE) of 0.048, and a within expected error (EE) envelope of 87.15%. For the Indian region, the MAE and EE are 0.05 and 95.8%, respectively. In summary, the proposed hybrid model demonstrates robust generalization capabilities, enabling accurate and stable aerosols estimation on a global scale.
Xiaohu Sun, Lin Sun 0001, Xiaole Fan
IEEE Trans. Geosci. Remote. Sens.1
2023 Multi-memory video anomaly detection based on scene object distribution
Hongjun Li 0003, Xiaohu Sun
Multim. Tools Appl.3
2023 MPAT: multi-path attention temporal method for video anomaly detection
Hongjun Li 0003, Xiaohu Sun, Xulin Shen, Zhengguang Xie
Multim. Tools Appl.2
2023 Video anomaly detection based on scene classification
Hongjun Li 0003, Xulin Shen, Xiaohu Sun
Multim. Tools Appl.3
2016 A partial least squares based ranker for fast and accurate age estimation
abstract
Facial age estimation is challenging due to complex dynamics in aging process, which render metric regression methods unfavorable. Rankers show better performance by exploiting the ordinal nature of ages. The difficulty of designing a ranker is that each binary classifier of a ranker has to be trained using highly unbalanced positive and negative data. This paper proposes a partial least squares based ranker (PLS-Ranker), which fully maintains the advantages of PLS and greatly boosts its performance on the ordinal problem. In PLS-Ranker, an adaptive threshold learning strategy is proposed to boost each of the binary classifiers learned from highly unbalanced data. Previous ranking approaches such as CS-OHRank suffer from heavy computations because dozens of binary classifiers are trained separately. However, in PLS-Ranker, they are jointly learned. Additionally, PLS-Ranker simultaneously reduces feature dimensions and ranks in high speed even for high-dimensional features. Experimental results on the age estimation problem show that PLS-Ranker outperforms the state-of-the-art methods in terms of both accuracy and speed. PLS-Ranker also achieves state-of-the-art performance on the multi-source cross-race-and-gender age estimation problem, which further demonstrates its robustness.
Xiaohu Sun
ICASSP2
2016 Linear canonical correlation analysis based ranking approach for facial age estimation
abstract
Facial age estimation is an important and challenging problem in computer vision and pattern recognition. Linear canonical correlation analysis (CCA) has been widely applied owing to low complexity, small and fixed amount of model parameters and good scalability. However, linear CCA based regression gets lower accuracy than its kernel version on the age estimation problem. The inexactness of metric distance information carried by age labels increases the complexity of using regression-based methods to estimate age. Hence, we propose a linear 2-norm regularized LS-CCA based ranking approach only exploiting the ordinal information carried by age labels. It gains the advantages of both linear 2-norm regularized LS-CCA and the ranking approach. Our method achieves competitive accuracy with the state-of-the-arts with sharply lower time cost and less amount of model parameters, which makes it appealing for real-time, large-scale applications or embedded systems. Additionally, experimental results on multi-source cross-population age estimation problem demonstrate that it is more robust against race and gender variations than the state-of-the-arts.
Xiaohu Sun
ICIP2
2016 Violence detection using Oriented VIolent Flows
Yuan Gao 0008, Hong Liu 0008, Xiaohu Sun, Can Wang 0006
Image Vis. Comput.3