VLDB 2026 Research / reviewers in the wild / expert
Chenyu Zhou 0005
dblp:215/8318-5
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2025
0009-0006-9220-7501ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MST-HA: Multi-Modal Signal Fusion with Bayesian Optimization for Robust Industrial Robot Joint Health AssessmentabstractThis paper presents a novel multi-modal deep learning framework for industrial robot joint health assessment and prediction, leveraging non-invasive signal fusion and Bayesian optimization. The proposed method addresses the challenges of comprehensive joint state monitoring in complex industrial environments without disrupting normal operations. We integrate Hall-effect current sensors, external accelerometers, and joint encoders to collect multi-modal data, including motor currents, vibrations, and kinematic information. A novel Adaptive Multi-Receptive Field Attention Network (AMRFAN) is employed to extract features from each modality, while a synchrosqueezing transform (SST) is utilized to capture time-frequency characteristics. An attention mechanism dynamically adjusts the weights of different modalities, and a bidirectional long short-term memory (BiLSTM) network models the temporal dependencies in the fused features. To enhance model performance and generalization, we implement a Bayesian optimization framework for hyperparameter tuning. Furthermore, we incorporate a Bayesian neural network to quantify prediction uncertainties, providing reliability metrics for decision-making processes. Experimental results on a four-axis industrial robot demonstrate that our framework achieves a 98.3% accuracy in joint health state classification and a mean absolute error of 4.2 in remaining useful life prediction, outperforming state-of-the-art single-modality methods. The proposed approach offers a robust, adaptable solution for real-time health monitoring and predictive maintenance of industrial robot joints, potentially improving manufacturing efficiency and reliability. Haoyu Wang 0011, Zilong Yin, Xiyue Yan, Chenyu Zhou 0005, Guangmeng Xue, Haichao Xu |
ICASSP | 5 |
| 2025 | Neural Synchronization and Analysis-Grounded Computational Model for Fine-Grained Sentiment Understanding
Zilong Yin, Chenyu Zhou 0005, Haoyu Wang 0011 |
ICONIP (1) | 5 |
| 2025 | Locate, enhance and fuse: a progressively optimized network for camouflaged object detection
Tianchi Qiu, Zhe Li 0030, Shaokang Ma, Kangwei Liu 0001, Chenyu Zhou 0005 |
Multim. Tools Appl. | 7 |
| 2024 | Prompt Fusion Interaction Transformer For Aspect-Based Multimodal Sentiment AnalysisabstractAspect-based multimodal sentiment analysis (ABMSA) is a recent and popular research area that uses multiple modalities like text and images to determine the sentiment orientation of opinion entities. The main challenge in multimodal sentiment analysis is dynamically modeling each modality and effectively fusing information across different modalities. Existing methods have not considered fine-grained texture features in images, and direct fusion introduces irrelevant and ineffective features unrelated to sentiment information. To address these limitations, we propose a new model for multimodal sentiment analysis, the multimodal prompt fusion interaction Transformer (MPFIT). We designed two key components: 1) the image assist module (IAM), which leverages the self-attention mechanism and statistical pooling to obtain weighted mean and standard deviation vectors, enabling the model to focus on image texture information and reduce image noise. 2) Multimodal prompt fusion (MPF) restricts multimodal fusion to interactions between small prompt tokens that capture vital information from different modalities, allowing the model to focus on features that are more relevant to sentiment information. Experimental results show that our model outperforms baseline models on two publicly available datasets, Twitter-2015 and Twitter-2017. We conducted ablation experiments to evaluate the impact of our key components. Zhe Li 0030, Chenyu Zhou 0005 |
ICME | 4 |
| 2024 | Multimodal Rumor Detection via Multimodal Prompt LearningabstractPre-trained vision-language (V-L) models exhibit significant generalization capabilities in detecting rumors. However, their reliance on single-modality prompts—either language or vision—limits their flexibility for dynamic adjustments in both representation spaces during rumor detection. To address these limitations, we propose a multimodal rumor detection framework that uses prompt learning in both the vision and language domains to align their representations better. Inspired by recent advances in efficiently tuning large language models, we introduce a set of trainable parameters in the input space, keeping the model backbone frozen. Additionally, we use distinct prompts at various early stages, which helps progressively model the relationships between features, enhancing comprehensive context learning. Extensive experiments with two real-world multimodal datasets demonstrate our framework’s superior ability to distinguish rumors from facts. Zhe Li 0030, Chenyu Zhou 0005, Jiabao Sheng |
IJCNN | 4 |
| 2024 | Camouflaged Object Detection using Multi-Level Feature Cross-FusionabstractCamouflaged object detection (COD) aims to segment objects that closely resemble their surroundings. Accurately recognizing camouflaged objects in these complex environments is challenging due to factors such as low illumination, object occlusion, small size, and similar background. To this end, we propose a novel network for camouflaged object detection, the Multi-Level Feature Cross-Fusion Network (MFCF-Net). This framework aims to learn and utilize background features at different scales through cross-fusion, thereby improving detection accuracy. The core of our approach is to use a modified version of the Pyramid Vision Transformer (PVTv2) as a backbone network to effectively capture contextual information at different scales. Then, we design the Multi-scale Feature Enhancement (MFE) module to optimize features at each scale. In addition, to enhance the model’s ability to recognize camouflaged objects in complex contexts, we cross-fused these enhanced features. Finally, we designed the Balanced Multilevel Feature Cross-Fusion (BMFCF) module. This module improves the accuracy of camouflaged object detection by deeply learning and effectively utilizing contextual feature information and cross-fusing these multi-scale features. Extensive research results show that our MFCF-Net significantly outperforms 18 leading methods on four widely used standard datasets. Tianchi Qiu, Chenyu Zhou 0005, Kangwei Liu 0001 |
IJCNN | 4 |
| 2024 | Enhancing Cross-Modal Alignment in Multimodal Sentiment Analysis via Prompt Learning
Zhe Li 0030, Chenyu Zhou 0005 |
PRCV (5) | 4 |
| 2024 | Enhancing Multimodal Rumor Detection with Statistical Image Features and Modal Alignment via Contrastive Learning
Chenyu Zhou 0005, Zhe Li 0030, Jiabao Sheng, Haoyu Wang 0011 |
PRICAI (3) | 1 |
| 2024 | Federated semi-supervised representation augmentation with cross-institutional knowledge transfer for healthcare collaboration
Zilong Yin, Haoyu Wang 0011, Hangling Sun, Anji Li 0002, Chenyu Zhou 0005 |
Knowl. Based Syst. | 8 |
| 2023 | Emphasizing Boundary-Positioning and Leveraging Multi-scale Feature Fusion for Camouflaged Object Detection
Zhe Li 0030, Chenyu Zhou 0005, Tianchi Qiu |
PRCV (12) | 5 |
| 2023 | Boundary Guided Feature Fusion Network for Camouflaged Object Detection
Tianchi Qiu, Kangwei Liu 0001, Chenyu Zhou 0005 |
PRCV (9) | 6 |
| 2023 | Multimodal Rumor Detection by Using Additive Angular Margin with Class-Aware Attention for Hard Samples
Chenyu Zhou 0005, Zhe Li 0030 |
PRCV (1) | 1 |