Chenyi Guo

dblp:135/4486 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
12since 2021 · last 2025
0000-0002-8222-5179ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Med-2E3: A 2D-Enhanced 3D Medical Multimodal Large Language Model
abstract
3D medical image analysis is essential for modern healthcare, yet traditional task-specific models are inadequate due to limited generalizability across diverse clinical scenarios. Multimodal large language models (MLLMs) offer a promising solution to these challenges. However, existing MLLMs have limitations in fully leveraging the rich, hierarchical information embedded in 3D medical images. Inspired by clinical practice, where radiologists focus on both 3D spatial structure and 2D planar content, we propose Med-2E3, a 3D medical MLLM that integrates a dual 3D-2D encoder architecture. To aggregate 2D features effectively, we design a Text-Guided Inter-Slice (TG-IS) scoring module, which scores the attention of each 2D slice based on slice contents and task instructions. To the best of our knowledge, Med-2E3 is the first MLLM to integrate both 3D and 2D features for 3D medical image analysis. Experiments on large-scale, open-source 3D medical multimodal datasets demonstrate that TG- IS exhibits task-specific attention distribution and sig-nificantly outperforms current state-of-the-art models. The code is available at: https://github.com/MSIIPlMed-2E3
Yiming Shi, Chenyi Guo, Miao Li 0003, Ji Wu 0002
BIBM5
2025 SSSL-HAR: Synthetic-Data-Driven Self-Supervised Learning for flexible IMU-Based Human Activity Recognition
abstract
The scarcity of labeled training data have significantly hindered the deployment of Inertial Measurement Unit-based Human Activity Recognition (IMU-HAR) in real-world scenarios. To address this limitation, we pro-pose SSSL-HAR (Synthetic-data-driven Self-Supervised-Learning HAR), a novel framework that leverages large amount synthetic IMU data for self-supervised pre-training, followed by fine-tuning with minimal real-world data. This approach mitigates the reliance on large-scale real IMU data collection compared to the traditional self-supervised-learning frameworks and also bypasses the need for labor-intensive annotation or cross-modality alignment compared to the traditional cross-modality methods. Furthermore, by adopting multi-view contrastive learning (MVCL) architectures, our method effectively captures the intrinsic relationships between synthetic sensor views, enabling robust generalization to diverse sensor placements and configurations. Experiments on the PAMAP2 dataset and a more complex custom fitness monitoring dataset demonstrate that SSSL-HAR achieves performance comparable to models pre-trained on real data, highlighting its potential for scalable and adaptive HAR deployment.
Timin Li, Zhuangzhuang Li, Ji Wu 0002, Yuepeng Chen, Xuefeng Feng, Chenyi Guo
IJCB9
2025 SEmgFormer: Muscle Synergy Channel Attention Enhanced Vision Transformer Network For sEMG Motion Recognition Using STFT Spectrogram
abstract
The recognition of motions based on surface electromyography (sEMG) has been extensively studied, yielding promising results from initial machine learning approaches to contemporary deep learning methods. However, most previous research has concentrated on the classification of movements from individual body parts, such as the widely used gesture dataset, NinaPro. Furthermore, much of work has been restricted to convolutional neural networks (CNNs) and their variants, without a thorough exploration of the synergistic effects of muscles from different body parts, often assigning equal weights to all muscles. This study presents the collection of electromyographic data from sixteen major muscles across the entire body, acquiring the MultiMotion-sEMG Dataset, which includes forty-three full-body movements from thirteen participants. According to the current knowledge, this is the first dataset designed to synchronize the capture of full-body surface electromyography (sEMG) signals. Based on this dataset, a novel sEMG recognition network, SEmgFormer, is proposed, which is augmented by a vision transformer (ViT). The short-time Fourier transform (STFT) is utilized to transform conventional time-domain sEMG signal recognition tasks into visual understanding tasks of time-frequency spectrograms, utilizing the Cutmix method for data augmentation. In addition, a novel Muscle Synergy Channel Attention (MS-CA) mechanism is introduced, improving the channel attention mechanism (CA). The results indicate that the proposed model surpasses other methods in performance, including CNN-based networks, achieving optimal accuracy. This validates the efficacy of the proposed ViT classifier using time-frequency spectrograms as input, enhancing the accuracy of sEMG recognition based on full-body signals, and paving new avenues for research in this field.
Zhuangzhuang Li, Chenyi Guo, Ji Wu 0002
IJCNN3
2025 Design of a Real-time Multi-individual Boxing Classification System Using Multiple Mode Sensing
abstract
In this paper, a real-time multi-individual boxing classification system, which is based on FPGA, using electromyography (EMG) signals and Inertial measurement units (IMUs) is presented. The proposed system is realized on the Xilinx Ultra96 single computer board, combining EMG signals and acceleration to classify different boxing actions. An improved attitude update algorithm based on extended Kalman filter (EKF) is used to calculate the change in roll angle of the IMU during a certain punch and a correlation-based method is then applied to classify between different types of punches. The presented system achieves a latency of about 21.7ms with an accuracy of 80.0% on a laptop with AMD R7-7735H, 16GM RAM and about 74.5ms with an accuracy of 68.3% on FPGA respectively in a multi-person boxing real-time monitoring scenario. The algorithm enables real-time boxing classification with high accuracy using only wearable sensors, which enhances the portability of the whole system and can be used in real-time scenarios.
Ziyao Zhao, Lingfeng Wu, Chenyi Guo, Milin Zhang 0001
ISCAS5
2025 DML-FitAR: A Deep Metric Learning Approach for IMU-Based Fitness Activity Recognition
abstract
This paper proposes DML-FitAR, a novel deep metric learning framework for IMU-based fitness action recognition, addressing critical challenges in real-world deployment. Unlike traditional transfer learning methods requiring fine-tuning for new action types, DML-FitAR achieves competitive accuracy on unseen actions through a retraining-free paradigm. Evaluated on a custom dataset (560+ fitness actions) and the MyoGYM dataset, DML-FitAR demonstrates superior performance over contrastive learning and visual backbone-based approaches, achieving cross-action-type recognition accuracy ranging from 80% to 90%. Besides that, the framework also exhibits robustness to sensor placement variations and noteworthy cross-dataset generalization.
Timin Li, Yuepeng Chen, Zhuangzhuang Li, Xuefeng Feng, Ji Wu 0002, Chenyi Guo
ICMR9
2025 sEMG-DGCN: Directed Graph Convolutional Network for Rehabilitation Action Difficulty Assessment Based on sEMG
abstract
Against the backdrop of an aging population and the high prevalence of chronic diseases, the demand for rehabilitation medical services has surged. However, traditional rehabilitation action difficulty assessment relies on expert experience, suffering from strong subjectivity and low reliability. Existing assessment methods struggle to cover the full-body kinematic characteristics, and existing models lack directed modeling of action difficulty relationships and fail to effectively capture muscle synergy. To address this, this study constructs a 16-channel full-body sEMG dataset, sEmgHuman-594, which includes 594 rehabilitation actions. This study proposes an sEMG-DGCN assessment method based on Directed Graph Convolutional Network (DGCN), which integrates 11-dimensional expert-annotated difficulty criteria to derive difficulty labels, employs directed graphs to model difficulty relationships between actions, and introduces an Anatomically Constrained Spatiotemporal Attention mechanism. Experimental results show that the sEMG-DGCN model achieves an accuracy of 93.61% in difficulty relationship classification, significantly outperforming comparative models. Ablation experiments verify the effectiveness of the attention mechanism, providing a new pathway for rehabilitation action difficulty assessment.
Zhuangzhuang Li, Xuefeng Feng, Chenyi Guo, Jian Ning
SMC5
2024 3D Human Pose Estimation via Non-causal Retentive Networks
Kaili Zheng, Feixiang Lu, Yihao Lv, Liangjun Zhang, Chenyi Guo, Ji Wu 0002
ECCV (33)5
2024 Dual Dynamic Attention Network for Flexible Job Scheduling with Reinforcement Learning
abstract
The flexible job shop problem (FJSP) is a classic combinatorial optimization problem that is strongly NP-hard. Recent studies have utilized deep reinforcement learning (DRL) methods for scheduling operations in FJSP problems, achieving results comparable to accurate methods such as OR tools. However, there are still limitations in extracting global representations of machines and operations. This paper proposes the Dual Dynamic Attention Network (DDAN), which addresses these limitations. The proposed method utilizes an in-channel dynamic attention mechanism to capture the global representation of machines and operations. This allows for accurate and efficient representation of complex dependencies between operations and machines, providing effective support for subsequent dispatching model. Assessments using synthetic datasets as well as public benchmarks corroborate the proposed approach’s superiority over traditional priority dispatching rules (PDRs) and state-of-the-art DRL algorithms. In certain cases, it even surpasses deterministic algorithms. Additionally, this method demonstrates superior performance and stronger generalization capabilities compared to current state-of-the-art DRL methods on large-scale FJSP problems that have not been previously encountered.
Yuepeng Chen, Ji Wu 0002, Chenyi Guo
IJCNN4
2024 A survey of label-noise deep learning for medical image analysis
Jialin Shi, Kailai Zhang, Chenyi Guo, Youquan Yang, Yali Xu, Ji Wu 0002
Medical Image Anal.3
2023 An iterative sinogram metal artifact reducdion based on UNet
abstract
In the practice of dentistry, oral dental CT images are frequently used to assist doctors in diagnosis. Filtered back projection (FBP) technique is widely employed in practice for the reconstruction of CT images obtained from X-ray calculations. However, when metal objects occur in a patient’s oral cavity, the CT images would show density discontinuities due to the metals’ “X-ray absorption coefficient is much larger than human tissues. When the FBP algorithm is applied to CT images with metals, severe metal artifacts would be obtained, which significantly reconstructed images. Therefore, metal artifact reduction (MAR) work is becoming an important problem in dentistry image processing. In this paper, we propose a novel iterative sinogram metal artifact reduction model (IS-MARM) to solve the problem. Inspired by the Diffusion model, we propose a new method to reduce metal artifacts and interpolate new data in sinogram of dentistry images iteratively. This approach reduces the difficulty of model learning and achieves good results. Secondly, we proposed a new simple method of iterative data generating to simulate real-world metals in CT sinogram images. Finally, we have demonstrated the effectiveness of our method through experiments on dental CT MAR work.
Zichong An, Xuemei Zhu, Xiangling Fu, Junqi Ma 0005, Chenyi Guo
IEEE Big Data5
2023 Image Captioning for Nantong Blue Calico Through Stacked Local-Global Channel Attention Network
Chenyi Guo
ICANN (2)1
2022 MPF-net: An effective framework for automated cobb angle estimation
Kailai Zhang, Nanfang Xu, Chenyi Guo, Ji Wu 0002
Medical Image Anal.3
2020 Multi-modal Feature Attention for Cervical Lymph Node Segmentation in Ultrasound and Doppler Images
Xiangling Fu, Mengke Zhang, Chenyi Guo, Ji Wu 0002
ICONIP (4)5