Rui Wang 0034

dblp:06/2293-34 · DBLP profile ↗
← Back
36ranked-venue papers
10as first author
30since 2021 · last 2026
0000-0002-7974-9510ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Computer networks · 5 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A novel cognitive diagnostic network with color and spatial cues for skin disease recognition
Ming Ju, Rui Wang 0034, Chunhua Qian
Eng. Appl. Artif. Intell.4
2026 PBDD: A Prompt-Based Learning Approach for Few-Shot Social Media Depression Detection
abstract
Automated detection of depressive moods from social media holds great promise for early mental health intervention, yet existing multimodal approaches typically require large quantities of annotated data and extensive feature engineering, impeding their deployment in real‐world settings where labels are scarce. To address this challenge, we propose prompt‐based depression detection (PBDD), a novel prompt‐based few‐shot learning framework that leverages frozen pretrained language and vision models to identify depression indicators from paired text‐image posts without fine‐tuning. Our method begins with rigorous data cleaning and sampling to construct a high‐quality few‐shot dataset, then encodes text via a masked language model and images via a self‐supervised rotation‐prediction task to capture deep semantic cues. Multimodal representations are seamlessly fused into a unified prompt template containing a [MASK] token, enabling the pre‐trained model to infer depressive states by language completion. Extensive experiments on both large‐scale and 1 % few‐shot subsets demonstrate that PBDD consistently outperforms state‐of‐the‐art baselines, achieving significant gains in accuracy and Macro‐F1. These results validate the effectiveness and scalability of our framework for depression detection under severe label scarcity, offering a practical solution for real‐time mental health monitoring in social media environments.
Rui Wang 0034, Heyang Feng, Erik Cambria, Kaize Shi, Xiaohan Yu 0001, Xuhui Fan 0001, Xianxun Zhu
IEEE Trans. Comput. Soc. Syst.1
2026 CIME: Contextual Interaction-Based Multimodal Emotion Analysis With Enhanced Semantic Information
abstract
Multimodal emotion analysis is pivotal in decoding complex human affect by integrating diverse data sources such as text, audio, and visual signals. In this article, we introduce contextual interaction-based multimodal emotion analysis with enhanced semantic information (CIME), a novel spatio-temporal interaction network that significantly improves emotion recognition accuracy and robustness. CIME employs a text-centric cross-modal attention mechanism to refine semantic representations, while simultaneously leveraging a graph convolutional network to model contextual dialog information by capturing both intraspeaker and interspeaker relationships. This dual approach enables the effective fusion of modality-specific cues and the mining of latent emotional associations across modalities. Extensive experiments conducted on benchmark datasets—including IEMOCAP and MOSEI—demonstrate that CIME consistently outperforms existing state-of-the-art methods in terms of overall classification accuracy and weighted F1-scores. Furthermore, detailed ablation studies underscore the critical contributions of both the cross-modal attention and graph-based contextual modules.
Rui Wang 0034, Chaopeng Guo, Mohammad Shabaz, Imad Rida, Erik Cambria, Xianxun Zhu
IEEE Trans. Comput. Soc. Syst.1
2026 A 3D Geometry-Based Stochastic Model for mmWave Massive MIMO Subway Train Channels in Tunnel Scenarios
abstract
With the rapid evolution of wireless communication technology, the demand for reliable communication within subway tunnels continues to grow. Future subway communication systems will necessitate higher data transmission rates, reduced latency, and enhanced reliability. Therefore, this paper presents a three-dimensional (3D) geometry-based stochastic channel model designed for millimeter-wave (mmWave) massive multiple-input multiple-output (MIMO) channels in tunnel environments. This work aims to address the limitations of previous studies by providing a more accurate and comprehensive model tailored to the unique propagation characteristics of tunnel scenarios. The material property parameters of tunnels in the ray-tracing (RT) model are calibrated using tunnel channel measurement data, thereby ensuring a strong alignment between the simulation results and the measurement outcomes. The calibrated RT model is subsequently utilized to generate extensive channel data, enabling the tracking and clustering of multipath components (MPCs). Cluster parameters are then derived for various cluster types. Finally, the proposed model’s validity is confirmed by comparing its channel characteristics with actual channel measurements in a new scenario. Simulation results show that the proposed model accurately represents the subway train channel characteristics in tunnel scenarios, offering essential theoretical support and experimental data for the design and optimization of subway wireless communication systems.
Yichen Feng, Rui Wang 0034, Xiaoyong Wang, Asad Saleem, Yicha Zhang, Guoxin Zheng
IEEE Trans. Wirel. Commun.2
2025 CEDT2M: text-driven human motion generation via cross-modal mixture of encoder-decoder
XiangYang Wang, Rui Wang 0034
Neural Comput. Appl.3
2025 Participatory Budget Project Selection Considering District Fairness: Two-Stage Large-Scale Multiattribute Group Decision Making Method
abstract
Participatory budgeting (PB) allows citizens with different backgrounds to participate directly in the decision-making process of budgeting and fund allocation. It can be considered as a multiattribute decision-making (MADM) process, which enhances the fairness, transparency, and effectiveness of decision-making. However, with the increasing number of participants and the complexity of the decision-making environment, how to effectively manage and optimize the selection of projects for large-scale PB with fairness has become a major challenge. To address this issue, this article proposes a two-stage MADM consensus method for large-scale participatory budget project selection problems. In the first stage, the ISODATA algorithm segments large participatory districts into subclusters, followed by an iterative algorithm based on group contribution theory to generate subcluster decision matrices. In the second stage, each subcluster is treated as a new participant, with district fairness considered to refine decision-making power through iterative adjustments, resulting in more manageable decisions. Finally, the iterative consensus-reaching algorithm is applied again to get the final group decision result. The applicability of the proposed methodology is verified through a case study in the Poland participant city project selection budget and comparative analysis. The result demonstrates that the methodology can integrate the preferences of different districts, facilitate budget consensus reaching, and improve the misrepresentations and fairness of decision-making.
Yannan Gou, Ying Ji 0001, Zeshui Xu, Rui Wang 0034
IEEE Trans. Comput. Soc. Syst.5
2025 A Multifactor Deep Forest Regression Stepwise Downscaling Framework for High-Resolution XCO2 in China
abstract
Satellite remote sensing, with its wide coverage, long time series, and high revisit frequency, has become an indispensable tool for CO2monitoring. However, mainstream XCO2datasets derived from satellite observations typically feature low spatial resolution, which limits the practical applications of these valuable satellite data. To overcome this limitation, we propose a multifactor Deep Forest Regression Stepwise Downscaling (DFRSD) Framework using the existing global monthly and Gap-Free column-averaged dry-air mole fraction of CO₂ (GF-XCO2) resolution of 0.1°. This approach introduces an intermediate resolution level (0.05°) between the initial resolution (0.1°) and target resolution (1km). Specifically, column-averaged dry-air mole fraction of CO₂ (XCO2) at resolution of 0.1° is first downscaled to that of 0.05°, and then further refined to 1 km resolution. The Deep Forest Regression (DFR) model is used to establish relationships between XCO2 and auxiliary variables at each downscaling step, including near-surface air pollutant, Normalized Difference Vegetation Index (NDVI), Temperature (TMP), and Precipitation (PRE). We successfully enhance the XCO2spatial resolution of 0.1° to that of 1 km, generating high-resolution monthly spatial resolution products of 1 km for the period from January 2016 to December 2020. To validate the accuracy of the reconstructed high-resolution XCO2, we compare it with field measurements from ground monitoring stations. The results demonstrate a strong agreement between the downscaled XCO2and field observations, with R² of 0.912 and RMSE of 1.10ppm, highlighting its reliability and precision. The results demonstrate the effectiveness of using SRF and Multifactor methods for XCO₂ downscaling. The proposed method not only significantly enhances spatial resolution but also preserves spatial distribution integrity, which satisfies the precision requirements of various research applications.
Ming Ju, Shiyan Sun, Rui Wang 0034, Simon Fong 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 A Geometric Algebra-Based Model for Enhanced Hyperspectral Anomaly Detection
abstract
Hyperspectral Anomaly Detection (HAD) is of great significance in remote sensing by identifying spectrally distinct pixels as anomalies without prior information, while the majority of pixels with similar spectral characteristics are classified as background. While existing approaches combine the strengths of deep learning in feature extraction and background suppression with the effectiveness of low-rank representation in background modeling, they fail to fully exploit rich spatial information and long-range spectral dependencies due to the limitations of conventional convolutional operations. To address this issue, we propose a Geometric Algebra (GA)-based deep neural network for HAD, named GA-HAD. The network constructs an encoder-decoder architecture using GA convolutional layers to extract spectral-spatial features, capturing multi-dimensional spatial information while preserving the intricate spectral characteristics of hyperspectral images. A specialized reconstruction error is designed to train the GA feature extraction network with high efficiency and accuracy. The encoder features are integrated with the low-rank representation algorithm and a Gaussian mixture model-based dictionary to improve background modeling and robustness. Finally, the detection maps output by the low-rank detection module are refined through an edge-preserving filter to produce the final detection result. Extensive experiments on four datasets demonstrate the superiority of our proposed GA-HAD framework in overall detection effect compared to thirteen representative baselines, implying our potential application in the field of HAD.
Rui Wang 0034, Ming Ju, Wei Xiang 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 A Geometric algebra-enhanced network for skin lesion detection with diagnostic prior
Ming Ju, Xianxun Zhu, Chunhua Qian, Rui Wang 0034
J. Supercomput.7
2025 AplusN: Progressively Integrating Attention and Normalization in Wavelet Domain for Pose Transfer
abstract
Pose-guided person image generation aims to synthesize images of human in various poses, often encountering issues such as occlusions and texture transfers. Previous methods have utilized attention mechanisms, flow field, normalization techniques, and diffusion model. Among them, flow field and attention are the two most commonly used methods. Flow fields are good at preserving detailed textures, while attention is better at generating reasonable semantic structures. Previous networks often used only one of the two and failed to make full use of their advantages. At the same time, the flow field and attention also showed complementary functions in the frequency domain. The flow field was good at preserving the high frequency information of the image with the detailed texture, while the semantic structure of attention was good at generating the image with the low frequency information, and few networks used this to improve the generation effect. Based on these facts, this paper introduces the AplusN network, which innovatively addresses the image generation problem by processing from low to high frequencies. For low-frequency information, a conditional large-kernel convolutional attention mechanism (CLA) is employed to capture the global information of the human body. High-frequency information is refined using a spatial-channel normalization module (SCN) to enhance the body's detailed textures. Additionally, we propose a wavelet loss function to align the frequency domain information of the generated images with the target images. Both qualitative and quantitative experiments demonstrate the superiority of our method over state-of-the-art (SOTA) methods, yielding better-defined overall body contours, local details, and higher-quality image generation.
Rui Wang 0034, Weizhi Yang, Wenjian Hu, Wei Xiang 0001
IEEE Trans. Multim.2
2025 Tic action recognition for children tic disorder with end-to-end video semi-supervised learning
Xiangyang Wang 0003, Rui Wang 0034, Jinhua Sun
Vis. Comput.4
2024 DFSTrack: Dual-stream fusion Siamese network for human pose tracking in videos
Xiangyang Wang 0003, Yuhui Tian, Fudi Geng, Rui Wang 0034
Image Vis. Comput.4
2024 TQRFormer: Tubelet query recollection transformer for action detection
Xiangyang Wang 0003, Rui Wang 0034, Jinhua Sun
Image Vis. Comput.4
2024 Emotion recognition based on brain-like multimodal hierarchical perception
Xianxun Zhu, Xiangyang Wang 0003, Rui Wang 0034
Multim. Tools Appl.4
2024 Enhanced keypoint information and pose-weighted re-ID features for multi-person pose estimation and tracking
Xiangyang Wang 0003, Tao Pei, Rui Wang 0034
Mach. Vis. Appl.3
2024 Self-supervised Siamese keypoint inference network for human pose estimation and tracking
Xiangyang Wang 0003, Yuhui Tian, Rui Wang 0034
Mach. Vis. Appl.3
2024 DPR-GAN: Dual-Stream Progressive Refinement for Adversarial 3D Point Cloud Generation
abstract
Abstract Point cloud generation aims to transfer a latent code to realistic 3D shapes through generative models. However, most of progressive generative methods ignore the spatial relationship among different stages and suffer from the loss of contextual information. To address this issue, we propose dual-stream progressive refinement adversarial network (DPR-GAN), which utilizes a dual-stream structure to establish the relationship between two adjacent stages. Such a mechanism can learn the spatial context and preserve more spatial details of point clouds at different stages. In addition, DPR-GAN adopts 3D gridding transformation to guide the shape deformation. In this way, 3D gridding transformation can learn a reasonable correspondence between the local regions of 3D shapes and latent codes. Benefiting from the uniformity and adaptability of the 3D grids, our proposed DPR-GAN can improve the quality and consistency of generated point clouds. We conduct comprehensive experiments to demonstrate that the proposed DPR-GAN is capable of generating pluralistic point clouds, as compared with state-of-the-art generation methods in terms of both visual and quantitative evaluations.
Xiangyang Wang 0003, Rui Wang 0034
Neural Process. Lett.3
2024 A 3D Non-Stationary Small-Scale Fading Model for 5G High-Speed Train Massive MIMO Channels
abstract
The use of fifth-generation (5G) communication technology by high-speed trains (HSTs) has a lot of potential to satisfy current needs for high data rates. Therefore, accurate modeling of the HST wireless channels is crucial for the design and performance assessment of the 5G systems. This paper proposes a general three-dimensional (3D) non-stationary small-scale fading model for 5G HST millimeter wave (mmWave) massive multiple-input multiple-output (MIMO) channels, which captures the wireless channel characteristics in different HST operating scenarios. The proposed channel model has two characteristics. Firstly, it incorporates the distribution of overhead line poles along the railway to characterize the periodic scattering components from the poles. Secondly, it represents complex HST operating scenarios as combinations of five types of scattering clusters, namely hills, trees, lakes, buildings, and concrete. Furthermore, the impact of environmental complexity (EC) on channel statistical properties in the 5G HST massive MIMO scenarios is investigated. Afterwards, based on the birth-death process of scattering clusters, the proposed channel model can characterize the channel non-stationarity in the space-time-frequency domain. Simulation results demonstrate that the proposed model effectively captures the channel non-stationarity, the scattering characteristics of overhead line poles, and the impact of different ECs on system performance. The accuracy and practicality of the proposed model are validated through a comparison with measurement results.
Yichen Feng, Rui Wang 0034, Guoxin Zheng, Asad Saleem, Wei Xiang 0001
IEEE Trans. Intell. Transp. Syst.2
2023 ILETC: Incremental learning for encrypted traffic classification using generative replay and exemplar
Xiuli Ma, Yanliang Jin, Rui Wang 0034
Comput. Networks4
2023 EETC: An extended encrypted traffic classification algorithm based on variant resnet network
Xiuli Ma, Jieling Wei, Yanliang Jin, Dongsheng Gu, Rui Wang 0034
Comput. Secur.6
2023 Human pose estimation based on lightweight basicblock
Xiangyang Wang 0003, Rui Wang 0034
Mach. Vis. Appl.4
2023 PCFN: Progressive Cross-Modal Fusion Network for Human Pose Transfer
abstract
The goal of human pose transfer is to transfer the human in the image from the original pose to the desired one. Existing methods utilizing progressive manner have achieved great success. However, they fail to remove background distraction and preserve appearance details in synthesized images since the correlation between the image and pose is not fully exploit. To this end, we propose a novel progressive cross-modal fusion network (PCFN), which consists of multiple cascaded cross-modal fusion blocks (CMFBs). Each CMFB comprises a feature fusion module (FFM) and a cross-modal module (CMM) to take full advantage of appearance and shape information. From an overall perspective, FFM fully exploits the correlation between image features and pose features through the residual gated convolution. Benefitting from feature integration and dynamic selection, CMFB can extract useful information from the image-pose stream. From a local perspective, CMM utilizes the feature-conditioned gated convolution and the pose-guided heterogeneous attention mechanism to update all codes in a crossing manner and enhance the interaction between fusion information and structural information. Qualitative and quantitative experiments demonstrate the superiority of PCFN compared to state-of-the-art methods, which can transfer the correct human features and increase the authenticity of the generated images. At the same time, PCFN can also be applied to supplement the dataset for person re-identification (ReID). PCFN works well for human pose transfer, and our usage of the gated convolution and the attention mechanism also provides references for other conditional generation tasks.
Rui Wang 0034, Wenming Cao 0001, Wei Xiang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Enhancing multi-scale information exchange and feature fusion for human pose estimation
Rui Wang 0034, Wanyu Wu, Xiangyang Wang 0003
Vis. Comput.1
2022 An approach to adaptive filtering with variable step size based on geometric algebra
abstract
Abstract Recently, adaptive filtering algorithms have attracted much more attention in the field of signal processing. By studying the shortcoming of the traditional real‐valued fixed step size adaptive filtering algorithm, this paper proposed the novel approach to adaptive filtering with variable step size based on Sigmoid function and geometric algebra (GA). First, the proposed approach to adaptive filtering with variable step size based on geometric algebra represents the multi‐dimensional signal as a GA multi‐vector for the vectorization process. Second, the proposed approach to adaptive filtering with variable step size based on geometric algebra solves the contradiction between the steady‐state error and the convergence rate by establishing a non‐linear function relationship between the step size and the error signal. Finally, the experimental results demonstrate that the proposed approach to adaptive filtering with variable step size based on geometric algebra achieves better performance than that of the existing adaptive filtering algorithms.
Yinmei He, Rui Wang 0034
IET Commun.4
2022 MTPose: Human Pose Estimation with High-Resolution Multi-scale Transformers
Rui Wang 0034, Fudi Geng, Xiangyang Wang 0003
Neural Process. Lett.1
2022 Learning Enriched Global Context Information for Human Pose Estimation
Rui Wang 0034, Xiangyang Wang 0003
Neural Process. Lett.1
2022 GA-CNN: Convolutional Neural Network Based on Geometric Algebra for Hyperspectral Image Classification
abstract
Convolutional neural networks (CNNs) have achieved state-of-the-art performance in hyperspectral images (HSIs) classification, which is widely used for the analysis of remotely sensed images. HSI includes spectral and spatial information from several hundreds of spectral data channels. Recent CNN models deal with various bands of HSIs as independent channels, which may lead to the loss of dependencies between different channels or the loss of associated information between each channel and the global. This article proposes a novel CNN model based on geometric algebra (GA), dubbed GA-CNN, to process the HSIs in a holistic way without losing the interrelationship among channels. Specifically, taking advantage of GA, different band images are represented as GA multivectors to capture the inherent structures and preserve the correlation of those channels. In particular, all the basic modules of our model, such as convolutional layers and the backpropagation algorithm, are extended to the GA domain. We evaluate the performance of the proposed GA-CNN model in classification tasks on four well-known HSI datasets. The experimental results indicate that our GA-CNN model outperforms traditional and state-of-the-art real-valued CNNs with higher classification accuracy and fewer model parameters.
Rui Wang 0034, Yi Wang 0063, Xiangyang Wang 0003, Wenming Cao 0001, Wei Xiang 0001
IEEE Trans. Geosci. Remote. Sens.3
2021 RGA-CNNs: convolutional neural networks based on reduced geometric algebra
Rui Wang 0034, Miaomiao Shen, Xiangyang Wang 0003, Wenming Cao 0001
Sci. China Inf. Sci.1
2021 Attention Refined Network for Human Pose Estimation
Xiangyang Wang 0003, Jiangwei Tong, Rui Wang 0034
Neural Process. Lett.3
2021 A Novel Approach of Intelligent Computing for Multiperson Pose Estimation with Deep High Spatial Resolution and Multiscale Features
abstract
Currently, human pose estimation (HPE) methods mainly rely on the design framework of Convolutional Neural Networks (CNNs). These CNNs typically consist of high‐to‐low‐resolution subnetworks (encoder) to learn semantic information and low‐to‐high subnetworks (decoder) to raise the resolution for keypoint localization. Because too low‐resolution feature maps in encoder will inevitably lose some spatial information, which cannot be recovered in the upsampling stages, keeping high spatial resolution features is critical for human pose estimation. On the other hand, due to scale variation of human body parts, multiscale features are also very important for human pose estimation. In this paper, a novel backbone network is proposed specifically for HPE, named High Spatial Resolution and Multiscale Networks (HSR‐MSNet), which maintain high spatial resolution features in deeper layers of the encoder and meanwhile construct multiscale features within one single residual block via subgroup splitting and fusion of feature maps. Experiments show that our approach outperforms other state‐of‐the‐art methods with more accurate keypoint locations on COCO dataset.
Xiangyang Wang 0003, Yijie Shi, Chunhua Qian, Rui Wang 0034
Wirel. Commun. Mob. Comput.6
2020 Enhancing feature fusion for human pose estimation
Rui Wang 0034, Jiangwei Tong, Xiangyang Wang 0003
Mach. Vis. Appl.1
2019 GA-SURF: A new Speeded-Up robust feature extraction algorithm for multispectral images based on geometric algebra
Rui Wang 0034, Yijie Shi, Wenming Cao 0001
Pattern Recognit. Lett.1
2018 Sparse Representation for Color Image Based on Geometric Algebra
abstract
Existing sparse representation models represent RGB channels separately without thinking about the relationship color channels, which lose some color structures inevitably. In this paper, we introduce a novel sparse representation model for color image based on geometric algebra (GA) theory and its corresponding dictionary learning algorithm, namely K-GASVD is proposed. The model represents the color image as a multivector with the spatial and spectral information in GA space, providing a kind of vectorial representation for the inherent color structures rather than a scalar representation via current sparse image models. The proposed sparse model is validated in the applications of color image denoising and reconstruction. The experimental results demonstrate that our sparse image model avoids the hue bias phenomenon successfully and retained the color structures completely. It shows its potential as a general and powerful tool in various applications of color image analysis.
Rui Wang 0034, Miaomiao Shen, Wenming Cao 0001
ICME1
2018 A New Singular Value Decomposition Algorithm for Octonion Signal
abstract
The singular value decomposition (SVD) has been considered as one of the most powerful tools in numerical algebra and has witnessed great success in a wide range of image processing tasks, such as principal component analysis, linear discriminant analysis and sparse representation. However, the existing SVD algorithms cannot directly applied in octonion signals. In this paper, we propose a novel singular value decomposition algorithm for octonion signal, namely OSVD. Firstly, a new real representation according to the components of the original octonion signal is formed and the real SVD for the real matrix is performed. Then with several largest singular values and the corresponding vectors in both left and right unitary matrices selected, the octonion signal can be reconstructed successfully. It is demonstrated by the denoising experiments multispectral image with seven spectral channels that our proposed algorithm significantly outperforms existing state-of-the-art algorithms in both quantitative and visual performance.
Miaomiao Shen, Rui Wang 0034
ICPR2
2017 Sparse fast Clifford Fourier transform
abstract
The Clifford Fourier transform (CFT) can be applied to both vector and scalar fields. However, due to problems with big data, CFT is not efficient, because the algorithm is calculated in each semaphore. The sparse fast Fourier transform (sFFT) theory deals with the big data problem by using input data selectively. This has inspired us to create a new algorithm called sparse fast CFT (SFCFT), which can greatly improve the computing performance in scalar and vector fields. The experiments are implemented using the scalar field and grayscale and color images, and the results are compared with those using FFT, CFT, and sFFT. The results demonstrate that SFCFT can effectively improve the performance of multivector signal processing.
Rui Wang 0034, Yi-xuan Zhou, Yanliang Jin, Wenming Cao 0001
Frontiers Inf. Technol. Electron. Eng.1
2007 Analysis of Higher Order Voronoi Diagram for Fuzzy Information Coverage
Weixin Xie, Rui Wang 0034, Wenming Cao 0001
MSN2