VLDB 2026 Research / reviewers in the wild / expert
Junhuai Li
dblp:81/7435
· DBLP profile ↗
48ranked-venue papers
6as first author
41since 2021 · last 2026
0000-0001-5483-5175ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 10 since 2021Computer networks · 10 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Security and privacy · 3 · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-domain Human Activity Recognition based on variational Bayesian Gaussian mixture model and Optimal Transport
Huaijun Wang, Changrui Cui, Junhuai Li, Wei Xiang 0001 |
Knowl. Based Syst. | 4 |
| 2026 | MDR-MSA: multi-perspective decoupled representation learning for multimodal sentiment analysis
Jiangying Du, Yuxing Zhi, Huaijun Wang, Kan Wang 0010, Junhuai Li |
Multim. Syst. | 5 |
| 2026 | Dual-driven optimization of collaborative multi-agent via case learning and curiosity
Ruizhu Chen, Rong Fei, Junhuai Li, Yalin Miao |
Neural Networks | 3 |
| 2026 | A graph contrastive learning network for change detection with heterogeneous remote sensing images
Zhiyong Lv, Sizhe Cheng, Linfu Xie, Junhuai Li, Minghua Zhao |
Pattern Recognit. | 4 |
| 2026 | CFSDBN: Emotion Recognition via Channel Feature Selection and Dynamic Brain NetworkabstractThe functional connectivity patterns revealed by emotion recognition closely align with neural pathways involved in emotional processing. Traditional methods often fail to adequately integrate the multidimensional spatiotemporal-spectral characteristics of Electroencephalography (EEG), making it challenging to accurately characterise dynamic emotional processes. Predefined brain regions and fixed thresholds yield coarse functional networks, which limit the accurate identification of critical connections. Furthermore, black-box models lack interpretability, providing little decision support or visualisable evidence of neural circuits for neuroscience and clinical applications. To address these limitations, this study proposes an emotion recognition via channel-feature selection and dynamic brain network (CFSDBN). First, spatial-spectral features are extracted using a residual network, while bidirectional gated recurrent units capture temporal dynamics, thereby improving feature utilisation. Next, these spatio-temporal-spectral features serve as node inputs to a graph attention network, where node attention weights enable adaptive channel selection and sparse functional connectivity learning, thereby overcoming localisation inaccuracies caused by coarse-grained processing. Finally, joint node embeddings and connection weights are used for emotion classification, and key channels and neural circuits are visualised to provide interpretable evidence for emotional neural mechanisms. By deeply coupling multidimensional features with brain network optimisation, CFSDBN achieves significant improvements in classification performance on the DEAP, MODMA, and SEED-V datasets. It enhances hierarchical interpretability from micro-level features to macro-level network interactions, offering a high-performance and explainable solution for EEG-based emotion recognition. Jiawei Du 0004, Yuxing Zhi, Junhuai Li, Huaijun Wang, Fangping Xia |
IEEE Trans. Affect. Comput. | 3 |
| 2026 | A Multimodal Sentiment Analysis Approach Based on Multiview Cross-Modal FusionabstractMultimodal sentiment analysis in social video is the foundation of affective computing and artificial intelligence. Due to the heterogeneity of multimodal data including facial expression, speech, and language, multimodal sentiment analysis remains a challenging problem. Existing methods mostly explore delicate fusion strategies to obtain consistent multimodal sentiment representation. There is still the problem of cross-modal bias due to distributional differences in multimodal data. To address this problem, a multimodal sentiment analysis method based on multiview cross-modal fusion (MVCF) is proposed in this article, which significantly improves the performance of extraction between heterogeneous architectures from three perspectives. The unimodal local context module is proposed to implement one to one pattern, reflecting independent change factors combined with specific information adaptively. We propose global union modules to learn the one-to-all pattern by projecting intermediate features into an aligned latent space, where modality-specific information is discarded. To align cross-modal affective prototypes, the all-to-all multimodal fusion module performs adaptive fusion for all modalities and interactions to eliminate perturbing information in modal commonality. Experimental results on multimodal sentiment analysis benchmarks show that our MVCF outperforms baselines in complete and incomplete modality settings, demonstrating the effectiveness and robustness of MVCF. Our code is publicly available onhttps://github.com/zhixingyu/MVCF. Yuxing Zhi, Junhuai Li, Huaijun Wang |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2026 | Joint Trajectory and Power Design With Cooperative Jamming UAV Assistance Based on Reinforcement LearningabstractWe examine a secure wireless communication system that is enabled by unmanned aerial vehicles (UAVs) in this research. In the wireless communication system with an eavesdropping UAV, we deploy a relay UAV to facilitate the transmission of confidential signals from the source station (denoted asS) to ground users. Additionally, we select an idle relay UAV to act as a cooperative jamming UAV, sending interference signals to the eavesdropping UAV. It is quite feasible that the eavesdropping UAV will leverage its mobility to improve the quality of its eavesdropping, making its trajectory unpredictable. First, to address the worst-case scenario for the ground user’s security performance, we assume the eavesdropping UAV approaches at the closest distance.We aim to maximize the worst secrecy rate under perfect CSI via designing the flight trajectory and transmission power of both the relay UAV and the jamming UAV. Second, we investigated the performance of the system’s secrecy outage probability under imperfect CSI. The presence of eavesdropping UAVs and the unpredictable nature of their environment makes traditional convex optimization methods mathematically complex for solving the trajectory optimization problem of the relay and jamming UAVs. To address this, we propose the Multi-Agent joint design trajectory and power (MAJDTP) algorithm based on the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm to optimize the flight trajectory and transmission power of both UAVs. During the design and training process, the relay and jamming UAVs are treated as agents to derive their optimal flight paths and transmission energy. Finally, our approach surpasses the benchmark algorithm, as demonstrated by the simulation results. Yingkun Wen, Fengshuan Wang, Hui-Ming Wang 0001, Junhuai Li, Kan Wang 0010, Huaijun Wang |
IEEE Trans. Wirel. Commun. | 4 |
| 2025 | Adversarial Training and Cross-modal Feature Fusion in Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis recognizes emotions through text, audio, and visual modalities, but data incompleteness is a major challenge. Existing methods often focus on specific types of deficiencies and perform poorly when multiple types of noise are present simultaneously. To address this issue, we propose a noise-prompted adversarial training framework with a multimodal interaction model to enhance the model’s robustness to missing modalities. The model first extracts common and unique features from each modality using a BERT text encoder and a shared-private encoder. Correlation measurements are then used to calculate the similarity between modalities, and a weighting mechanism is applied to the shared features. These features are deeply fused using a Transformer, and adversarial training combined with semantic reconstruction supervision helps the model learn a unified representation of noisy and clean data. Experimental results show that this method significantly improves the performance of multimodal sentiment analysis. Junhuai Li, Huaijun Wang, Yuxing Zhi, Tao Huang 0008 |
ICASSP | 1 |
| 2025 | HFedCWA: heterogeneous federated learning algorithm based on contribution-weighted aggregation
Jiawei Du 0004, Huaijun Wang, Junhuai Li, Kan Wang 0010, Rong Fei |
Appl. Intell. | 3 |
| 2025 | Rethinking probability volume for multi-view stereo: A probability analysis method
Zonghua Yu, Huaijun Wang, Junhuai Li, Haiyan Jin, Ting Cao 0002, Kuanhong Cheng |
Appl. Intell. | 3 |
| 2025 | Cross-domain human activity recognition based on deviation-graph constrained Non-Negative Matrix Factorization
Yuxing Zhi, Huaijun Wang, Kan Wang 0010, Lei Yu 0010, Rong Fei, Junhuai Li |
Eng. Appl. Artif. Intell. | 8 |
| 2025 | Cooperative Jamming Aided Secure Communication for RIS Enabled Symbiotic Radio SystemsabstractEnsuring signal confidentiality against eavesdroppers is particularly challenging, especially with imperfect channel state information (CSI). To address this, we propose a novel approach leveraging reconfigurable intelligent surfaces (RISs) to enhance security and optimize transmission performance. This paper focuses on secure communication in symbiotic radio (SR) systems by investigating cooperative jamming-assisted transmission with RISs, providing a robust solution to these challenges. RIS-I, acting as a secondary transmitter (STx), multicasts confidential signals from the primary transmitter (Alice) to a primary user (Bob), protecting against eavesdropping by Eve. Additionally, RIS-I transmits its own signals to a secondary user (SU) using backscattering radio technology. Meanwhile, RIS-II serves as a cooperative jammer, converting received confidential signals from Alice into jamming signals by strategically adjusting its reflection coefficients to disrupt Eve’s reception. These RISs can operate cooperatively; when RIS-II transmits as an STx, RIS-I functions as a cooperative jammer. We explore two scenarios: 1. With perfect CSI for the wiretap channel, we propose a joint SDR(Semi-definite relaxation)+MM(Minorization-maximization) optimization algorithm to simultaneously optimize Alice’s beamforming vector and the RISs’ reflection coefficients. 2. With imperfect CSI, we derive the secrecy outage probability formula and evaluate the scheme’s performance across different scenarios. Numerical results demonstrate that RIS-assisted cooperative jamming significantly enhances the secrecy rate and reduces the secrecy outage probability for Bob, outperforming traditional RIS-assisted SR systems. Yingkun Wen, Fengshuan Wang, Hui-Ming Wang 0001, Junhuai Li, Kan Wang 0010, Huaijun Wang |
IEEE Trans. Commun. | 4 |
| 2025 | Content-Adaptive Multi-Region Deep Network for Polarimetric SAR Image ClassificationabstractDeep learning methods excel in Polarimetric SAR (PolSAR) image classification. However, existing methods typically sample an image block for each pixel with a fixed-size square window, which always contains inconsistent/incomplete content with the central pixel, resulting in many misclassifications especially in boundary and heterogeneous regions. So, a size-fixed square window is not enough for representing various terrain objects. To address this issue, we develop a content-adaptive multi-region deep network to obtain contextual consistent sampling windows for diverse terrain objects. Firstly, a complex scene of PolSAR image is partitioned into homogeneous, heterogeneous and boundary regions. Then, sampling windows with adaptive direction and scale are designed for three distinct regions. Besides, windows with central and global regions are proposed to provide additional local and global information. Finally, a fusion network is designed to adaptively combine different sampling windows to enhance classification performance. Experimental results on three real data sets demonstrate that the proposed method can achieve superior performance in both edge details and heterogeneous terrain objects compared with the state-of-the-art methods. Junfei Shi, Shanshan Ji, Haiyan Jin, Junhuai Li, Maoguo Gong, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Hierarchical Feature Fusion Triple Network for Change Detection With Bitemporal Remote Sensing ImagesabstractAchieving land cover change detection (LCCD) through remotely sensed images (RSIs) is important in the observation of the changes on the Earth’s surface. In such detection, spectral-reflectance noise and the uncertainty of the imaging external conditions for the bitemporal RSIs usually cause some salt-and-pepper noisy pixels in the results and reduce the change detection accuracy. In this article, a hierarchical feature-fusion triple network (HFTN) is proposed to improve the performance of LCCD with RSIs. Overall, the proposed HFTN aims to learn representative features to improve change detection performance via two feature learning enhancement strategies and a hierarchical feature-fusion mechanism. First, an image feature difference model is proposed to generate the input feature for the middle branch and guide the learning performance. Second, a progressive denoising module (PDM) is proposed and applied to each temporal image to reduce the noise before feeding the features into the backbone of the proposed HFTN. Finally, a hierarchical feature-fusion module (HFFM) is proposed to fuse the learned deep feature for generating a change-magnitude image. Additionally, multiscale convolution, cross-scale fusion, and a shared weight are adopted in the backbone of the proposed HFTN to further enhance the feature learning performance. Compared with eight state-of-the-art methods, experimental results verified the feasibility and superiority of the proposed HFTN for LCCD with RSIs. For example, the proposed HFTN achieved improvement rates of approximately 0.43%–11.83% for overall accuracy (OA) and 0.11%–4.81% for false alarms (FAs) across six pairs of real RSIs. The code can be available athttps://github.com/ImgSciGroup/HFTN-NET.git. Zhiyong Lv, Tianyv Yang, Pingdong Zhong, Weiwei Sun 0005, Jón Atli Benediktsson, Junhuai Li |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Generative Adversarial Network-Aided Covert Communication for Cooperative Jammers in CCRNsabstractThis paper investigates a centralized cooperative cognitive radio network (CCRN) where a primary base station (PBS) transmits a message to a primary user while a secondary user transmitter (SU-Tx) function as a friendly jammer. The jammer sends jamming signals to protect the PBS’s messages from a potential eavesdropper (Eve). However, the SU-Tx also attempts to covertly transmit its own messages to a secondary user receiver using the allocated spectrum resource, contravening the PBS regulations. To address this issue, the PBS requests its partner CBS to help detect jammer’s behavior. Specifically, we propose a generative adversarial network (GAN) optimization framework that models the strategic game between the CBS monitoring and the covert transmission of cooperative jammers. We introduce a novel GAN-based beamforming design algorithm, termed GAN-BD, to determine the power allocation at the jammer for covert communication. Additionally, we develop the detection error probability (DEP) at the CBS and derive its expression using a hypothesis testing problem. Through extensive simulation results, we demonstrate that the proposed GAN-BD algorithm can achieve near-optimal solutions for conducting covert communication, leveraging knowledge of the current network environment and exhibiting rapid convergence capabilities. The simulation results highlight the effectiveness of our GAN-BD algorithm. Yingkun Wen, Yan Huo 0001, Junhuai Li, Kan Wang 0010 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | A Two-Step Cellular Network Traffic Forecasting Method Integrating Decomposition and Deep Neural Networks Based on Bayesian Joint Parameter OptimizationabstractAccurate cellular network traffic prediction is crucial for intelligent network planning and management in 6G. However, the non-stationary characteristics of cellular network traffic present significant challenges when training deep neural networks for traffic forecasting. To address this issue, we propose a two-stage deep learning framework, JO-DPNet, based on Bayesian joint parameter optimization, which integrates data decomposition techniques with Bayesian joint optimization to effectively mitigate the adverse impacts of non-stationarity and error accumulation on prediction accuracy. In the first stage, a data decomposition module uses Variational Mode Decomposition (VMD) to decompose the original data into network traffic subset series(TSS), thereby alleviating the negative effects of non-stationarity. In the second stage, a prediction and construction module leverages a bi-directional LSTM (Bi-LSTM) network to extract deep spatial-temporal features from the TSS in a bidirectional manner. A fully connected layer then captures the relationships between the TSS and reconstructs the predicted results into the final output. The JO module employs the Tree-structured Parzen Estimator based Bayesian optimization algorithm(TP-BO) simultaneously determines the optimal VMD mode number k and the hyperparameters of the Bi-LSTM network through probabilistic surrogate model. Extensive experiments on three real-world cellular traffic datasets demonstrate that the proposed method significantly mitigates the non-stationary characteristics of the traffic data. Compared to state-of-the-art methods, JO-DPNet achieves reductions in MAE by 29%, 3%, and 19% for three type prediction tasks on the Telecom Italia dataset. The source code is available to the public at: https://github.com/VicentZhang259/JO-DPNet. Pengfei Zhang 0012, Junhuai Li, Dong Ding 0002, Huaijun Wang, Kan Wang 0010, Xiaofan Wang 0002 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2024 | Prediction of miRNA family based on class-incremental learningabstractWith the development of deep sequencing, recent studies indicate that a miRNA precursor can generate multiple miRNA isoforms (isomiRs). The family prediction of canonical miRNAs and isomiRs could provide a basis for miRNA functional research. In this study, we propose a novel method for family identification of canonical miRNA and isomiRs based on incremental learning. First, a benchmark dataset is constructed by processing data based on miRNA sequences and their family annotation. Moreover, sequence embedding and RNN are used for capturing essential features and inherent dependencies within a sequence. Finally, incremental learning is applied to accommodate the continuous influx of miRNA sequencing data, enabling RNN to stay relevant and effective over time. Comparative experiments and ablation studies illustrate the effectiveness of our model, which can help to comprehensively understand miRNA’s function. Lulu Qiu, Rong Fei, Junhuai Li, Fang-Xiang Wu |
BIBM | 5 |
| 2024 | Activity Recognition Method Based on Kernel Supervised Laplacian EigenmapsabstractLaplacian dimensionality reduction can effectively achieve feature transformation and preserve the important structure of high-dimensional features. However, the trained model with this method usually require better generalization ability to new samples. Hence, a human activity recognition method based on kernel-supervised laplacian eigenmaps (KSLE) by combining the kernel method, laplacian mapping, and supervised learning is proposed. Firstly, the adjacency distance relationship of original samples features in the high-dimensional kernel space is obtained. Secondly, the category labels in the high-dimensional feature set are incorporated in the manifold learning for dimensionality reduction. Then, kernel trick is utilized to directly solve the low-dimensional embedding structure of the test feature set. Finally, the obtained low-dimensional feature set is input into the classifier for recognition. Extensive experiments conduct on two public datasets confirm the effectiveness of the proposed approach for improving the generalization ability of new samples. Pengjia Tu, Dandan Du, Junhuai Li, Huaijun Wang |
ICASSP | 4 |
| 2024 | A Fine-Grained Tri-Modal Interaction Model for Multimodal Sentiment AnalysisabstractThe methods based on multimodal representation learning enhance discriminable sentiment expression for multimodal sentiment analysis(MSA). The modal invariant and specific features serve different purposes in sentiment learning and the diversity of inter-sample and inter-category relationships takes less consideration in previous advances. In this paper, we propose a fine-grained tri-modal interaction model for MSA to refine and enhance the overall affective state at the uni/multi-modal level and label level. Concretely, we simultaneously focus on the contributions of both unimodal and multimodal views to improve holistic affective knowledge. The similarity measurement function and self-supervised learning are introduced to reduce the inherent modality gap and refine the invariant representations. We provide two specific constraints to strengthen the specific uniqueness and discriminative capacity. Moreover, we develop a multi-granularity fusion module to fully integrate two views in a coarse-to-fine form. Experimental results on two datasets show that our method fares better than the state-of-the-art model. Yuxing Zhi, Junhuai Li, Huaijun Wang, Ting Cao 0002 |
ICASSP | 2 |
| 2024 | CNN-Enhanced Deep Sparse Representation Network for Polarimetric SAR Image ClassificationabstractDeep learning networks can automatically acquire high-level semantic features for polarimetric SAR image classification, while it involves a blind learning procedure without explicit guidance. In contrast, sparse representation methods represent effective non-deep models with a robust mathematical mechanism serving as guidance. However, they can’t capture complex image features and semantic information. To address these issues, we propose a novel approach known as the CNN-enhanced Deep Sparse Representation Network (CE-DSRNet) for PolSAR image classification, which a Sparse Representation (SR) guided deep learning model. Initially, a sparse representation model is constructed for PolSAR images to capture essential features. Subsequently, to solve the sparse model, a Deep Sparse Representation Network (DSRNet) is devised by transforming the Soft Threshold Iterative (ISTA) optimization procedure into a network, enabling automatic learning of sparse coefficients as features. Finally, a CNN-enhanced DSRNet is introduced, integrating DSRNet with CNN to effectively extract deep semantic features and enhance classification accuracy. Experiments demonstrate the effectiveness of the proposed method compared to state-of-the-art approaches. Junfei Shi, Mengmeng Nie, Haiyan Jin, Junhuai Li, Yuanlin Zhang 0003 |
IGARSS | 4 |
| 2024 | Cross-Domain Activity Recognition Based on Stacked Transfer NetworkabstractHuman activity recognition (HAR) based on wearable sensors is a hot topic in health detection and motion management. Nonetheless, conventional identification methods necessitate substantial labeled datasets, and the acquisition of high-quality labeled data crucial for human activity recognition is both time-consuming and costly. To tackle this problem, transfer learning is used to annotating unlabeled or a few labeled target domains using labeled source domains. Meanwhile, individuals exhibit significant differences in amplitude, angles, and other aspects when performing the same action. Due to the limited expressive capacity of time series, effectively capturing various types of differences simultaneously poses a challenge, impacting the effectiveness of transfer. Therefore, in this paper, we propose a stacked transfer network (STN) for 2D modeling and feature extraction of sensor data in steps. It adaptively decomposes complex temporal variations into multiple intra- and inter-periodic variations using the Fast Fourier Transform (FFT), allowing the model to focus more on the trends of the activty and less on the effects of individual discrepancy. Our comprehensive cross-domain activity recognition (CDAR) experiments on three large public activity recognition datasets (i.e., OPPORTUNITY, PAMAP2, and UCI-DSADS) show that the STN achieves high accuracy in activity recognition ttransfer. Junhuai Li, Jingyi Cao, Yuxing Zhi, Huaijun Wang, Ting Cao 0002, Rong Fei |
IJCNN | 1 |
| 2024 | Generative Diffusion Model-Based Deep Reinforcement Learning for Uplink Rate-Splitting Multiple Access in LEO Satellite NetworksabstractThis work studies the joint transmit power control and receive beamforming in uplink rate splitting multiple access (RSMA)-based low earth orbit (LEO) satellite networks, using both generative diffusion model and proximal policy optimization (PPO) learning framework. In particular, using RSMA, interference is partially decoded and partially treated as noise, thereby improving the spectral efficiency, while the dynamics and uncertainty in LEO satellite networks would pose challenges to the real-time power control and receive beamforming optimization. First, a long-run sum data rate maximization problem is formulated, subject to the individual data rate requirement, and then the Markov decision process (MDP) is used to model it. Second, on the basis of MDP, a generative diffusion model-based proximal policy optimization (PPO) framework is proposed, where a denoising network is taken as the actor network in PPO to output the optimal continuous policy, thereby facilitating the hyperparameter tuning and improve the sample efficiency. Finally, experiments are conducted to show advantages of merging diffusion model into PPO, in terms of larger spectral efficiency, by comparing proposed framework with benchmarks. Xingjie Wang, Kan Wang 0010, Di Zhang 0004, Junhuai Li, Momiao Zhou, Timo Hämäläinen 0002 |
ISCC | 4 |
| 2024 | Action recognition method based on multi-stream attention-enhanced recursive graph convolution
Huaijun Wang, Bingqian Bai, Junhuai Li, Hui Ke, Wei Xiang 0001 |
Appl. Intell. | 3 |
| 2024 | Human Activity Recognition based on Local Linear Embedding and Geodesic Flow Kernel on Grassmann manifolds
Huaijun Wang, Changrui Cui, Pengjia Tu, Junhuai Li, Wei Xiang 0001 |
Expert Syst. Appl. | 5 |
| 2024 | CRPF-QC: An Efficient CSI Recurrence Plot-Based Framework for Queue CountingabstractQueue counting using WiFi channel state information (CSI) faces challenges due to susceptibility to external factors and relies on ideal testing environments for current methods. We propose an efficient CSI recurrence plot (RP)-based framework for queue counting (CRPF-QC), containing a transformation module and a recognition module. The conversion module transforms the CSI into RP, distinct from traditional models using a single signal point as the unit for feature extraction, utilizing the signal changes at different timestamps as units for feature extraction and effectively preserving the amplitude and phase relationships between any two time points. In the recognition module, the convolutional neural network (CNN) and the long short-term memory (LSTM) network are combined to profoundly understand the internal structure and changes within the image. The proposed integration framework is adept in the automatic extraction of amplitude and phase features, therefore improving image recognition accuracy. Meanwhile, we explore dynamic changes in the queuing crowd detection based on the Fresnel zone theory, identifying individuals’ entering and exiting behaviors at different positions within the Fresnel zone and updating the count accordingly, which makes up for the shortcomings of the static model. Intensive evaluations demonstrate that CRPF-QC, employing just two layers of CNN and one layer of LSTM, excels in adapting to dynamic environmental changes, outperforming traditional queue counting methods. Additionally, the dynamic model attains a perfect 100% accuracy in both scenarios. Rong Fei, Junhuai Li, Yuxin Wan, Zhongqi Zhao, Majid Habib Khan |
IEEE Internet Things J. | 3 |
| 2024 | A Dual-Scale Transformer-Based Remaining Useful Life Prediction Model in Industrial Internet of ThingsabstractWith recent advents of industrial Internet of Things (IIoT), the connectivity and data collection capabilities of industrial equipment have be significantly enhanced, yet bringing new challenges for the remaining useful life (RUL) prediction. To fulfill the RUL predicting demand in multivariate time series, this work proposes an encoder-decoder model termed as dual-scale transformer model (DSFormer), built upon the Transformer architecture. First, in the encoder part, a dual-attention module is designed for the weight feature extraction from both dimensions of the sensor and time series, aiming to compensate for the diverse impacts of different sensors on the prediction. Next, a temporal convolutional network (TCN) module is introduced to capture sequence features and alleviate the loss of positional information incurred by stacking blocks. Then, the feature decomposition module is integrated into the decoder for trend feature extraction from sequences, providing the model with additional sequence information. Finally, compared to existing models, the proposed method can obtain the superior performance in terms of the root mean square error (RMSE) and Score metrics on the FD001, FD002 and FD003 subsets of the C-MAPSS dataset, with an average improvement of 3.2% and 2.5% respectively. In particular, the ablation experiment further validates the effectiveness of proposed modules in handling multivariate time series and extracting features. Junhuai Li, Kan Wang 0010, Xiangwang Hou, Dapeng Lan, Yunwen Wu, Huaijun Wang, Lei Liu 0031, Shahid Mumtaz |
IEEE Internet Things J. | 1 |
| 2024 | A multi-scale no-reference video quality assessment method based on transformer
Yingan Cui, Zonghua Yu, Yuqin Feng, Huaijun Wang, Junhuai Li |
Multim. Syst. | 5 |
| 2024 | 3D human pose estimation method based on multi-constrained dilated convolutions
Huaijun Wang, Bingqian Bai, Junhuai Li, Hui Ke, Wei Xiang 0001 |
Multim. Syst. | 3 |
| 2024 | A Multimodal Sentiment Analysis Method Based on Fuzzy Attention FusionabstractAffective analysis is a technology that aims to understand human sentiment states, and it is widely applied in human–computer interaction and social sentiment analysis. Compared to unimodal, multimodal sentiment analysis (MSA) focuses more on the complementary information and differences from multimodalities, which can better represent the actual sentiment expressed by humans. Existing MSA methods usually ignore the problem of multimodal data ambiguity and the uncertainty of influence redundant features on the sentiment discriminability. To address these issues, we propose a fuzzy attention fusion-based MSA method, called FFMSA. FFMSA alleviates the heterogeneity of multimodal data through shared and private subspaces, and solves the ambiguity using a fuzzy attention mechanism based on continuous value decision making, in order to obtain accurate sentiment features for downstream tasks. The private subspace refines the latent features within each single modality through constraints on their uniqueness, while the shared subspace learns common features using a nonparametric independence criterion algorithm. By constructing sample pairs for unsupervised contrastive learning, we use fuzzy c-means to model uncertainty to constrain the similarity between similar samples to enhance the expression of shared features. Furthermore, we adopt a multiangle modeling approach to capture the consistency and complementarity of multimodalities, dynamically adjusting the interaction between different modalities through a fuzzy attention mechanism to achieve comprehensive sentiment fusion. Experimental results on two datasets demonstrate that our FFMSA outperforms state-of-the-art approaches in MSA and emotion recognition. The proposed FFMSA achieves sentiment binary classification accuracy of 85.8% and 86.4% on CMU-MOSI and CMU-MOSEI, respectively. Yuxing Zhi, Junhuai Li, Huaijun Wang, Wei Wei 0006 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | Detecting the Background-Similar Objects in Complex Transportation ScenesabstractWith the development of intelligent transportation systems, most human objects can be accurately detected in normal road scenes. However, the detection accuracy usually decreases sharply when the pedestrians are merged into the background with very similar colors or textures. In this paper, a camouflaged object detection method is proposed to detect the pedestrians or vehicles from the highly similar background. Specifically, we design a guide-learning-based multi-scale detection network (GLNet) to distinguish the weak semantic distinction between the pedestrian and its similar background, and output an accurate segmentation map to the autonomous driving system. The proposed GLNet mainly consists of a backbone network for basic feature extraction, a guide-learning module (GLM) to generate the principal prediction map, and a multi-scale feature enhancement module (MFEM) for prediction map refinement. Based on the guide learning and coarse-to-fine strategy, the final prediction map can be obtained with the proposed GLNet which precisely describes the position and contour information of the pedestrians or vehicles. Extensive experiments on four benchmark datasets, e.g., CHAMELEON, CAMO, COD10K, and NC4K, demonstrate the superiority of the proposed GLNet compared with several existing state-of-the-art methods. Bangyong Sun, Nianzeng Yuan, Junhuai Li |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Covert Communications Aided by Cooperative Jamming in Overlay Cognitive Radio NetworksabstractThis paper examines integrating jamming and secondary signals for covert communications in cognitive radio networks (CRNs), aiming to enhance covertness by using jamming and secondary signals in an overlay cooperative CRN. The scenario involves a primary base station (PBS) transmitting to a primary user (PU), with a secondary user transmitter (SU-Tx) acting as a cooperative jammer to obscure the message from a malevolent secondary user named “Willie.” During idle intervals on the primary channel, the SU-Tx opportunistically accesses it to transmit secondary signals, reinforcing the covert communication of primary signals. The study quantifies the detection error probability (DEP) experienced by Willie, considering perfect and statistical channel state information (CSI) scenarios. In the perfect CSI scenario, optimization has two phases. Phase I aims to maximize the signals-to-interference-plus-noise ratio (SINR) of the PU, subject to the warden DEP exceeding a specified threshold. Phase II uses an iterative search algorithm to optimize beamforming vectors, enhancing SINR. In the statistical CSI scenario, the goal is to maximize effective transmission throughput (ETT), measuring the information transmitted from PBS to PU under covert constraints. Numerical results validate the theoretical analysis. Yingkun Wen, Lei Liu 0031, Junhuai Li, Yilan Li, Kan Wang 0010, Shui Yu 0001, Mohsen Guizani |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | A high-performance deep-learning-based pipeline for whole-brain vasculature segmentation at the capillary resolutionabstractMOTIVATION: Reconstructing and analyzing all blood vessels throughout the brain is significant for understanding brain function, revealing the mechanisms of brain disease, and mapping the whole-brain vascular atlas. Vessel segmentation is a fundamental step in reconstruction and analysis. The whole-brain optical microscopic imaging method enables the acquisition of whole-brain vessel images at the capillary resolution. Due to the massive amount of data and the complex vascular features generated by high-resolution whole-brain imaging, achieving rapid and accurate segmentation of whole-brain vasculature becomes a challenge. RESULTS: We introduce HP-VSP, a high-performance vessel segmentation pipeline based on deep learning. The pipeline consists of three processes: data blocking, block prediction, and block fusion. We used parallel computing to parallelize this pipeline to improve the efficiency of whole-brain vessel segmentation. We also designed a lightweight deep neural network based on multi-resolution vessel feature extraction to segment vessels at different scales throughout the brain accurately. We validated our approach on whole-brain vascular data from three transgenic mice collected by HD-fMOST. The results show that our proposed segmentation network achieves the state-of-the-art level under various evaluation metrics. In contrast, the parameters of the network are only 1% of those of similar networks. The established segmentation pipeline could be used on various computing platforms and complete the whole-brain vessel segmentation in 3 h. We also demonstrated that our pipeline could be applied to the vascular analysis. AVAILABILITY AND IMPLEMENTATION: The dataset is available at http://atlas.brainsmatics.org/a/li2301. The source code is freely available at https://github.com/visionlyx/HP-VSP. Xuhua Liu, Xueyan Jia, Jianghao Wu 0005, Qianlong Zhang, Junhuai Li, Anan Li |
Bioinform. | 7 |
| 2023 | Novel Enhanced UNet for Change Detection Using Multimodal Remote Sensing ImageabstractLand cover change detection (LCCD) with bitemporal remote sensing images has been widely used in practical applications. However, when the bitemporal images are multimodal remote sensing images (MRSIs) which are acquired with different sensors, the change detection performance may be unsatisfactory, because MRSIs cannot be compared directly to generate a change magnitude and obtain a change detection map. Here a novel approach is proposed to overcome this problem, i.e., the Enhanced UNet (E-UNet) which learns deep shared features from MRSIs to achieve change detection with MRSIs. First, apre-event image to post-eventimage (P2P) transformation module based on classical Cycle-consistent Generative Adversarial Network (CGAN) is suggested to embed at the head of the proposed E-UNet to translate the pre-event image to a post-event image one. Then, multi-scale convolutions are added at each encoding layer to capture the various shapes and sizes of ground targets. Finally, a Polarized Self-Attention (PSA) module is employed before beginning the decoding progress of E-UNet with an aim to pay extra attention to changed areas. Compared with five typical state-of-the-art methods, experimental results based on two pairs of MRSIs well demonstrated the feasibility and advantages of the proposed E-UNet for LCCD with MRSIs in terms of visual observations and quantitative evaluations. For example, the improvement is 4.19% and 4.75% in terms of the overall accuracy for the Sardinia dataset and California dataset, respectively. The code of the proposed approach can be found at https://github.com/ImgSciGroup/E-UNet. Zhiyong Lv, Weiwei Sun 0005, Tao Lei 0003, Jón Atli Benediktsson, Junhuai Li |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2023 | Deep-block network for AU recognition and expression migration
Minghua Zhao, Yuxing Zhi, Junhuai Li, Jing Hu 0005, Shuangli Du, Zhenghao Shi |
Multim. Tools Appl. | 4 |
| 2023 | AngClust: Angle Feature-Based Clustering for Short Time Series Gene Expression ProfilesabstractWhen clustering gene expression, it is expected that correlation coefficients of genes in the same clusters are high, and that gene ontology (GO) enrichment analysis of most clusters will be significant. However, existing short-term gene expression clustering algorithms have limitations. To address this problem, we proposed a novel clustering process based on angular features for short-term gene expression. Our method (named AngClust) uses angular features to indicate the change of trend in gene expression levels at two neighboring time points. The changes of angles at multiple time points reflects the change of trend of the overall expression levels. Such changes are used to measure whether the expression trends of different genes are similar. To obtain functionally significant clusters from the clustering results, we evaluated numbers of genes in clusters, average correlation coefficient, fluctuation, and their correlation with GO term enrichment. The efficacy of AngClust outperform two other measures, Euclidean distance (ED) and dynamic time warping of correlation (DTW), on a dataset of yeast gene expression. The ratios of GO and pathway term-enriched of clusters of AngClust is higher than or equal to that of STEM and TMixClust on human, mouse, and yeast time series of gene expression. Junhuai Li, Saurav Mallik, Rong Fei, Hongfang Zhou |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | Novel Adaptive Region Spectral-Spatial Features for Land Cover Classification With High Spatial Resolution Remotely Sensed ImageryabstractSpectral-spatial features are important for ground target identification and classification with High Spatial Resolution Remotely Sensed (HSRRS) Imagery. In this paper, two novel features, named the Gaussian-Weighting Spectral (GWS) feature and the Area Shape Index (ASI) feature, are proposed to complement the deficiency of the basic image feature for land cover classification with HSRRS imagery. The proposed GWS feature is an adaptive region-based feature that aims to improve the spectral homogeneity of a local area surrounding a pixel. Additionally, it is well known that the spectral feature is inadequate for classifying HSRRS imagery. Therefore, one spatial feature called the ASI feature is proposed here to describe the relationship between the area and shape for an adaptive region around each pixel. The proposed GWS and ASI features coupled with the basic red-green-blue feature are fed into a supervised classifier to obtain the final classification map. Experiments based on four real HSRRS images demonstrate that the proposed GWS and ASI features are capable of improving classification accuracies compared with some cognate state of the art methods. Moreover, the experiments also reveal that the proposed spectral-spatial features can complement each other for enhancing the classification performance with HSRRS images. Zhiyong Lv, Pengfei Zhang 0012, Weiwei Sun 0005, Jón Atli Benediktsson, Junhuai Li |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Service Function Chaining in Industrial Internet of Things With Edge Intelligence: A Natural Actor-Critic ApproachabstractOwing to network function virtualization (NFV), each industrial application is constructed as a service function chain (SFC), concatenating the ordered service functions, to offer applications more flexibly in industrial Internet of Things (IIoT). When it comes to the emerging edge intelligence, the integration of NFV with edge in IIoT would enable more close-proximity services, yet also posing new challenges owing to more complicated environment. Although some efforts have been made to service function chaining in IIoT, the radio resource dynamics are not fully perceived. In this article, we investigate the radio-aware SFC deployment in the edge-enabled IIoT. First, a radio-aware deployment formulation is exhibited, steering the flow traversing both wireless and wired links. Next, Markov decision process is exhibited to track dynamics in both IIoT and radio resources. Afterwards, the natural gradient-based actor-critic SFC paradigm is introduced to adapt to network variation, by incorporating the curvature of parameter space into gradient information. To resolve the high-dimensionality in action space, we then recur to the norm penalty approach, reducing the space size by two orders of magnitude. Finally, numerical experiments are executed to uncover superiority of presented method, disclosing that the latency performance benefits from both the SFC routing between IIoT servers and elaborated wireless resource orchestration. Junhuai Li, Kan Wang 0010 |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | Guide Tracking Method Based On Particle Filter FusionabstractTraditional localization techniques cannot meet the requirements of indoor localization accuracy. Although PDR can obtain high localization accuracy in a relatively short period of time, its localization error will gradually accumulate as the user's walking distance increases. Therefore, this paper proposes a particle filter fusion-based guided trajectory tracking method, which combines pedestrian heading estimation and convolutional neural network-based landmark detection method to achieve real-time tracking of position and trajectory. The article uses the PDR method for estimation, including the number of steps and step lengths, to achieve the calculation of the position, collects the data of pre-set landmark points in the guide path by inertial sensors, and realizes the landmark point recognition based on CNN. The article designs two sets of experiments to analyze the fusion localization results, and compared with the traditional PDR method. The actual experimental results show that the localization error of this method is less than 1m, which effectively reduces the cumulative error of PDR. Junhuai Li |
TrustCom | 2 |
| 2022 | Non-intrusive load monitoring method with inception structured CNN
Dong Ding 0002, Junhuai Li, Huaijun Wang, Kan Wang 0010, Ting Cao 0002 |
Appl. Intell. | 2 |
| 2022 | Single image reflection removal through multi-scale gradient refinementabstractAbstract Removing the undesired reflection layer from images taken through glass windows is an important yet challenging task. Many existing CNN‐based methods try to utilize the gradient as an important clue to guide the training and achieve better separation. But the scene depth of real‐world scenarios is usually uncontrollable, leading to the uncertainty of smooth level in the transmission and reflection layers, which makes it a great challenge to model the two layers in the gradient domain. This paper proposes a multi‐scale gradient refinement network to resolve this problem. First, it is suggested that even the two layers are usually partially smooth, their gradients can still be sharp in the down‐scaled samples. To this end, the separation is conducted at four different scales by minimizing the similarity of the two layers to boost the gradient sharpness prior. Second, it is considered that the separation performance of downscaled samples is usually superior to that of the high‐resolution images because of the sharper edges. For this reason, a cascade architecture is designed that takes the down‐scaled predictions to promote the high‐resolution decomposition stage‐by‐stage to recover the full‐resolution results. Besides, the scale‐wise memory mechanism is introduced into the prediction network to resolve the detail loss issue caused by the multi‐stage upscaling refinement process. The experimental results on benchmark datasets indicate that the new model surpasses several state‐of‐the‐art methods. Kuanhong Cheng, Yilan Li, Junhuai Li |
IET Image Process. | 6 |
| 2022 | A Fractional Integral and Fractal Dimension-Based Deep Learning Approach for Pavement Crack Detection in Transportation Service ManagementabstractWith artificial intelligence prevailing in intelligent transportation system, pavement crack detection with deep learning has aroused wide attentions in both academia and transportation sector. Nevertheless, it still remains a challenge to accomplish crack detection due to the complexity in pavement background. Motivated by latest advents in computer vision research, a fractional integral-based filtering method is advocated to remove pavement noise, and a fractal dimension estimation method has also emerged to present shape feature at pixel level, with the multi-scale feature architecture. Therefore, we try to propose a deep learning method, integrating fractional integral with fractal dimension, for crack detection in transportation service management. Firstly, the crack image is taken as input in the bottom-up architecture to extract fractal dimension on multi-scale levels, and a per-level feature unit is built to incorporate maps to make context information flow. Secondly, after fed into a convolutional filter for dimension resizing, all the resized feature maps are next fused at each level to comprise a group network. Finally, extensive experiments are executed on different crack datasets, exhibiting that the proposed method surpasses existing cutting-edge ones in terms of both generalizability and accuracy, with the benefits from not only fractional integral filtering, but also multi-scale fractal dimension features. Ting Cao 0002, Lei Liu 0031, Kan Wang 0010, Junhuai Li |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2020 | Wearable Sensor-Based Human Activity Recognition Using Hybrid Deep Learning TechniquesabstractHuman activity recognition (HAR) can be exploited to great benefits in many applications, including elder care, health care, rehabilitation, entertainment, and monitoring. Many existing techniques, such as deep learning, have been developed for specific activity recognition, but little for the recognition of the transitions between activities. This work proposes a deep learning based scheme that can recognize both specific activities and the transitions between two different activities of short duration and low frequency for health care applications. In this work, we first build a deep convolutional neural network (CNN) for extracting features from the data collected by sensors. Then, the long short-term memory (LTSM) network is used to capture long-term dependencies between two actions to further improve the HAR identification rate. By combing CNN and LSTM, a wearable sensor based model is proposed that can accurately recognize activities and their transitions. The experimental results show that the proposed approach can help improve the recognition rate up to 95.87% and the recognition rate for transitions higher than 80%, which are better than those of most existing similar models over the open HAPT dataset. Huaijun Wang, Junhuai Li, Ling Tian, Pengjia Tu, Ting Cao 0002, Kan Wang 0010, Shancang Li |
Secur. Commun. Networks | 3 |
| 2020 | Joint V2V-Assisted Clustering, Caching, and Multicast Beamforming in Vehicular Edge NetworksabstractAs an emerging type of Internet of Things (IoT), Internet of Vehicles (IoV) denotes the vehicle network capable of supporting diverse types of intelligent services and has attracted great attention in the 5G era. In this study, we consider the multimedia content caching with multicast beamforming in IoV-based vehicular edge networks. First, we formulate a joint vehicle-to-vehicle- (V2V-) assisted clustering, caching, and multicasting optimization problem, to minimize the weighted sum of flow cost and power cost, subject to the quality-of-service (QoS) constraints for each multicast group. Then, with the two-timescale setup, the intractable and stochastic original problem is decoupled at separate timescales. More precisely, at the large timescale, we leverage the sample average approximation (SAA) technique to solve the joint V2V-assisted clustering and caching problem and then demonstrate the equivalence of optimal solutions between the original problem and its relaxed linear programming (LP) counterpart; and at the small timescale, we leverage the successive convex approximation (SCA) method to solve the nonconvex multicast beamforming problem, whereby a series of convex subproblems can be acquired, with the convergence also assured. Finally, simulations are conducted with different system parameters to show the effectiveness of the proposed algorithm, revealing that the network performance can benefit from not only the power saving from wireless multicast beamforming in vehicular networks but also the content caching among vehicles. Kan Wang 0010, Junhuai Li, Meng Li 0007 |
Wirel. Commun. Mob. Comput. | 3 |
| 2018 | Design and Implementation of High Concurrent Communication Server in Health Monitoring SystemabstractMany institutions and researchers have conducted in-depth research on mobile/tele-health monitoring in recent decades. This paper designs and implements a high concurrent data communication server, improving the concurrent processing ability of health monitoring gateway, based on I/O Completion Port (IOCP) communication model. Meanwhile, improve the communication efficiency by elevating the aspects of duplex communication, session connection pool, dynamic cache. Finally, the paper analyzes the optimal server packet size and the best number of concurrent services, through the response time, throughput and other indicators. Junhuai Li, Jingfei Fu, Huaijun Wang, Lei Yu 0010 |
SERA | 1 |
| 2013 | A Localization and Tracking Approach with Sparse Reference TagsabstractIn traditional localization systems, it is required that moving object carries a device to transmit or receive signals, and then localization system is able to locate an object based on signal strength it received. In this paper, we propose a new passive localization and tracking approach based on RFID with sparse reference tags, which can estimate the location of moving objects by detecting and analyzing signal strength distribution of target area. We firstly construct a signal fluctuation ellipse model between RFID reader and tag through the experiments, and then present a localization method based on this model. Then a tracking method based on Hidden Markov Model (HMM) is proposed to predict the trajectory of an object in a passive localization system with sparse reference tag. The experimental results show that our method not only reduces the computation complexity and cost but also ensures the accuracy of localization and tracking. Junhuai Li, Lei Yu 0010, Hailing Liu |
MSN | 1 |
| 2013 | Resource virtualization methodology for on-demand allocation in cloud computing systems
Xiaojun Chen 0007, Jing Zhang 0009, Junhuai Li |
Serv. Oriented Comput. Appl. | 3 |
| 2012 | Task mapper and application-aware virtual machine scheduler oriented for parallel computingabstractWe design a task mapper TPCM for assigning tasks to virtual machines, and an application-aware virtual machine scheduler TPCS oriented for parallel computing to achieve a high performance in virtual computing systems. To solve the problem of mapping tasks to virtual machines, a virtual machine mapping algorithm (VMMA) in TPCM is presented to achieve load balance in a cluster. Based on such mapping results, TPCS is constructed including three components: a middleware supporting an application-driven scheduling, a device driver in the guest OS kernel, and a virtual machine scheduling algorithm. These components are implemented in the user space, guest OS, and the CPU virtualization subsystem of the Xen hypervisor, respectively. In TPCS, the progress statuses of tasks are transmitted to the underlying kernel from the user space, thus enabling virtual machine scheduling policy to schedule based on the progress of tasks. This policy aims to exchange completion time of tasks for resource utilization. Experimental results show that TPCM can mine the parallelism among tasks to implement the mapping from tasks to virtual machines based on the relations among subtasks. The TPCS scheduler can complete the tasks in a shorter time than can Credit and other schedulers, because it uses task progress to ensure that the tasks in virtual machines complete simultaneously, thereby reducing the time spent in pending, synchronization, communication, and switching. Therefore, parallel tasks can collaborate with each other to achieve higher resource utilization and lower overheads. We conclude that the TPCS scheduler can overcome the shortcomings of present algorithms in perceiving the progress of tasks, making it better than schedulers currently used in parallel computing. Jing Zhang 0009, Xiaojun Chen 0007, Junhuai Li |
J. Zhejiang Univ. Sci. C | 3 |
| 2011 | Resource management framework for collaborative computing systems over multiple virtual machines
Xiaojun Chen 0007, Jing Zhang 0009, Junhuai Li |
Serv. Oriented Comput. Appl. | 3 |