Minglei Shu

dblp:84/7712 · DBLP profile ↗
← Back
45ranked-venue papers
1as first author
34since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 8 since 2021Computer networks · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Facial Privacy Protection for Remote Photoplethysmography
abstract
Remote photoplethysmography (rPPG) has emerged as a crucial technology for contactless health monitoring, providing a convenient and non-invasive method to measure physiological signals from skin videos. Because face videos are commonly used for rPPG measurements, privacy concerns arise due to the inherent sensitivity of facial biometric data. Concerns about privacy breaches in facial video recordings have hindered telemedicine advancements and limited the creation of large-scale medical datasets, restricting the development of rPPG-based technologies. Additionally, the necessity to transmit and store rPPG videos in these applications necessitates video compression as an indispensable step. However, existing facial privacy protection techniques and video compression methods tend to degrade the rPPG signal in videos. To address these challenges, this study proposes a straightforward yet effective face anonymization module-a plug-and-play component employing spatial pixel redistribution algorithms to achieve: 1) eliminating identifiable biometric features while preserving the physiological information; 2) facilitating video compression by a macroblock reassembly strategy based on chromaticity clustering. Experiments on three rPPG datasets illustrate that the proposed method preserves physiological information in anonymized videos while effectively facilitating video compression.
Jieying Wang, Caifeng Shan, Shuwang Zhou, Minglei Shu
IEEE J. Biomed. Health Informatics5
2026 From Few to More: Scribble-Based Medical Image Segmentation via Masked Context Modeling and Continuous Pseudo Labels
abstract
Scribble-based weakly supervised segmentation methods have shown promising results in medical image segmentation, significantly reducing annotation costs. However, existing approaches often rely on auxiliary tasks to enforce semantic consistency and use hard pseudo labels for supervision, overlooking the unique challenges faced by models trained with sparse annotations. These models must predict pixel-wise segmentation maps from limited data, making it crucial to handle varying levels of annotation richness effectively. In this paper, we propose MaCo, a weakly supervised model designed for medical image segmentation, based on the principle of "from few to more." MaCo leverages Masked Context Modeling (MCM) and Continuous Pseudo Labels (CPL). MCM employs an attention-based masking strategy to perturb the input image, ensuring that the model's predictions align with those of the original image. CPL converts scribble annotations into continuous pixel-wise labels by applying an exponential decay function to distance maps, producing confidence maps that represent the likelihood of each pixel belonging to a specific category, rather than relying on hard pseudo labels. We evaluate MaCo on three public datasets, comparing it with other weakly supervised methods. Our results show that MaCo outperforms competing methods across all datasets, establishing a new record in weakly supervised medical image segmentation.
Zhisong Wang, Yiwen Ye, Ziyang Chen 0003, Minglei Shu, Yanning Zhang 0001, Yong Xia 0001
IEEE J. Biomed. Health Informatics4
2025 Disruptive Attacks on Face Swapping via Low-Frequency Perceptual Perturbations
abstract
Deepfake technology, driven by Generative Adversarial Networks (GANs), poses significant risks to privacy and societal security. Existing detection methods are predominantly passive, focusing on post-event analysis without preventing attacks. To address this, we propose an active defense method based on low-frequency perceptual perturbations to disrupt face-swapping manipulation, reducing the performance and naturalness of generated content. Unlike prior approaches that used low-frequency perturbations to impact classification accuracy, our method directly targets the generative process of deepfake techniques.We combine frequency and spatial domain features to strengthen defenses. By introducing artifacts through low-frequency perturbations while preserving high-frequency details, we ensure the output remains visually plausible. Additionally, we design a complete architecture featuring an encoder, a perturbation generator, and a decoder, leveraging discrete wavelet transform (DWT) to extract low-frequency components and generate perturbations that disrupt facial manipulation models. Experiments on CelebA-HQ and LFW demonstrate significant reductions in face-swapping effectiveness, improved defense success rates, and preservation of visual quality.
Mengxiao Huang, Minglei Shu, Shuwang Zhou, Zhaoyang Liu 0002
IJCNN2
2025 Correction of medical image segmentation errors through contrast learning with multi-branch
Tianlei Gao, Lei Lyu 0001, Nuo Wei, Tongze Liu, Minglei Shu
Eng. Appl. Artif. Intell.5
2025 Advancing few-shot rolling element bearing diagnostics with fine-grained similarity learning across variable working environments
Tianlei Gao, Xiaoyun Xie, Nuo Wei, Minglei Shu
Neurocomputing4
2025 UAE: Universal Anatomical Embedding on multi-modality medical images
Fan Bai 0008, Xiaofei Huo, Jia Ge, Jingjing Lu, Xianghua Ye, Minglei Shu, Ke Yan 0006, Yong Xia 0001
Medical Image Anal.7
2025 Physiological Information Preserving Video Compression for rPPG
abstract
Remote photoplethysmography (rPPG) has recently attracted much attention due to its non-contact measurement convenience and great potential in health care and computer vision applications. Early rPPG studies were mostly developed on self-collected uncompressed video data, which limited their application in scenarios that require long-distance real-time video transmission, and also hindered the generation of large-scale publicly available benchmark datasets. In recent years, with the popularization of high-definition video and the rise of telemedicine, the pressure of storage and real-time video transmission under limited bandwidth have made the compression of rPPG video inevitable. However, video compression can adversely affect rPPG measurements. This is due to the fact that conventional video compression algorithms are not specifically proposed to preserve physiological signals. Based on this, we propose a video compression scheme specifically designed for rPPG application. The proposed approach consists of three main strategies: 1) facial ROI-based computational resource reallocation; 2) rPPG signal preserving bit resource reallocation; and 3) temporal domain up- and down-sampling coding. UBFC-rPPG, ECG-Fitness, and a self-collected dataset are used to evaluate the performance of the proposed method. The results demonstrate that the proposed method can preserve almost all physiological information after compressing the original video to 1/60 of its original size. The proposed method is expected to promote the development of telemedicine and deep learning techniques relying on large-scale datasets in the field of rPPG measurement.
Jieying Wang, Caifeng Shan, Shuwang Zhou, Minglei Shu
IEEE J. Biomed. Health Informatics5
2024 A Multi-Lead Electrocardiogram Signal Classification Method Based on Temporal and Multi-View Contrastive Learning
abstract
The rise in wearable devices has led to the generation of a large amount of unlabeled electrocardiogram (ECG) data. Effectively utilising this data has been a challenge. One approach to address this issue is contrastive learning. However, most existing contrastive learning methods based on data augmentation primarily utilize the augmented electro-cardiogram (ECG) signals for comparison. These approaches have certain drawbacks. While the augmented ECG signals may help highlight certain features, they could also potentially mask important variations in the original signal, leading to the loss of crucial information. Moreover, solely relying on augmented data may lead to the model relying too heavily on specific variation patterns, overlooking the authentic features present in the original signal, resulting in poor generalization and robustness for downstream tasks. To address this issue, we propose a multi-lead electrocardiogram signal classification method based on temporal and multi-view contrastive learning. The method utilizes the time invariance of ECG signals and the multi-view information provided by the original ECG signals and augmented signals for contrastive learning. Through this approach, it not only addresses the limitations of relying solely on augmented signals but also leverages the prior knowledge that ECG signal categories remain stable over short periods of time to fully model the time context of ECG signals, capturing dynamic features. The integration of the temporal context comparison module and the multi-view comparison module significantly enhances the performance of downstream classification tasks. The outcomes of our experiments show that our method performs better on four datasets than previous approaches, which even exceeds supervised performance. When we pretrain on the SPH dataset and fine-tune on the PTB-XL dataset, our approach shows the best performance. Specifically, our method's AUROC exceeds that of the best baseline model by 4.9% and surpasses supervised performance by 1.3%.
Luyao Li, Hui Liu 0046, Shuwang Zhou, Zhaoyang Liu 0002, Minglei Shu
SMC5
2024 Grouping Boundary Proposals for Fast Interactive Image Segmentation
abstract
Geodesic models are known as an efficient tool for solving various image segmentation problems. Most of existing approaches only exploit local pointwise image features to track geodesic paths for delineating the objective boundaries. However, such a segmentation strategy cannot take into account the connectivity of the image edge features, increasing the risk of shortcut problem, especially in the case of complicated scenario. In this work, we introduce a new image segmentation model based on the minimal geodesic framework in conjunction with an adaptive cut-based circular optimal path computation scheme and a graph-based boundary proposals grouping scheme. Specifically, the adaptive cut can disconnect the image domain such that the target contours are imposed to pass through this cut only once. The boundary proposals are comprised of precomputed image edge segments, providing the connectivity information for our segmentation model. These boundary proposals are then incorporated into the proposed image segmentation model, such that the target segmentation contours are made up of a set of selected boundary proposals and the corresponding geodesic paths linking them. Experimental results show that the proposed model indeed outperforms state-of-the-art minimal paths-based image segmentation approaches.
Li Liu 0065, Da Chen 0002, Minglei Shu, Laurent D. Cohen
IEEE Trans. Image Process.3
2024 Slippage Estimation via Few-Shot Learning Based on Wheel-Ruts Images for Wheeled Robots on Loose Soil
abstract
When a wheeled mobile robot (WMR) runs on loose soil (such as the planetary rover on surface of the planet), its wheels generally slip or skip. Since the slippage of the wheel directly affects the motion control and safety, it becomes urgent to effectively estimate the slippage. In this paper, an intuitive slippage estimation method is proposed based on wheel-ruts images, which aims to reduce the number of extra sensors and take advantages of the visual information effectively. Since the image samples sometimes are difficult to collect, a few-shot learning method is employed using distribution propagation graph network with dilated causal convolution layer (DCC-DPGN). The dilated causal convolution (DCC) layer is adopted in ResNet block to expand the receptive field and obtain the sequence information of wheel-ruts, which makes the model training more efficient. The proposed model is verified in the test set of images collected in real scene, which shows the potential of the proposed algorithm in slippage estimation.
Chao Chen 0009, Shibin Su, Minglei Shu, Chong Di 0001, Weihua Li 0008, Pengyao Xu, Junlong Guo, Ruotong Wang 0003, Yinglong Wang 0001
IEEE Trans. Intell. Transp. Syst.3
2024 Context Matters: Distilling Knowledge Graph for Enhanced Object Detection
abstract
The human visual system is capable of not only recognizing individual objects but also comprehending the contextual relationship between them in real-world scenarios, making it highly advantageous for object detection. However, in practical applications, such contextual information is often not available. Previous attempts to compensate for this by utilizing cross-modal data such as language and statistics to obtain contextual priors have been deemed sub-optimal due to a semantic gap. To overcome this challenge, we present a seamless integration of context into an object detector through Knowledge Distillation. Our approach intuitively represents context as a knowledge graph, describing the relative location and semantic relevance of different visual concepts. Leveraging recent advancements in graph representation learning with Transformer, we exploit the contextual information among objects using edge encoding and graph attention. Specifically, each image region propagates and aggregates the representation from its highly similar neighbors to form the knowledge graph in the Transformer encoder. Extensive experiments and a thorough ablation study conducted on challenging benchmarks MS-COCO, Pascal VOC and LVIS demonstrate the superiority of our method.
Aijia Yang, Sihao Lin, Chung-Hsing Yeh, Minglei Shu, Yi Yang 0001, Xiaojun Chang
IEEE Trans. Multim.4
2024 Active Learning for Deep Visual Tracking
abstract
Convolutional neural networks (CNNs) have been successfully applied to the single target tracking task in recent years. Generally, training a deep CNN model requires numerous labeled training samples, and the number and quality of these samples directly affect the representational capability of the trained model. However, this approach is restrictive in practice, because manually labeling such a large number of training samples is time-consuming and prohibitively expensive. In this article, we propose an active learning method for deep visual tracking, which selects and annotates the unlabeled samples to train the deep CNN model. Under the guidance of active learning, the tracker based on the trained deep CNN model can achieve competitive tracking performance while reducing the labeling cost. More specifically, to ensure the diversity of selected samples, we propose an active learning method based on multiframe collaboration to select those training samples that should be and need to be annotated. Meanwhile, considering the representativeness of these selected samples, we adopt a nearest-neighbor discrimination method based on the average nearest-neighbor distance to screen isolated samples and low-quality samples. Therefore, the training samples' subset selected based on our method requires only a given budget to maintain the diversity and representativeness of the entire sample set. Furthermore, we adopt a Tversky loss to improve the bounding box estimation of our tracker, which can ensure that the tracker achieves more accurate target states. Extensive experimental results confirm that our active-learning-based tracker (ALT) achieves competitive tracking accuracy and speed compared with state-of-the-art trackers on the seven most challenging evaluation benchmarks. Project website: https://sites.google.com/view/altrack/.
Di Yuan 0002, Xiaojun Chang, Qiao Liu 0001, Yi Yang 0001, Minglei Shu, Zhenyu He 0001, Guangming Shi
IEEE Trans. Neural Networks Learn. Syst.6
2024 Energy and Spectrum Efficient Federated Learning via High-Precision Over-the-Air Computation
abstract
Federated learning (FL) enables mobile devices to collaboratively learn a shared prediction model while keeping data locally. However, there are two major research challenges to practically deploy FL over mobile devices: (i) frequent wireless updates of huge size gradients v.s. limited spectrum resources, and (ii) energy-hungry FL communication and local computing during training v.s. battery-constrained mobile devices. To address those challenges, in this paper, we propose a novel multi-bit over-the-air computation (M-AirComp) approach for spectrum-efficient aggregation of local model updates in FL and further present an energy-efficient FL design for mobile devices. Specifically, a high-precision digital modulation scheme is designed and incorporated in the M-AirComp, allowing mobile devices to upload model updates at the selected positions simultaneously in the multi-access channel. Moreover, we theoretically analyze the convergence property of our FL algorithm. Guided by FL convergence analysis, we formulate a joint transmission probability and local computing control optimization, aiming to minimize the overall energy consumption (i.e., iterative local computing + multi-round communications) of mobile devices in FL. Extensive simulation results show that our proposed scheme outperforms existing ones in terms of spectrum utilization, energy efficiency, and learning accuracy.
Liang Li 0021, Chenpei Huang, Dian Shi, Hao Wang 0022, Xiangwei Zhou, Minglei Shu, Miao Pan
IEEE Trans. Wirel. Commun.6
2023 TransFS: Face Swapping Using Transformer
abstract
This paper proposes a Transformer based face swapping model, namely, TransFS. The proposed model mainly solves two current problems of face swapping: 1) the face swapping result does not fully preserve pose and expression of the target face as expected; 2) most of the existing models fail to accomplish high-quality face swapping on high-resolution images. To address these two challenges, we first propose a Cross- Window Face Encoder based on Swin Transformer that learns rich facial features including poses and expressions. Then, we devise an Identity Generator to reconstruct high-resolution images of specific identity with high quality while utilizing the Transformer attention mechanism to increase identity information retention. Finally, a Face Conversion Module is proposed to transform the source identity reconstructed image into the target face image to synthesize the final face swapping result while maintaining the details of pose and expression of the target face. Through extensive experiments, our method not only accomplishes face swapping for low-resolution images with arbitrary identities, but also accomplishes face swapping for high-resolution images. Furthermore, our method achieves the state-of-the-art performance in pose and expression controls compared to other methods.
Tianyi Wang 0006, Anming Dong, Minglei Shu
FG4
2023 Dynamic Facial Expression Recognition in Unconstrained Real-World Scenarios Leveraging Dempster-Shafer Evidence Theory
Tianyi Wang 0006, Shuwang Zhou, Minglei Shu
ICANN (2)4
2023 A lightweight 2-D CNN model with dual attention mechanism for heartbeat classification
Hongfu Xie, Shuwang Zhou, Tianlei Gao, Minglei Shu
Appl. Intell.5
2023 Learning automata-accelerated greedy algorithms for stochastic submodular maximization
Chong Di 0001, Fangqi Li 0001, Pengyao Xu, Ying Guo 0004, Chao Chen 0009, Minglei Shu
Knowl. Based Syst.6
2023 Geodesic Models With Convexity Shape Prior
abstract
The minimal geodesic models established upon the eikonal equation framework are capable of finding suitable solutions in various image segmentation scenarios. Existing geodesic-based segmentation approaches usually exploit image features in conjunction with geometric regularization terms, such as euclidean curve length or curvature-penalized length, for computing geodesic curves. In this paper, we take into account a more complicated problem: finding curvature-penalized geodesic paths with a convexity shape prior. We establish new geodesic models relying on the strategy of orientation-lifting, by which a planar curve can be mapped to an high-dimensional orientation-dependent space. The convexity shape prior serves as a constraint for the construction of local geodesic metrics encoding a particular curvature constraint. Then the geodesic distances and the corresponding closed geodesic paths in the orientation-lifted space can be efficiently computed through state-of-the-art Hamiltonian fast marching method. In addition, we apply the proposed geodesic models to the active contours, leading to efficient interactive image segmentation algorithms that preserve the advantages of convexity shape prior and curvature penalization.
Da Chen 0002, Jean-Marie Mirebeau, Minglei Shu, Xue-Cheng Tai, Laurent D. Cohen
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Curvilinear Structure Tracking Based on Dynamic Curvature-penalized Geodesics
Li Liu 0065, Shuwang Zhou, Minglei Shu, Laurent D. Cohen, Da Chen 0002
Pattern Recognit.4
2023 FGSQA-Net: A Weakly Supervised Approach to Fine-Grained Electrocardiogram Signal Quality Assessment
abstract
OBJECTIVE: Due to the lack of fine-grained labels, current research can only evaluate the signal quality at a coarse scale. This article proposes a weakly supervised fine-grained electrocardiogram (ECG) signal quality assessment method, which can produce continuous segment-level quality scores with only coarse labels. METHODS: A novel network architecture, i.e. FGSQA-Net, is developed for signal quality assessment, which consists of a feature shrinking module and a feature aggregation module. Multiple feature shrinking blocks, which combine residual CNN block and max pooling layer, are stacked to produce a feature map corresponding to continuous segments along the spatial dimension. Segment-level quality scores are obtained by feature aggregation along the channel dimension. RESULTS: The proposed method was evaluated on two real-world ECG databases and one synthetic dataset. Our method produced an average AUC value of 0.975, which outperforms the state-of-the-art beat-by-beat quality assessment method. The results are visualized for 12-lead and single-lead signals over a granularity from 0.64 to 1.7 seconds, demonstrating that high-quality and low-quality segments can be effectively distinguished at a fine scale. CONCLUSION: FGSQA-Net is flexible and effective for fine-grained quality assessment for various ECG recordings and is suitable for ECG monitoring using wearable devices. SIGNIFICANCE: This is the first study on fine-grained ECG quality assessment using weak labels and can be generalized to similar tasks for other physiological signals.
Hui Liu 0046, Tianlei Gao, Zhaoyang Liu 0002, Minglei Shu
IEEE J. Biomed. Health Informatics4
2022 A Robust Lightweight Deepfake Detection Network Using Transformers
Tianyi Wang 0006, Minglei Shu, Yinglong Wang 0001
PRICAI (1)3
2022 Comment recommendation based on graph bidirectional aggregation networks
Muchen Wang, Chao Chen 0009, Minglei Shu
Expert Syst. Appl.4
2022 IoT Device Friendly and Communication-Efficient Federated Learning via Joint Model Pruning and Quantization
abstract
Federated learning (FL) through its novel applications and services has enhanced its presence as a promising tool in the Internet of Things (IoT) domain. Specifically, in a multiaccess edge computing setup with a host of IoT devices, FL is most suitable since it leverages distributed client data to train high-performance deep learning (DL) models while keeping the data private. However, the underlying deep neural networks (DNNs) are huge, preventing its direct deployment onto resource-constrained computing and memory-limited IoT devices. Besides, frequent exchange of model updates between the central server and clients in FL could result in a communication bottleneck. To address these challenges, in this article, we introduce GWEP, a model compression-based FL method. It utilizes joint quantization and model pruning to reap the benefits of DNNs while meeting the capabilities of resource-constrained devices. Consequently, by reducing the computational, memory, and network footprint of FL, the low-end IoT devices may be able to participate in the FL process. In addition, we provide theoretical guarantees of FL convergence. Through empirical evaluations, we demonstrate that our approach significantly outperforms the baseline algorithms by being up to 10.23 times faster with 11 times lesser communication rounds, while achieving high-model compression, energy efficiency, and learning performance.
Pavana Prakash, Jiahao Ding, Rui Chen 0026, Xiaoqi Qin, Minglei Shu, Qimei Cui, Yuanxiong Guo, Miao Pan
IEEE Internet Things J.5
2022 Power-Efficient Data Collection Scheme for AUV-Assisted Magnetic Induction and Acoustic Hybrid Internet of Underwater Things
abstract
Power efficiency is a big concern in the Internet of Underwater Things (IoUT). The power consumption of underwater acoustic communications is typically in the scale of watts, which may drain the battery of underwater devices quickly. Whereas, the power consumption of underwater magnetic induction (MI) wireless communications is in the scale of milliwatt. Therefore, this article devotes to combine the underwater MI and acoustic communications to form a power-efficient underwater hybrid wireless network. Specifically, we investigate the power-efficient autonomous underwater vehicle (AUV) data collection schemes in an underwater MI and acoustic hybrid sensor network. We propose an alternating anchor nodes selection and flow routing (AANSFR) AUV data collection method, which alternately optimizes the AUV path planning and network data flow routing. The simulation results show that the proposed hybrid data collection scheme can significantly prolong the lifespan of underwater sensor networks.
Debing Wei, Chenpei Huang, Xuanheng Li, Bin Lin 0001, Minglei Shu, Jie Wang 0003, Miao Pan
IEEE Internet Things J.5
2022 Dynamic knowledge graph reasoning based on deep reinforcement learning
Hao Liu 0066, Shuwang Zhou, Changfang Chen, Tianlei Gao, Jiyong Xu, Minglei Shu
Knowl. Based Syst.6
2022 Trajectory Grouping With Curvature Regularization for Tubular Structure Tracking
abstract
Tubular structure tracking is a crucial task in the fields of computer vision and medical image analysis. The minimal paths-based approaches have exhibited their strong ability in tracing tubular structures, by which a tubular structure can be naturally modeled as a minimal geodesic path computed with a suitable geodesic metric. However, existing minimal paths-based tracing approaches still suffer from difficulties such as the shortcuts and short branches combination problems, especially when dealing with the images involving complicated tubular tree structures or background. In this paper, we introduce a new minimal paths-based model for minimally interactive tubular structure centerline extraction in conjunction with a perceptual grouping scheme. Basically, we take into account the prescribed tubular trajectories and curvature-penalized geodesic paths to seek suitable shortest paths. The proposed approach can benefit from the local smoothness prior on tubular structures and the global optimality of the used graph-based path searching scheme. Experimental results on both synthetic and real images prove that the proposed model indeed obtains outperformance comparing with the state-of-the-art minimal paths-based tubular structure tracing algorithms.
Li Liu 0065, Da Chen 0002, Minglei Shu, Huazhong Shu, Michel Pâques, Laurent D. Cohen
IEEE Trans. Image Process.3
2022 To Talk or to Work: Dynamic Batch Sizes Assisted Time Efficient Federated Learning Over Future Mobile Edge Devices
abstract
The coupling of federated learning (FL) and multi-access edge computing (MEC) has the potential to foster numerous applications. However, it poses great challenges to train FL fast enough with limited communication and computing resources of mobile edge devices. Motivated by recent development in ultra fast wireless transmissions and promising advances in artificial intelligence (AI) computing hardware of mobile devices, in this paper, we propose a time efficient FL over future mobile edge devices, called dynamic batch sizes assisted federated learning (DBFL) with convergence guarantee. The DBFL allows batch sizes to increase dynamically during training, which can unleash the computing potential of GPU’s parallelism for on- device training and effectively leverage the fast wireless transmissions (WiFi-6, 5G, 6G, etc.) of mobile edge devices. Furthermore, based on the derived DBFL’s convergence bound, we develop a batch size control scheme to minimize the total time consumption of FL over mobile edge devices, which trade-offs the “talking”, i.e., communication time, and “working”, i.e., computing time, by adjusting the incremental factor appropriately. Extensive simulations are conducted to validate the effectiveness of our proposed DBFL algorithm and demonstrate that our scheme outperforms existing time efficient FL approaches in terms of the total time consumption in various settings.
Dian Shi, Liang Li 0021, Maoqiang Wu, Minglei Shu, Rong Yu 0001, Miao Pan, Zhu Han 0001
IEEE Trans. Wirel. Commun.4
2021 To Talk or to Work: Delay Efficient Federated Learning over Mobile Edge Devices
abstract
Federated learning (FL), an emerging distributed machine learning paradigm, in conflux with edge computing is a promising area with novel applications over mobile edge devices. In FL, since mobile devices collaborate to train a model based on their own data under the coordination of a central server by sharing just the model updates, training data is maintained private. However, without the central availability of data, computing nodes need to communicate the model updates often to attain convergence. Hence, the local computation time to create local model updates along with the time taken for transmitting them to and from the server result in a delay in the overall time. Furthermore, unreliable network connections may obstruct an efficient communication of these updates. To address these, in this paper, we propose a delay-efficient FL mechanism that reduces the overall time (consisting of both the computation and communication latencies) and communication rounds required for the model to converge. Exploring the impact of various parameters contributing to delay, we seek to balance the trade-off between wireless communication (to talk) and local computation (to work). We formulate a relation with overall time as an optimization problem and demonstrate the efficacy of our approach through extensive simulations.
Pavana Prakash, Jiahao Ding, Maoqiang Wu, Minglei Shu, Rong Yu 0001, Miao Pan
GLOBECOM4
2021 A New Tubular Structure Tracking Algorithm Based On Curvature-Penalized Perceptual Grouping
abstract
In this paper, we propose a new minimal path-based framework for minimally interactive tubular structure tracking in conjunction with a perceptual grouping scheme. The minimal path models have shown great advantages in tubular structures tracing. However, they suffer from shortcuts or short branches combination problems especially in the case of tubular network with complicated structures or background. Thus, we utilize the curvature-penalized minimal paths and the prescribed tubular trajectories to seek the desired shortest path. The proposed approach benefits from the local smoothness prior on tubular structures and the global optimality of the graph-based path searching scheme. Experimental results on synthetic and real images prove that the proposed model indeed obtains outperformance to state-of-the-art minimal path-based algorithms.
Li Liu 0065, Da Chen 0002, Minglei Shu, Huazhong Shu, Laurent D. Cohen
ICASSP3
2021 DNANet: Dense Nested Attention Network for Single Image Dehazing
abstract
In this paper, we propose an innovative approach, called Dense Nested Attention Network (DNANet), to directly restore a clear image from a hazy image with a new topology of connection paths. Firstly, through dense nested connections from inside to outside, the DNANet can fuse both shallow and deep features from fine to coarse, then strengthen the feature propagation and reuse to a large extent. We use stacked dilated convolutions, as the basic operation, to alleviate the shortcomings of the traditional context information aggregation methods. Secondly, we examine the weakness of skipping connections by reasoning the existence of residual haze from the shallow to deep layers in the neural network. To address this problem, we use the attention mechanism to filter out the output of residual haze by capturing the information relations on the entire skip feature maps. Thirdly, we introduce an adjustable loss constraint on each block of the outermost nested structure to gather more accurate features. The result demonstrates that DNANet outperforms state-of-the-art methods by a large margin on the benchmark datasets in extensive experiments.
Dongdong Ren, Minglei Shu
ICASSP4
2021 SQuaFL: Sketch-Quantization Inspired Communication Efficient Federated Learning
Pavana Prakash, Jiahao Ding, Minglei Shu, Junyi Wang 0002, Wenjun Xu 0001, Miao Pan
SEC3
2021 A multi-stage denoising framework for ambulatory ECG signal based on domain knowledge and motion artifact detection
Xiaoyun Xie, Hui Liu 0046, Minglei Shu, Qing Zhu 0005, Anpeng Huang, Xiangpu Kong, Yinglong Wang 0001
Future Gener. Comput. Syst.3
2021 A Generalized Asymmetric Dual-Front Model for Active Contours and Image Segmentation
abstract
The Voronoi diagram-based dual-front scheme is known as a powerful and efficient technique for addressing the image segmentation and domain partitioning problems. In the basic formulation of existing dual-front approaches, the evolving contour can be considered as the interfaces of adjacent Voronoi regions. Among these dual-front models, a crucial ingredient is regarded as the geodesic metrics by which the geodesic distances and the corresponding Voronoi diagram can be estimated. In this paper, we introduce a new dual-front model based on asymmetric quadratic metrics. These metrics considered are built by the integration of the image features and a vector field derived from the evolving contour. The use of the asymmetry enhancement can reduce the risk for the segmentation contours being stuck at false positions, especially when the initial curves are far away from the target boundaries or the images have complicated intensity distributions. Moreover, the proposed dual-front model can be applied for image segmentation in conjunction with various region-based homogeneity terms. The numerical experiments on both synthetic and real images show that the proposed dual-front model indeed achieves encouraging results.
Da Chen 0002, Jack A. Spencer, Jean-Marie Mirebeau, Ke Chen 0002, Minglei Shu, Laurent D. Cohen
IEEE Trans. Image Process.5
2021 Geodesic Paths for Image Segmentation With Implicit Region-Based Homogeneity Enhancement
abstract
Minimal paths are regarded as a powerful and efficient tool for boundary detection and image segmentation due to its global optimality and the well-established numerical solutions such as fast marching method. In this paper, we introduce a flexible interactive image segmentation model based on the Eikonal partial differential equation (PDE) framework in conjunction with region-based homogeneity enhancement. A key ingredient in the introduced model is the construction of local geodesic metrics, which are capable of integrating anisotropic and asymmetric edge features, implicit region-based homogeneity features and/or curvature regularization. The incorporation of the region-based homogeneity features into the metrics considered relies on an implicit representation of these features, which is one of the contributions of this work. Moreover, we also introduce a way to build simple closed contours as the concatenation of two disjoint open curves. Experimental results prove that the proposed model indeed outperforms state-of-the-art minimal paths-based image segmentation approaches.
Da Chen 0002, Xinxin Zhang 0004, Minglei Shu, Laurent D. Cohen
IEEE Trans. Image Process.4
2020 Detection of Muscle Fatigue by Fusion of Agonist and Synergistic Muscle sEMG Signals
abstract
Muscle fatigue detection has a wide range of applications in the field of rehabilitation medicine. The existing methods only collect a single agonist signal for detection, whose accuracy and real-time detection are usually poor. To handle this issue, in this study collects a dataset containing both agonist and synergistic muscle surface electromyography (sEMG) signals, by analyzing the synergistic working principle of muscles. Through the fusion processing of different muscle group signals, the significance of detection of muscle state changes is improved. Moreover, the impact of signal non-stationarity and non-linearity on fatigue detection is reduced. Based on the collected dataset, a multichannel fusion recurrent attention network (MFRANet) is proposed. First, MFRANet enhances local anti-interference ability by fusing multi-channel EMG signals and reduces the impact of single channel signal noise on the overall detection performance. Second, MFRANet analyzes the signal from two dimensions, namely time domain and space domain. A gating mechanism is used to enhance the complex time correlation between channels, and an attention mechanism is employed to reconstruct the nonlinear relationship between channels, thus improving the generalization. Experiments show that the proposed signal fusion method of agonist and synergistic muscle sEMG signals significantly improves the accuracy of muscle fatigue detection, as well as reducing processing time compared to traditional machine learning methods.
Maoheng Li, Minglei Shu
CBMS3
2020 Epileptic Seizure Prediction Based on Region Correlation of EEG Signal
abstract
The existing methods of epileptic seizure prediction usually analyze the electroencephalogram (EEG) signals in the time domain, frequency domain or time-frequency domain. Although some good results have been achieved, the research and utilization of spatial information is still insufficient. Moreover, some studies extracted different features for different patients and achieved good results, but these methods are not universal and robust. Different from the previous methods, this paper propose a new feature processing method of EEG signal. All electrode signals on the scalp are considered as a whole, and fusing data from different regions to obtain spatial information. Then the correlation of first derivatives is used to obtain fluctuation information of signal caused by epilepsy, which further enlarge difference of signal in different seizures stages. In addition, we also design a post-processing strategy, which uses time-series information to rectify prediction results, so that the final result is more accurate. Finally, experimental results from the CHBMIT dataset show effectiveness of proposed method and strategy, while the extensive result confirms that our method is superior to several state-of-the-art methods in recent years.
Xuefei Liu, Minglei Shu
CBMS3
2020 Not All Areas Are Equal: A Novel Separation-Restoration-Fusion Network for Image Raindrop Removal
abstract
Abstract Detecting and removing raindrops from an image while keeping the high quality of image details has attracted tremendous studies, but remains a challenging task due to the inhomogeneity of the degraded region and the complexity of the degraded intensity. In this paper, we get rid of the dependence of deep learning on image‐to‐image translation and propose a separation‐restoration‐fusion network for raindrops removal. Our key idea is to recover regions of different damage levels individually, so that each region achieves the optimal recovery result, and finally fuse the recovered areas. In the region restoration module, to complete the restoration of a specific area, we propose a multi‐scale feature fusion global information aggregation attention network to achieve global to local information aggregation. Besides, we also design an inside and outside dense connection dilated network, to ensure the fusion of the separated regions and the fine restoration of the image. The qualitatively and quantitatively evaluations are conducted to evaluate our method with the latest existing methods. The result demonstrates that our method outperforms state‐of‐the‐art methods by a large margin on the benchmark datasets in extensive experiments.
Dongdong Ren, Minglei Shu
Comput. Graph. Forum4
2020 SCGA-Net: Skip Connections Global Attention Network for Image Restoration
abstract
Abstract Deep convolutional neural networks (DCNN) have shown their advantages in the image restoration tasks. But most existing DCNN‐based methods still suffer from the residual corruptions and coarse textures. In this paper, we propose a general framework “Skip Connections Global Attention Network” to focus on the semantics delivery from shallow layers to deep layers for low‐level vision tasks including image dehazing, image denoising, and low‐light image enhancement. First of all, by applying dense dilated convolution and multi‐scale feature fusion mechanism, we establish a novel encoder‐decoder network framework to aggregate large‐scale spatial context and enhance feature reuse. Secondly, the solution we proposed for skipping connection uses attention mechanism to constraint information, thereby enhancing the high‐frequency details of feature maps and suppressing the output of corruptions. Finally, we also present a novel attention module dubbed global constraint attention, which could effectively captures the relationship between pixels on the entire feature maps, to obtain the subtle differences among pixels and produce an overall optimal 3D attention maps. Extensive experiments demonstrate that the proposed method achieves significant improvements over the state‐of‐the‐art methods in image dehazing, image denoising, and low‐light image enhancement.
Dongdong Ren, Minglei Shu
Comput. Graph. Forum4
2019 Low Dose CT Image Denoising Using Multi-level Feature Fusion Network and Edge Constraints
abstract
Low-dose computed tomography image denoising is a challenging task that has been studied by many researchers. Current denoising methods based on deep learning tend to produce a blur effect on the final results, especially at high noise levels, which are prone to over-smoothed edges and loss of details. In this paper, we propose a deep learning approach based on deep convolutional and edge constraints to mitigate these problems. Firstly, to avoid the loss of shallow layers details while obtaining semantically-richer features information, we use dilated convolution instead of standard convolution, and fusion feature maps of different levels to aggregate information from different receptive field. Secondly, in order to improve the network's ability to distinguish between noise and image content, we have designed an attention block to adaptively recalibrate the information relationship of the fusion feature maps. Finally, we incorporate edge prior knowledge into LDCT image denoising task, guiding the network to pay more attention on texture and structure information by edge constraints loss. Extensive experiments demonstrate that the proposed method achieves significant improvements over the state-of-the-art methods.
Dongdong Ren, Lingli Li, Haiwei Pan, Minglei Shu
BIBM5
2018 A Transmission Power Control Algorithm for Wireless Body Area Networks
abstract
Energy efficiency is a key issue for wireless sensor nodes, especially for wireless body area networks (WBANs) that operate near the human body or in the human body. Aiming at the problem that WBAN system still has too fast energy consumption, we propose a ZigBee star network model with multiple sensing nodes as end nodes, and design an adaptive transmission power correction control algorithm with adjustment factors to select the appropriate transmit power to reduce the energy consumption. Experiments show that the proposed power control algorithm reduces the overall energy consumption by reasonably controlling the transmission power in the ZigBee star network model.
Zhuoran Zhengl, Xiangwei Zheng 0001, Jie Tian 0003, Minglei Shu
CSCWD4
2017 A Study on the Second Order Statistics of \kappa κ - \mu μ Fading Channels
Changfang Chen, Minglei Shu, Yinglong Wang 0001, Nuo Wei
WASA2
2016 IRIBE: Intrusion-resilient identity-based encryption
Jia Yu 0003, Rong Hao, Huawei Zhao, Minglei Shu, Jianxi Fan
Inf. Sci.4
2015 Physiological-Signal-Based Key Negotiation Protocols for Body Sensor Networks: A Survey
abstract
Body sensor networks (BSNs) are deployed around the human body to measure and process physiological signals in real time, which have wide application prospects in intelligent healthcare. Physiological signals measured and processed by BSNs involve individual privacy, thus security mechanisms must be developed to secure BSNs, and therein adoption of key negotiation protocols is fundamental. Due to stringently limited operation resources, BSNs require these protocols to be low-energy and highly efficient. Recent development has discovered that certain physiological signals can be used for efficiently negotiating common keys among biosensor nodes. These signals and fuzzy technology are used to design lightweight key negotiation protocols, and many solutions have been proposed. In this paper, we explore and classify these solutions, and evaluate their performance by analyzing their merits and demerits. Finally, we present open research issues that should be solved in the future.
Huawei Zhao, Ruzhi Xu, Minglei Shu, Jiankun Hu
ISADS3
2015 Planning the obstacle-avoidance trajectory of mobile anchor in 3D sensor networks
Minglei Shu, Huanqing Cui, Yinglong Wang 0001, Cheng-Xiang Wang 0001
Sci. China Inf. Sci.1
2015 Hierarchical Adaptive Path-Tracking Control for Autonomous Vehicles
abstract
This paper presents a hierarchical controller for an autonomous vehicle to track a reference path in the presence of uncertainties in both tire-road condition and external disturbance. The hierarchical control architecture consists of three layers: high, low, and intermediate levels. The upper-layer module deals with the vehicle motion control objective, which generates the desired longitudinal/lateral forces and yaw moment. The low-level module handles the braking control for each wheel based on the wheel slip dynamics. The intermediate-level controller generates the longitudinal slip reference for the low-level brake control module and the front-wheel steering angles. To cope with the unknown and nonuniform road condition parameters appearing in the actuator models, an adaptive law is designed for each wheel, and the convergence of the adaptive parameters is guaranteed under a certain persistency-of-excitation condition. The stability of the integrated control system is analyzed by utilizing a Lyapunov function approach. Simulation results are included to illustrate the proposed control scheme.
Changfang Chen, Yingmin Jia, Minglei Shu, Yinglong Wang 0001
IEEE Trans. Intell. Transp. Syst.3