Mingkai Chen 0001

dblp:174/9971-1 · DBLP profile ↗
← Back
32ranked-venue papers
6as first author
22since 2021 · last 2026
0000-0002-8826-2587ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 20 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Explainable Artificial Intelligence Enhance Image Semantic Communication System in 6G-IoT
abstract
The emerging 6G-IoT paradigm is driving communication toward intelligent services, semantic communication enables efficient semantic sharing via artificial intelligence (AI), significantly boosting communication efficiency. However, current semantic systems suffer from black-box decision-making, while existing explainable artificial intelligence (XAI) methods face two key challenges: explicability granularity mismatch and closed-loop optimization gap. To address these, we propose a semantic communication framework integrated with XAI (XAI-SCS). Specifically, we first design an explainable semantic codec architecture enhanced by Kolmogorov–Arnold Networks (KAN), where traditional fixed activation functions are replaced with learnable and parameterized ones, enabling function-level visualization to improve model explainability. Second, we develop an explainable semantic transmission module driven by contrastive learning that enhances the robustness of semantic transmission, and incorporating a semantic separability metric to quantify channel impacts on semantic integrity. Third, we introduced a KAN-enhanced causal semantic decoder, which integrates counterfactual interventions to generate pixel-level difference maps. We also propose a contrastive explanation consistency metric to evaluate the sensitivity of key features, enhancing the quality of the reconstruction. The experimental results show that our approaches enhance explainability across the entire decision process, achieve a significant accuracy improvement of up to 55% on the CIFAR-10 dataset with a bandwidth compression ratio of 1/25, and also obtain competitive image reconstruction quality in increasing compression levels. The source code is publicly available at: https://github.com/guyuangui/XAI-SCS.git.
Mingkai Chen 0001, Yuangui Gu, Xiaoming He 0004, Feng Huang 0007, Lei Wang 0009
IEEE Internet Things J.1
2026 Covertness and Reliability Analysis of UAV Relay Communications Aided by a Friendly Jammer Under Co-Channel Interference
abstract
This paper investigates the covertness and reliability of unmanned aerial vehicle (UAV) relay communications assisted by a friendly jammer in the presence of co-channel interference (CCI). In the considered scenario, communication from a source to a destination is relayed by a UAV, while a malicious node attempts to monitor the transmission behaviors of both the source and the UAV. To enhance the system’s covertness, we propose a friendly jammer aided covert communication (FJACC) scheme, which leverages a friendly jammer to degrade the detection capability of the malicious monitoring node by sending artificial noise. We analyze the covertness and reliability of the FJACC scheme by deriving closed-form expressions of the outage probability and detection error probability under CCI, considering that the wireless channels are characterized by Nakagami-mfading for UAV-to-ground links and Rayleigh fading for terrestrial links. Numerical results confirm the superior covertness of the proposed FJACC scheme over traditional non-friendly jammer aided transmission scheme. Furthermore, increasing the number of co-channel interferers or their transmit power exhibits a dual effect: it enhances covertness at the expense of reliability.
Bin Li 0022, YuLong Zou, Mingkai Chen 0001, Dingcheng Li, Yunxiao Tang, Peishun Yan, Weifeng Cao
IEEE Internet Things J.3
2026 Initial Access Beam Management Framework for LEO Satellite Networks Integrated With 5G NR
abstract
This paper considers a Low-Earth-Orbit (LEO) satellite communication system integrated with 5G New Radio, which performs conventionally initial access beam management through multiple measurements and reports. However, such management scheme may lead to outdated measurement results and large terminal access delay, in the presence of significant propagation losses, high dynamics, and limited number of antennas on LEO satellites. Addressing these issues, we propose an initial access beam management framework based on the dynamic coordination of signaling and service beams. This framework includes the optimization of the beam alignment process and the analysis of the coordination mechanism of signaling and service beams with a queueing model. Specifically, this queueing model is further used to assist in modeling an optimization problem with respect to the number and direction of signaling and service beams, to achieve the dynamic balance between coverage capability and service efficiency. Furthermore, a Lyapunov based initial access beam management (LYP-IABM) algorithm is specially designed to find the optimized solution of the modeled problem in a low complexity and online manner. The mathematical derivation is also provided to prove the correctness of the designed algorithm. Simulation results show that, the proposed online LYP-IABM algorithm not only outperforms the traditional offline algorithms, but also demonstrates the superiority comparing with the online benchmarking algorithms, in terms of reducing the time complexity by more than 90%.
Yanjin Zhang, Xiaojin Ding, Mingkai Chen 0001
IEEE Internet Things J.4
2026 GAI-Enabled Task-Driven Semantic Communication for Surveillance Video
abstract
With the development of surveillance cameras, more bandwidth is required to transmit surveillance videos. Since surveillance videos contain a large amount of redundant information, it causes a waste of bandwidth. Meanwhile, previous video compression methods with the fixed compression standards are unable to handle asymmetric information effectively. To address these problems, we propose Task-driven Semantic Communication with Unsupervised Semantic Segmentation (TSCUSS) for surveillance video assisted by Generative Artificial Intelligence (GAI), to improve efficiency. First, at the transmitter, we segment the videos into the foreground semantic and background models. Second, in the transmission side, we transmit the extracted semantic information in two-stage semantic communication, which greatly reduces redundant information. Third, at the receiver, we merge the foreground and background semantic models through the diffusion model to recover the original semantic content. Finally, our experiment shows that our method not only achieves 78.34% average video compression rate and improves bandwidth utilization, but also dominates in both semantic segmentation accuracy and generative foreground background merge similarity.
Mingkai Chen 0001, Lei Wang 0009, Wael Bazzi, Kezhi Wang, Shahid Mumtaz
IEEE Trans. Commun.1
2026 Toward Emotion-Preserving Speech Semantic Communication in Affective Social Systems
abstract
With the rapid development of artificial intelligence (AI) and ubiquitous connectivity, speech communication is becoming increasingly important in achieving intelligent interactions between humans, machines, and objects in computational social systems. However, current neural network-based semantic communication frameworks primarily focus on transmitting semantic information while largely overlooking the emotion features in speech communication, which is vital for naturalness and effective interaction in computational social systems. In this article, we propose an emotion-enhanced speech semantic communication system, which effectively enhances the expressiveness and robustness of emotion AI for speech communication. First, we propose an emotion fusion encoding module at the transmitter, where features are dynamically fused via attention mechanisms and subsequently encoded through a channel encoder. Then, we introduce an emotion orthogonal decoding module at the receiver, which reconstructs the fused features via a channel decoder followed by an orthogonally constrained disentanglement network. In addition, a conditional diffusion model guided by emotion features reconstructs high-fidelity speech with enriched emotional expressiveness. Finally, experimental evaluations demonstrate that the proposed framework significantly improves both the bit error rate and the mean opinion score over state-of-the-art models. Furthermore, the system achieves notable reductions in transmission dimensionality.
Taojie Zhu, Mingkai Chen 0001, Lei Wang 0009, M. Shamim Hossain
IEEE Trans. Comput. Soc. Syst.2
2026 Foundation Model Empowered Real-Time Video Conference With Semantic Communications
abstract
With the development of real-time video conferences, interactive multimedia services have proliferated, leading to a surge in traffic. Interactivity becomes one of the main features on future multimedia services, which brings a new challenge to Computer Vision (CV) for communications. In addition, many directions for CV in video, like recognition, understanding, saliency segmentation, coding, and so on, do not satisfy the demands of the multiple tasks of interactivity without integration. Meanwhile, with the rapid development of the foundation models, we apply task-oriented semantic communications to handle them. Therefore, we propose a novel framework, called Real-Time Video Conference with Foundation Model (RTVCFM), to satisfy the requirement of interactivity in the multimedia service. Firstly, at the transmitter, we perform the causal understanding and spatiotemporal decoupling on interactive videos, with the Video Time-Aware Large Language Model (VTimeLLM), Iterated Integrated Attributions (IIA) and Segment Anything Model 2 (SAM2), to accomplish the video semantic segmentation. Secondly, in the transmission, we propose a two-stage semantic transmission optimization driven by Channel State Information (CSI), which is also suitable for the weights of asymmetric semantic information in real-time video, so that we achieve a low bit rate and high semantic fidelity in the video transmission. Thirdly, at the receiver, RTVCFM provides multidimensional fusion with the whole semantic segmentation by using the Diffusion Model for Foreground Background Fusion (DMFBF), and then we reconstruct the video streams. Finally, the simulation result demonstrates that RTVCFM can achieve a compression ratio as high as 95.6%, while it guarantees high semantic similarity of 98.73% in Multi-Scale Structural Similarity Index Measure (MS-SSIM) and 98.35% in Structural Similarity (SSIM), which shows that the reconstructed video is relatively similar to the original video.
Mingkai Chen 0001, Mujian Zeng, Xiaoming He 0004, Jian Xiong 0005, Lei Wang 0009, Anwer Adel Al-Dulaimi, Shahid Mumtaz
IEEE Trans. Image Process.1
2025 Energy-Efficient Personalized Federated Learning for Establishing Green Iot
abstract
Green Internet of Things (Green IoT) is a technique that intends to reduce energy consumption and carbon emissions of Internet of Things (IoT) devices by optimizing hardware design, communication protocols, and data processing. One of the most promising schemes to realize Green IoT is personalized Federated Learning (pFL). Unfortunately, existing pFL methods still need further improvement in achieving Green IoT from the following aspects. 1) Computational energy consumption: model training on IoT devices generates a substantial amount of computational energy consumption. 2) Model performance: the dynamic role differences in each layer of the trained deep neural network need to be considered. Jointly considering these aspects, we present a novel pFL framework named Energy-Efficient personalized Federated Learning (EE-pFL) for establishing Green IoT. Specifically, an IoT device serves as an edge server. Each IoT device produces a customized model through a model training phase and a model aggregation phase. In the model training phase, a threshold-based sparsification strategy is introduced to reduce the computational energy consumption of IoT devices by selectively executing parameter updates. In the model aggregation phase, layer aggregation and an Adaptive Weight Calculation (AWC) mechanism are proposed to capture dynamic role differences in different layers of a deep neural network. Experimental results demonstrate that EEpFL shows lower computational energy consumption and higher classification accuracy than advanced benchmarks.
Yingchi Mao, Xiaoming He 0004, Mingkai Chen 0001, Saba Al-Rubaye
ICC6
2025 A GAI-Based Haptic Transmission Architecture for Extending Headset Lifespan in Haptic-Enhanced XR
abstract
Moving computing components from headsets to cloud servers is a promising approach to increasing headsets' lifespan and comfort for Extended Reality (XR) users. However, latency is an unavoidable challenge for XR services due to longdistance transmission, especially when haptic feedback (usually requires 1 ms latency) is involved. To address this challenge, we leverage the Latent Diffusion Model (LDM) to propose a novel transmission architecture, which can accurately generate the future potential haptic feedback from current or previous video frames. Moreover, we also introduce an acceleration architecture to accelerate the haptic feedback generation process. The simulations indicate that the lifespan of headsets can be tripled by moving computing resources to the cloud. In addition, our proposed architecture can accurately generate future potential haptic feedback at least 170 ms before contact, which can satisfy the 1 ms latency requirement of haptic feedback.
Zhe Zhang 0010, Mingkai Chen 0001, Anqi Tong, Chung-Horng Lung, Joel J. P. C. Rodrigues
ICC3
2025 Efficient Semantic Codec for Real-time Vibrotactile Transmission
abstract
Nowadays, haptic data has gained a fast-growing volume with enormous interaction points during human-computer interaction and embodied AI. In the near future, the massive haptic signals -encompassing both kinesthetic and vibrotactile signals- will place significant demands on both communication and computing resources. To address this challenge, we propose the first task-oriented sematic codec of low-delay vibrotactile transmission, namely, vibrotactile semantic codec (VTSC). Specifically, we design a perception-based vibrotactile semantic extraction mechanism (PSEM) that considers the high and low thresholds of vibrotactile perception in effective semantic coding while adhering to the low delay constraint. Inspired by this principle, we then propose a vibrotactile semantic encoder (VSE) with local and global semantic extractors, which can efficiently extract and preserve semantic features within the short frame context. Besides, we present a semantic distribution loss function to enhance the learning of meaningful representations. Comprehensive experiments demonstrate the superiority of our VTSC, achieving significantly higher task accuracy than the state-of-the-art vibrotactile codecs at the same compression ratio (CR), e.g. 60% improvement when CR=256. When compared to transferred audio-visual sematic codecs, our VTSC also shows promising improvements, validating the effectiveness our approach.
Runjie Wang, Kemi Chen, Shuijie Li, Mingkai Chen 0001, Tiesong Zhao
ACM Multimedia4
2025 Visual-Tactile Fusion for Multimodal Semantic Communication with Foundation Models
abstract
Integrating vision and touch is key to understanding the physical world, but it faces two main challenges: effective multimodal fusion and high-fidelity tactile representation. This paper proposes a multimodal semantic communication framework based on foundation models through visual-tactile fusion. First, a multimodal enhancement fusion network extracts deep features from video to improve tactile recognition and semantic understanding. Second, a CLIP-driven framework, grounded in a tactile knowledge base, enhances the accuracy of tactile information transmission. An end-to-end model with joint source-channel coding further improves transmission efficiency. Finally, we introduce a tactile generative reconstruction method using ImageBind, which ensures high similarity in both visual features and pressure distribution. Experimental results confirm the effectiveness of our approach in semantic tactile reconstruction. Overall, the proposed method enables efficient, low-bit-rate communication with high semantic fidelity, offering a promising solution for visual-tactile fusion in real-world applications.
Zhuorui Wang, Mingkai Chen 0001, Xiaoming He 0004, Haitao Zhao 0004, Yun Lin 0005, Mariam Hussain, Shahid Mumtaz
VTC2025-Spring2
2025 Two-stage Reinforcement Learning Empowered Wireless Semantic Video Transmission
abstract
With the rapid development of video conferencing technology, interactive multimedia services proliferate, resulting in a surge in business traffic. Remote work and online collaboration will become the mainstream office in the future, which poses new challenges to computer vision technology in the field of communication. At the same time, the appearance of semantic communication has solved the problem that the traditional video coding technology can not deal with semantic redundancy. Therefore, we propose the Semantic Transmission Optimization of Video Conference(STOVC) architecture to meet the interactivity requirements of multimedia services. We first use Iterated Integrated Attribution(IIA) and Segment Anything Model 2(SAM2) to understand the video and complete the semantic segmentation of the video. We then propose Neural Video Compression with Actor-Critic(NVCAC) to optimize video encoding and transmission strategies to balance video compression with semantic fidelity. Finally, the$\alpha$-fusion method is used to reconstruct the video at the receiving end. In addition, we use DeepCache accelerated stable diffusion to generate conference video to enrich the data set. The experimental results show that STOVC has a better performance than H.264 and H.265 at low bit rate.
Mujian Zeng, Kaipeng Zheng, Wenqin Zhuang, Mingkai Chen 0001
WCNC5
2025 Federated-Learning-Enabled Cross-Modal Semantic Communication for 6G
abstract
In view of super-large scale access and dynamic connectivity requirements in 6G, the number of users and data is increasing exponentially, which makes it difficult to achieve sustainable development of communications. Meanwhile, with the development of cross-modal processing, semantic information is highly considered accurate and intelligent, which is expected to bring new ideas for 6G. Therefore, in this study, we propose a novel framework of cross-modal semantic communications to improve the efficiency of the multimodal processing, especially the tactile modal. Meanwhile, we have redesigned the significant aspects of artificial intelligence (AI), such as encoding, transmission, and processing. As federated learning (FL) inherently supports multiple privacy-preserving and security measures, we also introduce FL to assist AI training in cross-modal semantic communications, including the expansion of multimodal data, cross-modal semantic extraction, comprehensive decision-making, and privacy protection. First, in encoding aspects, we propose a hybrid coding for haptic signal coding by deep learning (DL) according to the semantic association between material identification and tactile code optimization. Second, in transmission aspects, we propose a modal-aware resource allocation for the fairness optimization between transmission requirements and network resources with deep reinforcement learning (DRL). Third, in signal processing, low-rank signal reconstruction and immersive quality of experience (QoE) evaluation by DL and machine learning (ML) are provided to improve and quantify users’ experience. In addition, the computational experiments of those key technologies are shown individually and the specification of their probability is discussed separately. Finally, future research topics related to the above issues are suggested.
Ruochen Huang, Chen Qiu 0004, Mingkai Chen 0001, Changwei Zhang, Hongbo Zhu 0002
IEEE Internet Things J.3
2025 Edge Computing Enabled Large-Scale Traffic Flow Prediction With GPT in Intelligent Autonomous Transport System for 6G Network
abstract
The Intelligent Autonomous Transport System in 6G (6G-IATS) refers to the coordination of 6G, Artificial Intelligence (AI), and intelligent transportation systems, which is expected to revolutionize future intelligent transportation systems. In 6G-IATS, large-scale traffic flow prediction, affiliated with time series prediction, holds significant value for transportation planning and urban management. As an emerging AI method, Large Language Models (LLMs) have emerged prominently in time series forecasting. Unfortunately, it is challenging to achieve accurate and efficient large-scale traffic flow prediction by LLMs in 6G-IATS, due to the two issues: a) these LLMs fail to capture the spatio-temporal correlations in a large-scale road network, leading to limited prediction accuracy, and b) they process a substantial amount of training data on the central server, which imposes low training efficiency. Jointly considering the two concerns, this paper proposes a novel LLM and edge computing-based architecture for large-scale traffic flow prediction in 6G-IATS, called Spatio-Temporal Generative Large Language Model on Edge (STGLLM-E). In this architecture, we first decompose the entire large-scale road network into several subgraphs. To capture the spatio-temporal correlations, an LLM-based method named Spatio-Temporal Generative Large Language Model (STGLLM) including Spatio-Temporal Module (STM) and Generative Large Language Model (GLLM) is proposed. Secondly, to improve the training efficiency of the STGLLM-E, an edge training strategy based on edge servers is devised. Experiments are conducted on two real-world traffic flow datasets. The experimental results illustrate that the STGLLM-E is superior to the baselines in the prediction accuracy and the efficiency of training.
Yingchi Mao, Huajun Cui, Xiaoming He 0004, Mingkai Chen 0001
IEEE Trans. Intell. Transp. Syst.5
2025 QoE-Driven Proactive Caching With DRL in Sustainable Cloud-to-Edge Continuum
abstract
Cloud-enabled edge computing scenarios can intelligently cache and update the content on a periodic basis, thereby enhancing users' overall perception of quality, which is called quality of experience (QoE). To enhance the QoE, we aim to the multi-objective optimization, which maximizes the cache hit ratio while simultaneously minimizing traffic load and time latency. To address this issue, we focus on employing an innovative algorithm named HT-PAD, which provides a complete solution for prediction and decision-making for proactive caching. First, to improve the prediction accuracy of the cached content, we use the encoding layer in hyperdimensional computing to extract the information features. Second, HD-Transformer, as the prediction part of HT-PAD, is proposed to make predictions based on user preferences, historical information, and popular information. HD-Transformer uses DNN to predict user preferences and process time series data by combining hyperdimensional computation with Transformer. Third, to avoid error in the prediction content, we employ PER-MADDPG as the decision-making part of HT-PAD, which consists of Multi-Agent Deep Deterministic Policy Gradient (MADDPG) and Prioritized Experience Replay (PER). We use MADDPG to enhance the content decision-making and utilized PER to select appropriate training samples for PER-MADDPG. Finally, our experiments have shown that our proposed approach achieves the strong performance in terms of the edge hit ratio, the latency, and the traffic load, thus improving the QoE
Xiaoming He 0004, Huajun Cui, Yinqiu Liu, Mingkai Chen 0001, Maher Guizani, Shahid Mumtaz
IEEE Trans. Mob. Comput.5
2025 Exploring MEC Server Strategy in Blockchain Networks: Mining for Mobile Users or for Self
abstract
Blockchain-based decentralized applications (DApps) offer enhanced security and decentralization features; nevertheless, their maintenance demands substantial computational resources and poses challenges for deployment in mobile networks. To address this, a number of studies have explored offloading blockchain mining tasks from mobile users to mobile edge computing (MEC) servers. However, the existing literature overlooks the fact that MEC servers can not only mine for mobile users but also mine for themselves, potentially explaining why MEC mining offloading has not gained broad acceptance within the industry. In this work, we exploit a more practical case and rethink the question of whether MEC servers lease computing power to mobile users by taking into account that MEC servers can mine for themselves. We establish a game model and apply backward induction to analytically characterize Nash equilibria for mining strategies adopted by MEC and mobile users. Our findings suggest that, if MEC can mine for self, MEC would mine for mobile users only under specific conditions where mobile users possess superior information gathering capability (at least better than the MEC server) or the whole blockchain system exhibits significant network value. We further provide a series of simulations to verify our conclusion and illustrate the impact of network parameters on the strategies of both sides.
Xintong Ling, Weihang Cao, Mingkai Chen 0001, Jiaheng Wang 0001, Zhi Ding 0001, Xiqi Gao 0001
IEEE Trans. Mob. Comput.4
2025 Cross-Modal Haptic Compression Inspired by Embodied AI for Haptic Communications
abstract
Haptic data compression has gradually become a key issue for emerging real-time haptic communications in Tactile Internet (TI). However, it is challenging to achieve a trade-off between high perceptual quality and compression ratio in haptic data compression scheme. Inspired by the perspective of embodied AI, we propose a cross-modal haptic compression scheme for haptic communications to improve the perception quality on TI devices in this paper. Since multimodal fusion is routinely employed to improve the ability of system in cognition, we assume that haptic codec is guided by visual semantics to optimize parameter settings in the coding process. We first design a multi-dimensional tactile feature fusion network (MTFFN) relying on multi-head attention mechanism. The MTFFN extracts the multi-dimensional features from the material surface and maps them to infer the coding parameters. Secondly, we provide second-order difference and linear interpolation to establish an criterion for the determination of optimal codec parameters, which are customized by the material categories so as to give high robustness. Finally, the simulation results reveal that our compression scheme can efficiently make a personalized codec procedure for different materials, obtaining more than 17% improvement in terms of compression ratio with high perceptual quality at the same time.
Xinmeng Tan, Mingkai Chen 0001, Zhe Zhang 0010, Xin Wei 0001, Tiesong Zhao
IEEE Trans. Multim.3
2024 Federated Knowledge Distillation Enabled Image Semantic Communication
abstract
The transition from 5G to beyond 5G (B5G) heralds a shift towards more pervasive and intelligent communication trends. This evolution necessitates a departure from traditional information theory towards semantic communication (SemCom), propelled by artificial intelligence (AI), aimed at enhancing capacity and optimizing resources. Concurrently, image SemCom (ISC) emerges to empower image applications. However, ISC demands suitable devices and ample computing resources to support complex neural network models and their training, presenting a significant challenge. In response, we propose a federated semantic feature distillation (FedSFD) architecture to enhance the overall performance of ISC. Combining federated learning (FL) and feature distillation (FD), FedSFD facilitates the transfer of group feature knowledge. Specifically, the powerful server model refines itself by approximating the distance between its middle layer features and those of the devices via FD. Subsequently, lightweight device ISC models leverage FD to incorporate the server model’s knowledge during training. This iterative process is bolstered by the information bottleneck (IB)-based loss function, enhancing image compression and reconstruction capabilities. Notably, this architecture operates without necessitating a unified model, thereby offering improved privacy protection. Simulation experiments demonstrate that compared to the baseline, our approach can achieve superior integration capability for ISC at the noisy edge.
Xinmian Xu, Kaipeng Zheng, Haie Dou, Mingkai Chen 0001, Lei Wang 0009
GLOBECOM4
2024 Federated KD-Assisted Image Semantic Communication in IoT Edge Learning
abstract
The evolution from fifth-generation mobile communications (5G) to beyond 5G (B5G) will lead to more ubiquitous and smarter paradigms in the Internet of Things (IoT). Communication will shift from the classical information theory to a semantic communication (SemCom) paradigm driven by artificial intelligence (AI) to enhance the capacity and optimize resources. Image SemCom (ISC) will empower IoT applications, such as drone image acquisition. However, ISC requires suitable devices and sufficient computing resources to support complex neural network models, posing a significant challenge. To address this, we propose a federated semantic feature distillation (FedSFD) architecture to improve the global performance of ISC by combining federated learning (FL) and feature distillation (FD) for the feature knowledge transfer. First, lightweight IoT device models in the edge group and the powerful server model alternately minimize losses to update parameters. After learning middle-layer features of all the edge models, the server can guide the individual device models. Second, we incorporate the information bottleneck (IB) concept into the design of loss functions to balance compression and reconstruction. Third, focusing on the tradeoff between the local training and knowledge interaction, FedSFD achieves image semantic reconstruction without sharing the private data, ensuring personalization and privacy protection in the FL framework. Finally, compared to the baseline, the simulation experiments show that the proposed approach achieves better ISC reconstruction and noise robustness during group ISC.
Xinmian Xu, Yikai Xu, Haie Dou, Mingkai Chen 0001, Lei Wang 0009
IEEE Internet Things J.4
2024 Tactile Codec with Visual Assistance in Multi-modal Communication for Digital Health
Mingkai Chen 0001, Xinmeng Tan, Huiyan Han
Mob. Networks Appl.1
2023 Modal-Aware Resource Allocation for Cross-Modal Collaborative Communication in IIoT
abstract
With the development of human–machine interactions, users are increasingly evolving toward an immersion experience with multidimensional stimuli. Facing this trend, cross-modal collaborative communication is considered an effective technology in the Industrial Internet of Things (IIoT). In this article, we focus on open issues about resource reuse, pair interactivity, and user assurance in cross-modal collaborative communication to improve Quality of Service (QoS) and users’ satisfaction. Therefore, we propose a novel architecture of modal-aware resource allocation to solve these contradictions. First, taking all the characteristics of multimodal into account, we introduce network slices to visualize resource allocation, which is modeled as a Markov decision process (MDP). Second, we decompose the problem by the transformation of probabilistic constraint and Lyapunov Optimization. Third, we propose a deep reinforcement learning (DRL) decentralized method in the dynamic environment. Meanwhile, a federated DRL framework is provided to overcome the training limitations of local DRL models. Finally, numerical results demonstrate that our proposed method performs better than other decentralized methods and achieves superiority in cross-modal collaborative communications.
Mingkai Chen 0001, Lindong Zhao, Xin Wei 0001, Mohsen Guizani
IEEE Internet Things J.1
2023 In-Network Caching for ICN-Based IoT (ICN-IoT): A Comprehensive Survey
abstract
The Internet of Things (IoT) has already emerged as one of the most popular directions in today’s information and communication technology (ICT) domain. With its advancement over different application areas, such as smart home, smart healthcare, industry 4.0, etc., a huge amount of data has been generated by billions of IoT devices, which aggravates the shortcomings of the network layer (IP)-based networks, such as limited expressiveness of IP addressing, inefficient support for mobility, and in-network caching. Building IoT on top of information-centric networking (ICN) is believed to be a promising solution to tackle the above challenge, especially the in-network caching of ICN can significantly benefit IoT in terms of reducing data and saving IoT devices’ energy. However, caching IoT data is more challenging than caching traditional Internet content, e.g., video, because IoT data are usually valid within a certain period of time, and IoT devices are typically constrained with battery. Hence, in this survey, we first review the current implementation proposals of ICN-based IoT (ICN-IoT). Next, we present the conventional caching decision policies and replacement policies which could be adopted to mitigate the aforementioned challenges, e.g., reducing IoT traffic, saving energy, and reducing data retrieval latency. Further, since leveraging machine learning (ML) techniques have the potential to further improve the caching efficiency by dealing with uncertainties, e.g., predicting unknown information, adaptively interacting with the environment, we also demonstrate the recently proposed ML-based caching schemes for ICN-IoT. In addition, we outline the open research issues and point out the future opportunities of caching in ICN-IoT.
Zhe Zhang 0010, Chung-Horng Lung, Xin Wei 0001, Mingkai Chen 0001, Subhajit Chatterjee, Zhicai Zhang
IEEE Internet Things J.4
2023 Resource Allocation for Multi-Traffic in Cross-Modal Communications
abstract
Cross-modal communications that incorporate audio-visual and tactile signals will bring a more holistic immersive experience to people. However, due to the different transmission requirements of these signals, it is a challenging task to rationalize the allocation of transmission resources. Therefore, this work proposes a joint transmission scheme to deal with the resource allocation problem of diverse signals. Network slicing and puncturing architecture are introduced in the scheme to achieve flexible resource allocation and reduce the wasting of resources. To reduce the negative impact of puncturing transmission on users, we construct the optimization problem related to transmission rate and reliability. This problem can realize the reasonable allocation of radio resources and meet the transmission requirements of the two types of signals. Next, we divide the optimization problem into two parts: video traffic resources allocation and tactile traffic puncturing resources allocation. To solve both of the problems, we leverage the channel matching (CM) algorithm and puncturing resource allocation (PRA) algorithm. In addition, we discuss the advantages and disadvantages of both ways to occupy puncturing resources, namely, occupy resources proportionally (ORP) and occupy resources blocks for transmission (ORB). Finally, the effectiveness of the proposed scheme is verified by comparing the excepted rate and resources loss ratio of system users with different schemes.
Lei Wang 0009, Anmin Yin, Xue Jiang 0003, Mingkai Chen 0001, Kapal Dev, Nawab Muhammad Faseeh Qureshi, Jiming Yao, Baoyu Zheng
IEEE Trans. Netw. Serv. Manag.4
2020 Deep-broad Learning System for Traffic Flow Prediction toward 5G Cellular Wireless Network
abstract
Nowadays, accurate traffic flow prediction toward 5G cellular wireless network has become an indispensable part for future artificial intelligence (AI)-assisted network. Meanwhile, low delay communication is also the essential part in the upcoming 5G era. However, traditional deep learning models applied in traffic flow prediction have many drawbacks, such as too much running time and computational resources. To tackle these issues, especially jointly considering effectiveness and efficiency, we design a deep-broad learning system (DBLS) for traffic flow prediction. Specifically, based on broad learning system (BLS), we firstly adopt deep representative learning to extract meaningful information from raw data in mapped feature nodes. Then, to further improve the performance of prediction, we add some other nodes i.e., enhancement nodes generated from mapped features as extra inputs to enhance the representative capability. Finally, taking mapped features nodes and enhancement nodes as inputs of the last-layer neural network, we apply ridge regression to compute the final weights quickly. Experimental results demonstrate that our proposed DBLS can make full use of advantages of both deep neural network and traditional BLS to increase the accuracy of traffic flow prediction, meanwhile, maintaining low complexity and running time.
Mingzi Chen, Xin Wei 0001, Liqi Huang, Mingkai Chen 0001, Bin Kang
IWCMC5
2020 Federated Quantile Regression over Networks
abstract
In order to solve the issue of isolated data islands and data security and personal privacy in the development of artificial intelligence, federated machine learning effectively solves the problem of sharing knowledge while protecting user privacy and data security. In a wireless sensor network, a secure learning framework is particularly needed, so that each sensor node can jointly learn knowledge without leaking local node data. Compared to traditional regression analysis algorithms, quantile regression can more fully describe the relationship between response values and its covariates by estimating conditional quantile sequences rather than a single value (such as the mean). In this paper, we propose a quantile regression federated learning framework that applies quantile regression with federated learning frameworks to wireless sensor networks, studies the performance of the algorithms, and the effectiveness of the algorithms is verified by simulations.
Liqi Huang, Xin Wei 0001, Peikang Zhu, Mingkai Chen 0001, Bin Kang
IWCMC5
2020 Computation Offloading With Reinforcement Learning in D2D-MEC Network
abstract
With the deployment of compute-intensive applications on mobile devices, some work aims to support these applications by enhancing mobile edge computing (MEC) using device-to-device (D2D) communication technology. Different from previous work, we investigate the energy consumption optimization of MEC assisted by idle user equipment (UE) in a mobile environment when combining the two technologies, and propose an optimization method for continuous time. In this paper, we build a D2D-MEC model with user mobility. To optimize the processing decision to save the energy consumption in continuous time under this model, we define long-term costs and formulate the problem of minimizing long-term costs as a markov decision process (MDP) problem. The complex MDP problem is decomposed into two sub-problems, where the explicit cost is minimized first and then the long-term cost. Due to the uncertainty of the environment caused by the mobility of the UE and the high dimensionality of the environmental information, we use the reinforcement learning method based on the neural network approximation to minimize the long-term cost.
Gaibin Li, Mingkai Chen 0001, Xin Wei 0001, Wenqin Zhuang
IWCMC2
2020 QoE-Driven Distributed Content Segments Sharing with Service Differentiation in D2D Network
abstract
Content sharing via device to device (D2D) communication is considered as a promising method to improve the performance of cellular network. However, in D2D networks, the placement of multimedia content is still an complex and urgent issue. Hence, it is significant to figure out a fine-grained content collaboration placement and delivery strategy in D2D networks. For the efficient utilization of the storage and downloading capacity of user devices, we introduce a distributed content segment sharing strategy. In order to improve QoE with limited network resource, such strategy provides users multimedia service with differentiated quality. Then, we formulate QoE-driven D2D content segment placement with service differentiation as a submodular maximization problem, and a polynomial time algorithm is proposed. Simulation results shows that proposed algorithm outperforms the algorithms without distributed content segments sharing mechanism in terms of QoE. It also demonstrates that proposed algorithm has better on the performance of cache content diversity and fairness.
Yiming Song, Mingkai Chen 0001, Xin Wei 0001
IWCMC2
2020 MEC-enabled video streaming in device-to-device networks
abstract
By offloading video streaming from the centralised cloud to the edge, mobile edge computing (MEC) servers offer new opportunities for real‐time video transmission. Deploying on the edge of users can ensure low latency transmission, however, the limited storage and computing ability cannot adapt to the currently used video transmission technologies such as video transcoding or simulcast. To solve this problem, a more flexible video transmission architecture needs to be considered. Under this motivation, the authors propose a device‐to‐device (D2D) assisted video streaming scheme, which fuses the technical advantages of MEC and scalable video coding. Specifically, they first construct a novel architecture for delay‐sensitive live video streaming services in edge‐enabled wireless heterogeneous networks named MEC‐enabled goodput‐aware (MEGA) model. Then they present a mathematical formulation for optimising the aggregation goodput performance of video traffic including both cellular and D2D links. Finally, they derive a three‐step solution based on a distributed heuristic algorithm. Numerical simulation results show that MEGA outperforms existing models in terms of goodput, end‐to‐end delay, effective loss rate, and users' quality‐of‐experience.
Huangda Lin, Mingkai Chen 0001, Bin Kang, Lei Wang 0009
IET Commun.3
2019 Intelligent Content Sharing Based on Cooperative Crowdsensing
abstract
Mobile crowdsensing (MCS) has become a promising solution to support the location-based content sharing applications. To meet users' demand on personalized content sharing, a general region of interest (RoI) distribution model that allows each user to have its specific RoI needs to be considered. In this context, how to deal with the asymmetry of cooperation caused by different RoI distribution is of significance for achieving the full benefits of personalized content sharing. Thus motivated, we propose an intelligent content sharing scheme based on cooperative crowdsensing, which ensures both efficiency and fairness. Specifically, users' decision-making of whether to participate in MCS is cast as a MCS participation game (MPG). The game captures the impact of different RoI distributions on the collective cooperation of MCS. By computing the Nash equilibrium of MPG with desirable properties, we develop a cooperation scheme that maximizes the overall system utility and is acceptable to all users. The system efficiency of the proposed scheme is further quantified by numerical simulations over various parameters.
Lindong Zhao, Lei Wang 0009, Mingkai Chen 0001, Bin Kang, Baoyu Zheng
ICC3
2019 Delay Constrainted-Rate Allocation for SVC over Device-to-Device Networks
abstract
Device-to-Device (D2D) multicast content sharing is becoming a promising technology to alleviate video traffic overload and can improve the quality of local area services. Whereas existing studies mainly focus on the delay or throughput performance. However, for delay sensitive real-time video traffic, throughput as an indicator of the network-layer cannot properly indicate the benefits of upper-layer applications. Thus, in this paper we propose a video multicast scheme for D2D cooperative scalable video coding (SVC) streaming distribution to cope with the difference between multicast channels firstly. Then, we have provided analytical expressions of the goodput in heterogeneous multicast networks based on D2D collaboration, and a distributed heuristic algorithm is proposed to solve this NP-hard optimization problem. Our results show that the proposed scheme can effectively reduce end-to-end delay, effective loss rate and improve the goodput in the system.
Lei Wang 0009, Huangda Lin, Mingkai Chen 0001, Bin Kang, Wenqin Zhuang
IWCMC3
2019 Sidelobe interference reduced scheduling algorithm for mmWave device-to-device communication networks
Lei Wang 0009, Siran Liu, Mingkai Chen 0001, Guan Gui 0001, Hikmet Sari
Peer-to-Peer Netw. Appl.3
2018 Reflection Based Resource Allocation for Indoor mmWave D2D Communications
abstract
The abundant spectrum resources in millimeterwave (mmWave) frequency band enable high throughput for indoor communications. However, the huge path loss and high blockage probability result in vulnerability during transmission. The performance of propagation is severely influenced by blocked odds according to the reflections off smooth surfaces. In this paper, we propose a random blockage model to characterize the reflections off the walls and ceiling. To adjust the variant transmission which is suffering different reflections, we first propose a combinational temperature random algorithm (CTRA) to arrange resource blocks with plausible availability. The CTAR includes two phases: partition and matching, which highly enhance the average throughput. The main conclusions are that the modified algorithm brings about higher system throughput with appropriate temperature factors and performances much better.
Lei Wang 0009, Xiaoting Yu, Yanshan Chen, Mingkai Chen 0001
APCC4
2017 QoE-Driven D2D Media Services Distribution Scheme in Cellular Networks
abstract
Device-to-device (D2D) communication has been widely studied to improve network performance and considered as a potential technological component for the next generation communication. Considering the diverse users’ demand, Quality of Experience (QoE) is recognized as a new degree of user’s satisfaction for media service transmissions in the wireless communication. Furthermore, we aim at promoting user’s Mean of Score (MOS) value to quantify and analyze user’s QoE in the dynamic cellular networks. In this paper, we explore the heterogeneous media service distribution in D2D communications underlaying cellular networks to improve the total users’ QoE. We propose a novel media service scheme based on different QoE models that jointly solve the massive media content dissemination issue for cellular networks. Moreover, we also investigate the so-called Media Service Adaptive Update Scheme (MSAUS) framework to maximize users’ QoE satisfaction and we derive the popularity and priority function of different media service QoE expression. Then, we further design Media Service Resource Allocation (MSRA) algorithm to schedule limited cellular networks resource, which is based on the popularity function to optimize the total users’ QoE satisfaction and avoid D2D interference. In addition, numerical simulation results indicate that the proposed scheme is more effective in cellular network content delivery, which makes it suitable for various media service propagation.
Mingkai Chen 0001, Lei Wang 0009, Xin Wei 0001
Wirel. Commun. Mob. Comput.1