EDBT 2026 Demo / reviewers in the wild / expert
Wenhan Yu
dblp:330/2166
· DBLP profile ↗
18ranked-venue papers
10as first author
18since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 11 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Benchmarking multi-step legal reasoning and analyzing Chain-of-Thought effects in Large Language Models
Wenhan Yu, Xinbo Lin, Lanxin Ni, Jinhua Cheng, Lei Sha |
Inf. Process. Manag. | 1 |
| 2026 | TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety AlignmentabstractZhewen Tan, Wenhan Yu, Jianfeng Si, Tongxin Liu, Kaiqi Guan, Huiyan Jin, Jiawen Tao, Xiaokun Yuan, Xiangzheng Zhang, Duohe Ma, Tong Yang, Lin Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhewen Tan, Wenhan Yu, Jianfeng Si, Tongxin Liu, Kaiqi Guan, Huiyan Jin, Jiawen Tao, Xiaokun Yuan, Xiangzheng Zhang, Duohe Ma, Tong Yang 0003, Lin Sun 0010 |
ACL (1) | 2 |
| 2026 | Multiagent Deep Reinforcement Learning for Device-Enhanced Distributed Task Scheduling in Terminal-Edge Collaborative Computing NetworksabstractDevice-enhanced mobile edge computing (MEC) is an emerging technology designed to handle intensive and delay-sensitive tasks through device-to-device (D2D) communication. In this paper, we present a terminal-edge collaborative computing network to investigate device-enhanced distributed tasks scheduling (DDTS) with specific application deployment. Our optimization focuses on offloading choices, bandwidths, and computing frequencies, aiming to minimize execution costs, including processing delay and energy consumption. We decouple the joint multiple goals optimization problem into several sub-problems which are solved by math optimization methods except the NP-hard offloading choices sub-problem. This NP-hard problem is modeled as a multitask scheduling game (MTSG), which we demonstrate to be a potential game with at least one Nash equilibrium solution. However, considering further the dynamic nature of real-world application deployment and the complexity of large-scale games, the problem evolves into a stochastic game with a Markov policy (SGMP). Thus, we propose a multi-agent DDTS algorithm based on a dueling double deep Q-network (D3QN) to approximate an optimal solution. Extensive experiments confirm the feasibility and efficiency of our approach. Yukun Sun, Wenhan Yu, Jun Zhao 0007, Xing Zhang 0001 |
IEEE Internet Things J. | 2 |
| 2026 | PrivTuner With Homomorphic Encryption and LoRA: A P3EFT Scheme for Privacy-Preserving Parameter-Efficient Fine-Tuning of AI Foundation ModelsabstractAI foundation models have recently demonstrated impressive capabilities across a wide range of tasks. Fine-tuning (FT) is a method of customizing a pre-trained AI foundation model by further training it on a smaller, targeted dataset. In this paper, we initiate the study of the Privacy-Preserving Parameter-Efficient FT (P3EFT) framework, which can be viewed as the intersection of Parameter-Efficient FT (PEFT) and Privacy-Preserving FT (PPFT). PEFT modifies only a small subset of the model’s parameters to achieve FT (i.e., adapting a pre-trained model to a specific dataset), while PPFT uses privacy-preserving technologies to protect the confidentiality of the model during the FT process. There have been many studies on PEFT or PPFT, but very few on their fusion, which motivates our work on P3EFT to achieve both parameter efficiency and model privacy. To exemplify our P3EFT, we present thePrivTunerscheme, which incorporates Fully Homomorphic Encryption (FHE) enabled privacy protection into LoRA (short for “Low-Rank Adapter”), a popular PEFT solution published in ICLR 2021 [1]. Intuitively speaking, PrivTuner allows the model owner and the external data owners to collaboratively implement PEFT with encrypted data. After describing PrivTuner in detail, we further investigate its energy consumption and privacy protection. Then, we consider a PrivTuner system over wireless communications and formulate a joint optimization problem to adaptively minimize energy while maximizing privacy protection, with the optimization variables including FDMA bandwidth allocation, wireless transmission power, computational resource allocation, and privacy protection. A resource allocation algorithm is devised to solve the problem. Experiments demonstrate that our algorithm can significantly reduce energy consumption while adapting to different privacy requirements. Yang Li 0187, Wenhan Yu, Jun Zhao 0007 |
IEEE Trans. Wirel. Commun. | 2 |
| 2026 | User-Centric Heterogeneous-Action Deep Reinforcement Learning for Virtual Reality in the Metaverse Over Wireless NetworksabstractThe Metaverse emerging as maturing technologies are empowering the different facets. Virtual Reality (VR) technologies serve as the backbone of the virtual universe within the Metaverse to offer a highly immersive user experience. As mobility is emphasized in the Metaverse context, VR devices reduce their weights at the sacrifice of local computation abilities. In this paper, for a system consisting of a Metaverse server and multiple VR users, we consider two cases of (i) the server generating frames and transmitting them to users, and (ii) users generating frames locally and thus consuming device energy. As Metaverse emphasizes on the accessibility for all users anywhere and anytime, the users can have totally different characteristics, devices and demands. In this paper, the channel access arrangement (including the decisions on frame generation location), and transmission powers for the downlink communications from the server to the users are jointly optimized by our proposed user-centric Deep Reinforcement Learning (DRL) algorithm, namely User-centric Critic with Heterogenous Actors (UCHA). Comprehensive experiments demonstrate that our UCHA algorithm leads to remarkable results under various requirements and constraints. Wenhan Yu, Terence Jie Chua, Jun Zhao 0007 |
IEEE Trans. Wirel. Commun. | 1 |
| 2026 | Optimization for 6G Wireless Communications With Heterogeneous VR and Non-VR 360° Videos: A Differentiated Reinforcement Learning ApproachabstractVirtual Reality (VR) and its reliance on 360° videos are pivotal in delivering a seamless, immersive experience. In the era of emerging 6G technology, where diverse mobile devices are increasingly prevalent, optimizing Quality of Experience (QoE) becomes critical. This is especially true in applications that integrate both VR and non-VR modes. The focus on 6G highlights its capacity to cater to these varying requirements, ensuring high-quality video transmission across different platforms and user experiences. This paper introduces two novel algorithms: Separated Input Differentiated Output (SIDO) and Merged Input Differentiated Output (MIDO). These algorithms are designed to optimize resolution and power allocations in downlink wireless communication, catering to both non-VR and VR users within a chunk-based structure. By encapsulating diverse parameters like subjective perceptual video quality, chunk success rate, and cybersickness into our comprehensive QoE model, we present an approach to address challenges inherent to 360° video optimization. Our deep reinforcement learning algorithms, SIDO and MIDO, further refine the optimization process. Extensive experiments reveal the efficacy of our methodologies. This work, at its core, aims to bridge the gap between technical optimization and user experience, ensuring seamless integration of users. Wenhan Yu, Jun Zhao 0007 |
IEEE Trans. Wirel. Commun. | 1 |
| 2025 | Play to Earn in Augmented Reality With Mobile Edge Computing Over Wireless Networks: A Deep Reinforcement Learning ApproachabstractPlay-to-earn (P2E) games have been gaining popularity as they enable players to earn in-game tokens which can be translated to real-world profits. With the advancements in augmented reality (AR) technologies, AR play-to-earn games become compute-intensive. In-game graphical scenes need to be offloaded from mobile devices to an edge server for computation. In this work, we consider an optimization problem where the Mobile edge computing Service Provider (MSP)’s objective is to reduce downlink transmission latency of in-game graphics, the latency of uplink data transmission, and the worst-case (greatest) battery charge expenditure of user equipments (UEs), while maximizing the worst-case (lowest) UE resolution-influenced in-game earning potential through optimizing the downlink UE-Mobile edge computing Base Station (UE-MBS) assignment, downlink, and the uplink transmission power selection. The downlink and uplink transmissions are executed asynchronously. We propose a Multi-Asynchronous-Agent, Loss-Sharing (MALS) reinforcement learning model to tackle the asynchronous and asymmetric problem. We then compare the MALS model with other baseline models and show its superiority over other methods. Finally, we conduct multi-variable optimization weighting analyses and show the viability of using our proposed MALS algorithm to tackle joint optimization problems. Terence Jie Chua, Wenhan Yu, Jun Zhao 0007 |
IEEE Trans. Wirel. Commun. | 2 |
| 2025 | Mobile Edge Adversarial Detection for Digital Twinning to the Metaverse: A Deep Reinforcement Learning ApproachabstractDigital Twinning of physical world scenes onto the Metaverse is necessary for augmented reality (AR)-assisted driving. In AR-assisted driving, physical environment scenes are first captured by AR vehicles and are uploaded to the Metaverse for the construction of the Metaverse Map. However, the development of AR-assisted driving applications invites adversaries. These attackers may place adversarial patches on physical objects, seeking to contort the Metaverse Map. As real-time, accurate detection of adversarial patches is compute-intensive, these physical world scenes have to be offloaded to the Metaverse Map Base Station (MMBS) for computation. Therefore, we considered a scenario where AR vehicles capture physical world scenes and upload these scenes in real-time to the MMBSs. We formulated an optimization problem where the MMSP’s objective is to maximize adversarial patch detection mean Average Precision (mAP), while minimizing the computed AR scene uplink transmission latency and minimizing the worst-case (largest) AR vehicle’s uplink transmission battery charge consumption, through optimizing the AR vehicle-MMBS allocation, AR vehicle uplink scene resolution selection, and AR vehicle uplink power output selection. We proposed a Heterogeneous Action (HA) algorithm to tackle the proposed problem. Extensive experiments show our HA models outperforms baseline models when compared against key metrics. Terence Jie Chua, Wenhan Yu, Jun Zhao 0007 |
IEEE Trans. Wirel. Commun. | 2 |
| 2024 | Counterfactual Reward Estimation for Credit Assignment in Multi-agent Deep Reinforcement Learning over Wireless Video Transmission
Wenhan Yu, Liangxin Qian, Terence Jie Chua, Jun Zhao 0007 |
ICDCS | 1 |
| 2024 | Orchestration of Emulator Assisted 6G Mobile Edge Tuning for AI Foundation Models: A Multi-Agent Deep Reinforcement Learning ApproachabstractThe efficient deployment and fine-tuning of foundation models are pivotal in contemporary artificial intelligence. In this study, we present a groundbreaking paradigm inte-grating 6G Mobile Edge Computing (MEC) with foundation models, specifically designed to enhance local task performance on user equipment (UE). Central to our approach is the innovative Emulator-Adapter architecture, segmenting the foundation model into two cohesive modules. This design not only conserves computational resources but also ensures adaptability and fine-tuning efficiency for downstream tasks. Additionally, we introduce an advanced resource allocation mechanism that is fine-tuned to the needs of the Emulator-Adapter structure in decentralized settings. To address the challenges presented by this system, we employ a hybrid multi-agent Deep Reinforcement Learning strategy, adept at handling mixed discrete-continuous action spaces, ensuring dynamic and optimal resource allocations. Our comprehensive simulations and validations underscore the practical viability of our approach, demonstrating its robustness, efficiency, and scalability. Collectively, this work offers a fresh perspective on deploying foundation models and balancing computational efficiency with task proficiency. Wenhan Yu, Terence Jie Chua, Jun Zhao 0007 |
VTC Spring | 1 |
| 2024 | Human-Centric Resource Allocation in the Metaverse Over Wireless CommunicationsabstractThe Metaverse will provide numerous immersive applications for human users, by consolidating technologies like extended reality (XR), video streaming, and cellular networks. Optimizing wireless communications to enable the human-centric Metaverse is important to satisfy the demands of mobile users. In this paper, we formulate the optimization of the system utility-cost ratio (UCR) for the Metaverse over wireless networks. Our human-centric utility measure for virtual reality (VR) applications of the Metaverse represents users’ perceptual assessment of the VR video quality as a function of the data rate and the video resolution and is learned from real datasets. The variables jointly optimized in our problem include the allocation of both communication and computation resources as well as VR video resolutions. The system cost in our problem comprises the energy consumption and delay and is non-convex with respect to the optimization variables. To solve the non-convex optimization, we develop a novel fractional programming technique, which contributes to optimization theory and has broad applicability beyond our paper. Our proposed algorithm for the system UCR optimization is computationally efficient and finds a stationary point to the constrained optimization. Through extensive simulations, our algorithm is demonstrated to outperform other approaches. Jun Zhao 0007, Liangxin Qian, Wenhan Yu |
IEEE J. Sel. Areas Commun. | 3 |
| 2023 | Heterogeneous 360 Degree Videos in Metaverse: Differentiated Reinforcement Learning ApproachesabstractAdvanced video technologies are driving the development of the futuristic Metaverse, which aims to connect users from anywhere and anytime. As such, the use cases for users will be much more diverse, leading to a mix of 360-degree videos with two types: non-VR and VR 360° videos. This paper presents a novel Quality of Service model for heterogeneous 360° videos with different requirements for frame rates and cybersickness. We propose a frame-slotted structure and conduct frame-wise optimization using self-designed differentiated deep reinforcement learning algorithms. Specifically, we design two structures, Separate Input Differentiated Output (SIDO) and Merged Input Differentiated Output (MIDO), for this heterogeneous scenario. We also conduct comprehensive experiments to demonstrate their effectiveness. Wenhan Yu, Jun Zhao 0007 |
GLOBECOM | 1 |
| 2023 | Mobile Edge Adversarial Detection for Digital Twinning to the Metaverse with Deep Reinforcement LearningabstractReal-time Digital Twinning of physical world scenes onto the Metaverse is necessary for a myriad of applications such as augmented-reality (AR) assisted driving. In AR assisted driving, physical environment scenes are first captured by Internet of Vehicles (IoVs) and are uploaded to the Metaverse. A central Metaverse Map Service Provider (MMSP) will aggregate information from all IoVs to develop a central Metaverse Map. Information from the Metaverse Map can then be downloaded into individual IoVs on demand and be delivered as AR scenes to the driver. However, the growing interest in developing AR assisted driving applications which relies on digital twinning invites adversaries. These adversaries may place physical adversarial patches on physical world objects such as cars, signboards, or on roads, seeking to contort the virtual world digital twin. Hence, there is a need to detect these physical world adversarial patches. Nevertheless, as real-time, accurate detection of adversarial patches is compute-intensive, these physical world scenes have to be offloaded to the Metaverse Map Base Stations (MMBS) for computation. Hence in our work, we considered an environment with moving Internet of Vehicles (IoV), uploading real-time physical world scenes to the MMBSs. We formulated a realistic joint variable optimization problem where the MMSPs' objective is to maximize adversarial patch detection mean average precision (mAP), while minimizing the computed AR scene up-link transmission latency and IoVs' up-link transmission idle count, through optimizing the IoV-MMBS allocation and IoV up-link scene resolution selection. We proposed a Heterogeneous Action Proximal Policy Optimization (HAPPO) (discrete-continuous) algorithm to tackle the proposed problem. Extensive experiments shows HAPPO outperforms baseline models when compared against key metrics. Terence Jie Chua, Wenhan Yu, Jun Zhao 0007 |
ICC | 2 |
| 2023 | Virtual Reality in Metaverse Over Wireless Networks with User-Centered Deep Reinforcement LearningabstractThe Metaverse and its promises are fast becoming reality as maturing technologies are empowering the different facets. One of the highlights of the Metaverse is that it offers the possibility for highly immersive and interactive socialization. Virtual reality (VR) technologies are the backbone for the virtual universe within the Metaverse as they enable a hyper-realistic and immersive experience, and especially so in the context of socialization. As the virtual world 3D scenes to be rendered are of high resolution and frame rate, these scenes will be offloaded to an edge server for computation. Besides, the metaverse is user-center by design, and human users are always the core. In this work, we introduce a multi-user VR computation offloading over wireless communication scenario. In addition, we devised a novel user-centered deep reinforcement learning approach to find a near-optimal solution. Extensive experiments demonstrate that our approach can lead to remarkable results under various requirements and constraints. Wenhan Yu, Terence Jie Chua, Jun Zhao 0007 |
ICC | 1 |
| 2023 | Semantic Communications, Semantic Edge Computing, and Semantic Caching with Applications to the Metaverse and 6G Mobile NetworksabstractThe increasing popularity of applications like the Metaverse has led to the exploration of new, more effective ways of communication. Semantic communication, which focuses on the meaning behind transmitted information, represents a departure from traditional communication paradigms. As mobile devices become increasingly prevalent, it is important to explore the potential of edge computing to aid the semantic encoding/decoding process, which requires significant computing power and storage capabilities. However, establishing knowledge bases (KBs) for domain-oriented communication can be time-consuming. To address this challenge, this paper proposes a semantic caching model in edge computing system that caches domain-specialized general models and user-specific individual models. This approach has the potential to reduce the time and resources required to establish individual KBs while accurately capturing the semantics behind users' messages, ultimately leading to more efficient and accessible semantic communication. Wenhan Yu, Jun Zhao 0007 |
ICDCS | 1 |
| 2023 | Mobile Edge Computing and AI Enabled Web3 Metaverse over 6G Wireless Communications: A Deep Reinforcement Learning ApproachabstractThe Metaverse is gaining attention among academics as maturing technologies empower the promises and envisagements of a multi-purpose, integrated virtual environment. An interactive and immersive socialization experience between people is one of the promises of the Metaverse. In spite of the rapid advancements in current technologies, the computation required for a smooth, seamless and immersive socialization experience in the Metaverse is overbearing, and the accumulated user experience is essential to be considered. The computation burden calls for computation offloading, where the integration of virtual and physical world scenes is offloaded to an edge server. This paper introduces a novel Quality-of-Service (QoS) model for the accumulated experience in multi-user socialization on a multichannel wireless network. This QoS model utilizes deep reinforcement learning approaches to find the near-optimal channel resource allocation. Comprehensive experiments demonstrate that the adoption of the QoS model enhances the overall socialization experience. Wenhan Yu, Terence Jie Chua, Jun Zhao 0007 |
VTC2023-Spring | 1 |
| 2023 | Asynchronous Hybrid Reinforcement Learning for Latency and Reliability Optimization in the Metaverse Over Wireless CommunicationsabstractTechnology advancements in wireless communications and high-performance Extended Reality (XR) have empowered the developments of the Metaverse. The demand for the Metaverse applications and hence, real-time digital twinning of real-world scenes is increasing. Nevertheless, the replication of 2D physical world images into 3D virtual objects is computationally intensive and requires computation offloading. The disparity in transmitted object dimension (2D as opposed to 3D) leads to asymmetric data sizes in uplink (UL) and downlink (DL). To ensure the reliability and low latency of the system, we consider an asynchronous joint UL-DL scenario where in the UL stage, the smaller data size of the physical world images captured by multiple extended reality users (XUs) will be uploaded to the Metaverse Console (MC) to be construed and rendered. In the DL stage, the larger-size 3D virtual objects need to be transmitted back to the XUs. We design a novel multi-agent reinforcement learning algorithm structure, namely Asynchronous Actors Hybrid Critic (AAHC), to optimize the decisions pertaining to computation offloading and channel assignment in the UL stage and optimize the DL transmission power in the DL stage. Extensive experiments demonstrate that compared to proposed baselines, AAHC obtains better solutions with satisfactory training time. Wenhan Yu, Terence Jie Chua, Jun Zhao 0007 |
IEEE J. Sel. Areas Commun. | 1 |
| 2022 | Detection of Uncertainty in Exceedance of Threshold (DUET): An Adversarial Patch LocalizerabstractDevelopment of defenses against physical world attacks such as adversarial patches is gaining traction within the research community. We contribute to the field of adversarial patch detection by introducing an uncertainty-based adversarial patch localizer which localizes adversarial patch on an image, permitting post-processing patch-avoidance or patch-reconstruction. We quantify our prediction uncertainties with the development of Detection of Uncertainties in the Exceedance of Threshold (DUET) algorithm. This algorithm provides a framework to ascertain confidence in the adversarial patch localization, which is essential for safety-sensitive applications such as self-driving cars and medical imaging. We conducted experiments on localizing adversarial patches and found our proposed DUET model outperforms baseline models. We then conduct further analyses on our choice of model priors and the adoption of Bayesian Neural Networks in different layers within our model architecture. We found that isometric gaussian priors in Bayesian Neural Networks are suitable for patch localization tasks and the presence of Bayesian layers in the earlier neural network blocks facilitates top-end localization performance, while Bayesian layers added in the later neural network blocks contribute to better model generalization. We then propose two different well-performing models to tackle different use cases. Terence Jie Chua, Wenhan Yu, Chang Liu 0093, Jun Zhao 0007 |
BDCAT | 2 |