Ying He 0006

dblp:39/2405-6 · DBLP profile ↗
← Back
66ranked-venue papers
14as first author
57since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 31 · 10 first-author · 25 since 2021Artificial intelligence and machine learning · 20 · 1 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 14 since 2021Systems, architecture and hardware · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ForeDiffusion: Foresight-Conditioned Diffusion Policy via Future View Construction for Robot Manipulation
abstract
Diffusion strategies have advanced visual motor control by progressively denoising high-dimensional action sequences, providing a promising method for robot manipulation. However, as task complexity increases, the success rate of existing baseline models decreases considerably. Analysis indicates that current diffusion strategies are confronted with two limitations. First, these strategies only rely on short-term observations as conditions. Second, the training objective remains limited to a single denoising loss, which leads to error accumulation and causes grasping deviations. To address these limitations, this paper proposes Foresight-Conditioned Diffusion (ForeDiffusion), by injecting the predicted future view representation into the diffusion process. As a result, the policy is guided to be forward-looking, enabling it to correct trajectory deviations. Following this design, ForeDiffusion employs a dual loss mechanism, combining the traditional denoising loss and the consistency loss of future observations, to achieve the unified optimization. Extensive evaluation on the Adroit suite and the MetaWorld benchmark demonstrates that ForeDiffusion achieves an average success rate of 80% for the overall task, significantly outperforming the existing mainstream diffusion methods by approximately 20% in high difficulty tasks, while maintaining more stable performance across the entire tasks.
Weize Xie, Ying He 0006, Leilei Wang, Binwen Bai, Zheyi Zhao, Chenyang Wang 0001, F. Richard Yu
AAAI3
2026 PSPO: Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy Optimization
abstract
Reinforcement Fine-tuning (RFT) methods such as Group Relative Policy Optimization (GRPO) have demonstrated strong capabilities in aligning Large Language Models with human preferences. However, these approaches often suffer from limited data efficiency, necessitating extensive on-policy rollouts to maintain competitive performance. We propose PSPO (Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy Optimization), a lightweight yet effective enhancement to GRPO that improves training stability and sample efficiency through two complementary techniques. First, we introduce an experience-weighted reward smoothing mechanism, which uses exponential moving averages to track group-level reward statistics for each prompt. This enables more stable advantage estimation across training steps without storing entire trajectories, allowing the model to capture historical reward trends in a lightweight and memory-efficient manner. Second, we adopt a prompt-level prioritized sampling strategy, which is an online data selection method inspired by prioritized experience replay. It dynamically emphasizes higher-impact prompts based on their relative advantages, thereby improving data efficiency. Experiments on multiple mathematical reasoning benchmarks and models show that PSPO achieves comparable or better accuracy than GRPO, while significantly accelerating convergence, and maintaining low computational and memory overhead.
Ying He 0006, Haowen Hou, Ruichong Zhang, Nianbo Zeng, Yulin Peng, Jiongfeng Fang, F. Richard Yu
AAAI2
2026 Heterogeneous Multi-Agent Reinforcement Learning for Energy-Aware Resource Scheduling in Cloud Environments
Ying He 0006, Peijie Xian, F. Richard Yu, Guangzheng Zhang, Jianbo Du
ICC1
2026 Restoring neural radiance fields performance under adverse weather conditions
Ying He 0006, Gan Chen, F. Richard Yu, Ming Li 0073, Fei Ma 0006, Guang Zhou
Eng. Appl. Artif. Intell.1
2026 Joint Service Caching and Computation Offloading in Mobile Edge Networks: A Hierarchical DRL Approach With Active Inference
abstract
Mobile edge computing (MEC) is a promising paradigm that provides abundant computation and storage resources at the edge close to mobile devices (MDs). In MEC networks, MDs offload compute-heavy tasks to nearby edge servers (ESs) for delay-sensitive processing, where relevant services are stored to support task execution. However, the limited computation and storage capacities of ESs make joint optimization of service caching and computation offloading challenging due to coupled decisions, a large solution space, and dynamic environments. In this paper, we investigate the joint optimization of service caching and computation offloading in MEC networks, aiming to maximize the cache hit ratio and minimize the average service latency. To tackle this problem, the original formulation is decomposed into two hierarchical subproblems, namely high-level service caching and low-level computation offloading. We propose a novel hierarchical deep reinforcement learning (DRL) algorithm with active inference, termed HADRL. At the high-level, we adopt a deep deterministic policy gradient (DDPG) based DRL approach to maximize the cache hit ratio. At the low-level, we employ an active inference based DRL approach to minimize the average service latency. Unlike conventional DRL, the active inference based DRL approach selects policies by minimizing expected free energy instead of relying only on explicit rewards, making it well suited for highly dynamic low-level computation offloading. According to the simulation outcomes, the HADRL scheme surpasses the benchmark algorithms with respect to cache hit ratio as well as average service latency.
Zhenjie Lv, Yuhang Wang 0019, Ying He 0006, Weiwei Fang, F. Richard Yu
IEEE Internet Things J.4
2025 ABM++: Learning Generalizable Manipulation Policies with a Mask-Guided World Model
abstract
Achieving robust generalization across diverse scenarios is crucial for advancing the practical application of robotics. Existing approaches typically rely on interaction object masks as visual inputs to predict subsequent actions, gaining a certain degree of generalization capability. However, these methods primarily map visual inputs and task-relevant object masks to expert actions, overlooking the environmental dynamics that govern physical interactions among objects during manipulation. To overcome these limitations, we introduce ABM++, a novel framework that leverages pre-trained VLMs to build a mask-guided world model (MGWM) within an imitation learning paradigm for generalized robotic manipulation. Specifically, we extend a world model into a coarse-to-fine imitation learning framework to reconstruct future mask-dominated visual features, which enables the model to capture state transitions between the current and next states based on predicted actions, effectively modeling environmental dynamics. Comprehensive experiments demonstrate that ABM++ significantly surpasses established baselines in both simulation and real-world environments, achieving a relative improvement of 12.2% across 8 complex tasks, which underscores the superiority of our method.
Fan Zhuo, Ying He 0006, F. Richard Yu, Pengshuai Yin, Fei Ma 0006
ECAI2
2025 PAFT: Prompt-Agnostic Fine-Tuning
abstract
Fine-tuning large language models (LLMs) often causes overfitting to specific prompt wording, where minor phrasing variations drastically reduce performance.To address this, we propose Prompt-Agnostic Fine-Tuning (PAFT), a method that enhances robustness through dynamic prompt variation during training.PAFT first generates diverse synthetic prompts, then continuously samples from this set to construct training instances, forcing models to learn fundamental task principles rather than surface-level patterns.Across systematic evaluations using both supervised fine-tuning (SFT) and reinforcement learning fine-tuning (RLFT), PAFT demonstrates substantially improved prompt robustness, achieving 7% higher generalization accuracy on unseen prompts than standard methods.In addition to enhanced robustness, PAFT consistently yields superior overall performance on established benchmarks for question answering, mathematical reasoning, and tool use.Notably, models trained with PAFT attain 3.2× faster inference speeds due to reduced prompt sensitivity.Ablation studies further validate effectiveness of PAFT, while theoretical analysis reveals that PAFT can effectively enhance the cross-domain generalization ability of LLM.
Chenxing Wei, Mingwen Ou, Ying He 0006, Yao Shu, F. Richard Yu
EMNLP3
2025 DEP-SLAM: A Dynamic Environment Perception SLAM System with Large Language Models
abstract
Inderscience is a global company, a dynamic leading independent journal publisher disseminates the latest research across the broad fields of science, engineering and technology; management, public and business administration; environment, ecological economics and sustainable development; computing, ICT and internet/web services, and related areas.
Ying He 0006, F. Richard Yu, Fei Ma 0006, Ming Li 0073, Guang Zhou
ICASSP1
2025 Resource Allocation for Semantic Segmentation Tasks in Autonomous Driving: A Likelihood Active Inference Approach
abstract
The latest Segment Anything Model enables realtime scene annotation and understanding for autonomous driving systems, enhancing driving safety. However, effectively allocating resources for real-time performance and accuracy remains challenging in edge-cloud architectures. Traditional reinforcement learning struggles with poor generalization and the explorationexploitation dilemma, making it difficult to define clear reward functions. To address this, we propose a likelihood active inference approach to optimize resource allocation and improve system resource utilization. We use "intelligence" as a high-level indicator to quantify the efficiency of cognition in active inference, evaluating the difference between predicted and actual states during policy exploration. Experimental results show our algorithm outperforms mainstream deep reinforcement learning algorithms, improving sample efficiency and suitability for dynamically changing task workloads.
F. Richard Yu, Ying He 0006
ICASSP3
2025 Congestion Control for Blockchain-enabled SDN in Web 4.0: A Reinforcement Learning Approach through Active Inference
abstract
Web 4.0 is characterized by decentralized intelligence and blockchain integration, which introduces significant challenges in congestion management for software-defined networking (SDN). Traditional reinforcement learning (RL)-based approaches encounter inefficiencies due to limited adaptability to decentralized and delayed online learning capabilities. To address these issues, we propose an Active Inference-based Reinforcement Learning (AIRL) framework that integrates generative modelling with RL for enhanced decision-making in congestion control. By leveraging blockchain-enabled secure model trading and predictive intelligence, AIRL ensures adaptive policy optimization while maintaining transparency and trust in decentralized network environments. The proposed method demonstrates substantial improvements in delay reduction, packet loss, and efficient utilization of network resources under various dynamic scenarios.
Chenyang Wang 0001, Xiaoxu Ren, Ying He 0006, F. Richard Yu, Victor C. M. Leung
ICDCS3
2025 Inter3D: A Benchmark and Strong Baseline for Human-Interactive 3D Object Reconstruction
abstract
Recent advancements in implicit 3D reconstruction methods, e.g., neural rendering fields and Gaussian splatting, have primarily focused on novel view synthesis of static or dynamic objects with continuous motion states. However, these approaches struggle to efficiently model a human-interactive object with n movable parts, requiring 2^n separate models to represent all discrete states. To overcome this limitation, we propose Inter3D, a new benchmark and approach for novel state synthesis of human-interactive objects. We introduce a self-collected dataset featuring commonly encountered interactive objects and a new evaluation pipeline, where only individual part states are observed during training, while part combination states remain unseen. We also propose a strong baseline approach that leverages Space Discrepancy Tensors to efficiently modelling all states of an object. To alleviate the impractical constraints on camera trajectories across training states, we propose a Mutual State Regularization mechanism to enhance the spatial density consistency of movable parts. In addition, we explore two occupancy grid sampling strategies to facilitate training efficiency. We conduct extensive experiments on the proposed benchmark, showcasing the challenges of the task and the superiority of our approach. The code and data are publicly available at https://github.com/Inter3D-ui/Inter3D.
Gan Chen, Ying He 0006, Mulin Yu, F. Richard Yu, Fei Ma 0006, Ming Li 0073, Guang Zhou
IJCAI2
2025 TagGuideBot: Enhancing Robot Intelligence with Object Tags and VLMs
abstract
This research aims to enhance the interaction between humans and robots, especially in environments with multiple similar objects or semantic ambiguities. Traditional command-based interactions typically require users to provide precise descriptions, which often poses a significant challenge. To address this issue, we propose a framework named Tag-GuideBot, which leverages Visual Language Models (VLMs) and utilizes object markers to help locate and identify objects in the environment. By integrating positional point prompts of the target objects with robot motion planning models, we aim to achieve a more accurate understanding and execution of complex commands, thus improving the efficiency and naturalness of interactions. Experimental results demonstrate that TagGuideBot effectively addresses the challenges posed by complex commands and environmental complexities, achieving an accuracy of 66.3% on user instructions extended beyond the training set, providing solid support for further optimization of human-robot interaction.
Ying He 0006, F. Richard Yu
IROS2
2025 JAM: Keypoint-Guided Joint Prediction after Classification-Aware Marginal Proposal for Multi-Agent Interaction
abstract
Predicting the future motion of road participants is a critical task in autonomous driving. In this work, we address the challenge of low-quality generation of low-probability modes in multi-agent joint prediction. To tackle this issue, we propose a two-stage multi-agent interactive prediction framework named keypoint-guided joint prediction after classification-aware marginal proposal (JAM). The first stage is modeled as a marginal prediction process, which classifies queries by trajectory type to encourage the model to learn all categories of trajectories, providing comprehensive mode information for the joint prediction module. The second stage is modeled as a joint prediction process, which takes the scene context and the marginal proposals from the first stage as inputs to learn the final joint distribution. We explicitly introduce key waypoints to guide the joint prediction module in better capturing and leveraging the critical information from the initial predicted trajectories. We conduct extensive experiments on the real-world Waymo Open Motion Dataset interactive prediction benchmark. The results show that our approach achieves competitive performance. In particular, in the framework comparison experiments, the proposed JAM outperforms other prediction frameworks and achieves state-of-the-art performance in interactive trajectory prediction. The code is available at https://github.com/LinFunster/JAM to facilitate future research.
Fangze Lin, Ying He 0006, F. Richard Yu, Hong Zhang 0013
IROS2
2025 Audio-Driven Talking Face Video Generation with Joint Uncertainty Learning
abstract
Talking face video generation with arbitrary speech audio is a significant challenge within the realm of digital human technology. The previous studies have emphasized the significance of audio-lip synchronization and visual quality. Currently, limited attention has been given to the learning of visual uncertainty, which creates several issues in existing systems, including inconsistent visual quality and unreliable performance across different input conditions. To address the problem, we propose a Joint Uncertainty Learning Network (JULNet) for high-quality talking face video generation, which incorporates a representation of uncertainty that is directly related to visual error. Specifically, we first design an uncertainty module to individually predict the error map and uncertainty map after obtaining the generated image. The error map represents the difference between the generated image and the ground truth image, while the uncertainty map is used to predict the probability of incorrect estimates. Furthermore, to match the uncertainty distribution with the error distribution through a KL divergence term, we introduce a histogram technique to approximate the distributions. By jointly optimizing error and uncertainty, the performance and robustness of our model can be enhanced. Extensive experiments demonstrate that our method achieves superior high-fidelity and audio-lip synchronization in talking face video generation compared to previous methods.
Fei Ma 0006, Yi Bin, Ying He 0006, F. Richard Yu
ICMR4
2025 ReDit: Reward Dithering for Improved LLM Policy Optimization
abstract
DeepSeek-R1 has successfully enhanced Large Language Model (LLM) reasoning capabilities through its rule-based reward system. While it's a ''perfect'' reward system that effectively mitigates reward hacking, such reward functions are often discrete. Our experimental observations suggest that discrete rewards can lead to gradient anomaly, unstable optimization, and slow convergence. To address this issue, we propose ReDit (Reward Dithering), a method that dithers the discrete reward signal by adding simple random noise. With this perturbed reward, exploratory gradients are continuously provided throughout the learning process, enabling smoother gradient updates and accelerating convergence. The injected noise also introduces stochasticity into flat reward regions, encouraging the model to explore novel policies and escape local optima. Experiments across diverse tasks demonstrate the effectiveness and efficiency of ReDit. On average, ReDit achieves performance comparable to vanilla GRPO with only approximately 10% the training steps, and furthermore, still exhibits a 4% performance improvement over vanilla GRPO when trained for a similar duration. Visualizations confirm significant mitigation of gradient issues with ReDit. Moreover, theoretical analyses are provided to further validate these advantages.
Chenxing Wei, Jiarui Yu, Ying He 0006, Hande Dong, Yao Shu, F. Richard Yu
NeurIPS3
2025 A Survey of DDoS Attack and Defense Technologies in Multiaccess Edge Computing
abstract
Multiaccess edge computing (MEC) is a novel service paradigm located at the network’s edge, where servers with computation and storage capabilities are placed in proximity to network endpoints to meet the high-speed computation and low-latency requirements of these endpoints. MEC faces numerous security challenges, with Distributed Denial of Service (DDoS) attacks being one of the primary threats. On one hand, the edge server layer (ESL) encounters a greater number of attack sources and has relatively fewer defense resources, making traditional defense techniques less applicable. On the other hand, the ESL can integrate with new technologies for earlier detection and interception of attack traffic, achieving more real-time attack mitigation. To provide comprehensive insights into the latest research developments and inspire new DDoS defense solutions, this article conducts an extensive survey and synthesis. This article begins with a summary of the basic concepts, application scenarios, and security vulnerabilities of MEC networks. It then introduces the types and principles of DDoS attacks faced by MEC networks. Subsequently, various security solutions for DDoS attacks in MEC were detailed and extensively compared, followed by an introduction to current application cases of DDoS defense deployment in practical MEC scenarios. Finally, open issues and future research directions are listed for further exploration.
Yong Ma 0005, Zhiquan Liu 0001, Fagen Li, Qilin Xie, Kaiwei Chen, Chenyang Lv, Ying He 0006
IEEE Internet Things J.8
2025 A Review of Human Emotion Synthesis Based on Generative Technology
abstract
Human emotion synthesis is a crucial aspect of affective computing. It involves using computational methods to mimic and convey human emotions through various modalities, with the goal of enabling more natural and effective human-computer interactions. Recent advancements in generative models, such as Autoencoders, Generative Adversarial Networks, Diffusion Models, Large Language Models, and Sequence-to-Sequence Models, have significantly contributed to the development of this field. However, there is a notable lack of comprehensive reviews in this field. To address this problem, this paper aims to address this gap by providing a thorough and systematic overview of recent advancements in human emotion synthesis based on generative models. Specifically, this review will first present the review methodology, the emotion models involved, the mathematical principles of generative models, and the datasets used. Then, the review covers the application of different generative models to emotion synthesis based on a variety of modalities, including facial images, speech, and text. It also examines mainstream evaluation metrics. Additionally, the review presents some major findings and suggests future research directions, providing a comprehensive understanding of the role of generative technology in the nuanced domain of emotion synthesis.
Fei Ma 0006, Yukan Li, Ying He 0006, Fuji Ren, F. Richard Yu, Shiguang Ni
IEEE Trans. Affect. Comput.4
2025 Industrial Internet of Things With Large Language Models (LLMs): An Intelligence-Based Reinforcement Learning Approach
abstract
Large Language Models (LLMs), as advanced AI technologies for processing and generating natural language text, bring substantial benefits to the Industrial Internet of Things (IIoT) by enhancing efficiency, decision-making, and automation. Nevertheless, their deployment faces significant obstacles due to high computational and energy demands, which often exceed the capabilities of many industrial devices. To overcome these challenges, edge-cloud collaboration has become increasingly essential, assisting in offloading LLMs tasks to reduce the computational load. However, traditional reinforcement learning (RL)-based strategies for LLMs task offloading encounter difficulties with generalization ability and defining explicit, appropriate reward functions. Therefore, in this paper, we propose a novel framework for offloading LLMs inference tasks in IIoT, utilizing a Decentralized Identifier (DID)-based identity management system for trusted task offloading. Furthermore, we introduce an intelligence-based RL (IRL) approach, which sidesteps the need for defining specific reward functions. Instead, it uses “intelligence” as a metric to evaluate cognitive improvements and adapt to varying environmental preferences, significantly improving generalizability. In our experiments, we employ the GPT-J-6B model and utilize the Human Eval dataset to assess its ability to tackle programming challenges, demonstrating the superior performance of our proposed solution compared to existing methods.
Yuzheng Ren, Haijun Zhang 0001, F. Richard Yu, Wei Li 0240, Pincan Zhao, Ying He 0006
IEEE Trans. Mob. Comput.6
2025 Intelligence-Based Reinforcement Learning for Dynamic Resource Optimization in Edge Computing-Enabled Vehicular Networks
abstract
Intelligent transportation systems demand efficient resource allocation and task offloading to ensure low-latency, high-bandwidth vehicular services. The dynamic nature of vehicular environments, characterized by high mobility and extensive interactions among vehicles, necessitates considering time-varying statistical regularities, especially in scenarios with sharp variations. Despite the widespread use of traditional reinforcement learning for resource allocation, its limitations in generalization and interpretability are evident. To overcome these challenges, we propose an Intelligence-based Reinforcement Learning (IRL) algorithm. This algorithm utilizes active inference to infer the real world and maintain an internal model by minimizing free energy. Enhancing the efficiency of active inference, we incorporate prior knowledge as macro guidance, ensuring more accurate and efficient training. By constructing an intelligence-based model, we eliminate the need for designing reward functions, aligning better with human thinking, and providing a method to reflect the learning, information transmission and intelligence accumulation processes. This approach also allows for quantifying intelligence to a certain extent. Considering the dynamic and uncertain nature of vehicular scenarios, we apply the IRL algorithm to environments with constantly changing parameters. Extensive simulations confirm the effectiveness of IRL, significantly improving the generalization and interpretability of intelligent models in vehicular networks.
Yuhang Wang 0019, Ying He 0006, F. Richard Yu, Kaishun Wu, Shanzhi Chen
IEEE Trans. Mob. Comput.2
2024 Resource Allocation for Video Diffusion Task Offloading in Cloud-Edge Networks: A Deep Active Inference Approach
abstract
With the growing popularity and demand for text-to-video generation applications on mobile devices, resource-constrained mobile terminals struggle to efficiently perform video diffusion inference tasks. Cloud-edge computing networks, with the enhanced computational capabilities, flexibility and spectrum utilization, offers a promising solution for improved performance. Traditional Deep Reinforcement Learning (DRL) based methods have been employed for video diffusion inference tasks. However, existing DRL solutions suffer from low data efficiency, insensitivity to latency, and inability to adapt to task load variations, which degrade the performance. In this paper, we propose a novel deep active inference approach for video diffusion inference task offloading and resource allocation in cloud-edge computing networks. Simulation results demonstrate that our method outperforms mainstream DRL in terms of data utilization efficiency and adaptability to cloud-edge networks, better coping with dynamic task load scenarios.
Jiongfeng Fang, Ying He 0006, F. Richard Yu, Jianbo Du
GLOBECOM2
2024 A Novel Meta-Hierarchical Active Inference Reinforcement Learning Approach for QoS-Driven Resource Allocation in Dynamic Clouds
abstract
The cloud computing environment is highly dynamic due to a variety of external factors such as seasonal changes, market trends, and social events. In this case, tenant behavior patterns exhibit significant variability. Therefore, the cloud computing resource allocation model must continuously adapt to these changes. To this end, we propose an adaptive algorithm for dynamic cloud resource allocation based on meta-hierarchical active inference reinforcement learning (MHAIRL). The algorithm combines active inference with meta-hierarchical reinforcement learning. It improves the overall performance of the algorithm, as well as quickly adapts to environmental changes, and improves the generalization performance. In addition, we design a novel polling scheduling framework combined with long short-term memory (LSTM) network. The framework ensures scheduling fairness and flexibility while greatly reducing the state and action space dimensions of the agent. Extensive simulation results show that our method outperforms baseline algorithms in quality of service (QoS) metrics and significantly improves system performance in highly dynamic cloud resource allocation.
Peijie Xian, Ying He 0006, F. Richard Yu, Jianbo Du
GLOBECOM2
2024 A Quantum Temporal Difference Learning Method Based on Quantum World Model
abstract
Based on quantum parallelism theory and quantum phenomena such as superposition and entanglement, quantum reinforcement learning (QRL) has the potential to surpass classical reinforcement learning (RL). Although some excellent works have been done on QRL, existing RL methods encounter chanllenges when performing in environments with sparse rewards. In this paper, we provide a new perspective on conducting temporal difference (TD) learning in quantum computing, which can eliminate redundant exploration steps compared to classical methods. Specifically, we first use environment information to construct a world model with quantum circuits, enabling it to interact in a quantum way. Then, we perform the learning process by using the quantum world model and Grover’s algorithm to query backwards how to reach the recorded states with high TD-errors. Simulation results show that our proposed method has superior performance compared to classical RL algorithms.
Peigen Zeng, Ying He 0006, F. Richard Yu, Jianbo Du
GLOBECOM2
2024 LBVP: Lightweight Blockchain-Based Vehicle Platooning Scheme for Secure and Efficient Platoon Management
Zhiquan Liu 0001, Ying He 0006, Xia Feng, Jianfeng Ma 0001
ICA3PP (6)4
2024 OTOcc: Optimal Transport for Occupancy Prediction
Pengteng Li, Ying He 0006, F. Richard Yu, Pinhao Song, Xingchen Zhou, Guang Zhou
IJCAI2
2024 ABM: Attention before Manipulation
Fan Zhuo, Ying He 0006, F. Richard Yu, Pengteng Li, Zheyi Zhao, Xilong Sun
IJCAI2
2024 PP-TIL: Personalized Planning for Autonomous Driving with Instance-based Transfer Imitation Learning
abstract
Personalized motion planning holds significant importance within urban automated driving, catering to the unique requirements of individual users. Nevertheless, prior endeavors have frequently encountered difficulties in simultaneously addressing two crucial aspects: personalized planning within intricate urban settings and enhancing planning performance through data utilization. The challenge arises from the expensive and limited nature of user data, coupled with the scene state space tending towards infinity. These factors contribute to overfitting and poor generalization problems during model training. Henceforth, we propose an instance-based transfer imitation learning approach. This method facilitates knowledge transfer from extensive expert domain data to the user domain, presenting a resolution to these issues. We initially train a pre-trained model using large-scale expert data. Subsequently, during the fine-tuning phase, we feed the batch data, which comprises expert and user data. Employing the inverse reinforcement learning technique, we extract the style feature distribution from user demonstrations, constructing the regularization term for the approximation of user style. In our experiments, we conducted extensive evaluations of the proposed method. Compared to the baseline methods, our approach mitigates the overfitting issue caused by sparse user data. Furthermore, we discovered that integrating the driving model with a differentiable nonlinear optimizer as a safety protection layer for end-to-end personalized fine-tuning results in superior planning performance. The code will be available at https://github.com/LinFunster/PP-TIL.
Fangze Lin, Ying He 0006, F. Richard Yu
IROS2
2024 LLaKey: Follow My Basic Action Instructions to Your Next Key State
abstract
In 3D object manipulation, collecting expert data for end-to-end imitation learning becomes a mainstream method. Though successful, previous works neglect the guiding role of language in action execution. These methods lack the understanding of action semantics, in which multiple action sequences are guided by a category of instructions, resulting in overlearned object semantics and vague action semantics. To address the above limitation, we introduce a novel framework named LLaKey, which breaks down skill commands into more detailed action instructions based on key states for fine-grained action control. Specifically, LLaKey first leverages the knowledge encoded in pre-trained large-scale models to fine-tune an action instruction conductor. Then, these instructions are executed by a downstream action model. Comprehensive experiments show that LLaKey significantly surpasses baselines with a relative improvement of 15% in nine complex and varied skill tasks, demonstrating the superiority of our method.
Zheyi Zhao, Ying He 0006, F. Richard Yu, Pengteng Li, Fan Zhuo, Xilong Sun
IROS2
2024 A Language-Driven Navigation Strategy Integrating Semantic Maps and Large Language Models
abstract
Accurate perception of semantic and spatial information is crucial for robots performing language-driven navigation tasks. Existing approaches utilize visual-language models to extract semantic information from the environment and construct maps. However, constrained by the generalization and accuracy of these models themselves, the constructed maps may not be accurate and comprehensive, thereby affecting the accuracy of navigation tasks. Inspired by foundational models’ outstanding classification and segmentation capabilities, this study introduces a semantic map constructed using foundational models. We leverage a foundational model to semantically segment objects in the robot’s video stream and fuse semantics onto the map. Furthermore, this map is used in conjunction with large language models (LLMs) that receive natural language instructions to complete the navigation task. A substantial number of experiments in a simulated environment demonstrate that our method outperforms existing ones in language-driven navigation tasks.
Zhengjun Zhong, Ying He 0006, Pengteng Li, F. Richard Yu, Fei Ma 0006
IROS2
2024 OptEx: Expediting First-Order Optimization with Approximately Parallelized Iterations
abstract
First-order optimization (FOO) algorithms are pivotal in numerous computational domains, such as reinforcement learning and deep learning. However, their application to complex tasks often entails significant optimization inefficiency due to their need of many sequential iterations for convergence. In response, we introduce first-order optimization expedited with approximately parallelized iterations (OptEx), the first general framework that enhances the time efficiency of FOO by leveraging parallel computing to directly mitigate its requirement of many sequential iterations for convergence. To achieve this, OptEx utilizes a kernelized gradient estimation that is based on the history of evaluated gradients to predict the gradients required by the next few sequential iterations in FOO, which helps to break the inherent iterative dependency and hence enables the approximate parallelization of iterations in FOO. We further establish theoretical guarantees for the estimation error of our kernelized gradient estimation and the iteration complexity of SGD-based OptEx, confirming that the estimation error diminishes to zero as the history of gradients accumulates and that our SGD-based OptEx enjoys an effective acceleration rate of Θ(√N ) over standard SGD given parallelism of N, in terms of the sequential iterations required for convergence. Finally, we provide extensive empirical studies, including synthetic functions, reinforcement learning tasks, and neural network training on various datasets, to underscore the substantial efficiency improvements achieved by our OptEx in practice.
Yao Shu, Jiongfeng Fang, Ying He 0006, F. Richard Yu
NeurIPS3
2024 SPP-SLAM: Dynamic Visual SLAM with Multiple Constraints based on Semantic Masks and Probabilistic Propagation
abstract
Visual simultaneous localization and mapping (vS-LAM) has attracted great attentions in mobile robots. Most vSLAM systems assume that the objects are stationary in static environments. However, in the real world, there are many objects that are non-stationary in dynamic environments, which will cause performance degradation of vSLAM systems. In this paper, we propose a novel vSLAM system suitable for dynamic environments, named as SPP-SLAM, which is based on semantic masks and probabilistic propagation. Prior motion probabilities of feature points are obtained using semantic mask constraints and multi-view geometric constraints. Then the dynamic probability of each feature point is obtained via the probability propagation model, and highly dynamic feature points are rejected. In addition, we propose a missed detection compensation module combined with inertial measurement unit (IMU) information to recover the semantic masks of the missed objects. Experimental results on the OpenLORIS-Scen and TUM RGB-D datasets demonstrate that the proposed approach can improve the performance of vSLAM systems in a variety of challenging scenarios.
Run Qiu, Ying He 0006, F. Richard Yu, Guang Zhou
WCNC2
2024 Intelligence-based Reinforcement Learning for Continuous Dynamic Resource Allocation in Vehicular Networks
abstract
The rapid advancement of intelligent transportation systems necessitates efficient resource allocation for low-latency and high-bandwidth vehicular services. While traditional reinforcement learning has been widely utilized for resource allocation, it suffers from limitations such as poor generalization and interpretability. To overcome these challenges, we propose a novel Intelligence-based Reinforcement Learning (IRL) al-gorithm, which uses active inference to infer the real world and maintain an internal model of the world by minimizing free energy. We address the inefficiency of active inference by incorporating prior knowledge as macro guidance, ensuring more accurate and efficient training. By constructing the intelligence-based model, we eliminate the need for designing reward functions, which aligns better with human thinking and provides a method to reflect the learning, information transmission, and intelligence accumulation processes. Considering the dynamic and uncertain nature of vehicular scenarios, we apply the IRL algorithm to continuously evolving environments where environmental parameters are not fixed. Extensive simulations confirm the effectiveness of IRL, significantly enhancing the generalization and interpretability of intelligent models.
Yuhang Wang 0019, Ying He 0006, F. Richard Yu, Kaishun Wu
WCNC2
2024 A Novel Internet of Things Web Attack Detection Architecture Based on the Combination of Symbolism and Connectionism AI
abstract
The rapid advancement and wide application of the Internet of Things technology (IoT) have brought unprecedented convenience to people’s production and life. A great number of devices are connected to the IoT network to provide various services for people, which also makes the IoT more vulnerable to various cyber-attacks. This paper designs a novel IoT web attack detection architecture, which combines the powerful knowledge expression ability and high interpretability of symbolic artificial intelligence (AI) with the adaptive learning ability of connectionist AI to form a closed loop of knowledge embedding and extraction, effectively improve the detection ability of web attacks. The architecture solves the “black box” feature of deep learning models and can obtain knowledge from the trained detection model and add it to the training process of the new model to improve detection capabilities. It also uses the advantages of blockchain technology to realize intelligent sharing between different detection systems, solve the problem of difficult detection model updates and training data acquisition “bottlenecks”. To better detect web attacks, we propose a semi-supervised learning method based on an interpretable convolutional neural network (CNN) to reduce misjudgments during self-training and improve detection accuracy. Additionally, we propose a new feature method to extract the features of web logs in IoT devices, which can help the system to detect web attacks in IoT more quickly and accurately. Simulation results on two different datasets show that the proposed architecture and method can effectively detect web attacks in IoT and reduce the false positive rate.
Yufei An, F. Richard Yu, Ying He 0006, Jianqiang Li 0001, Jianyong Chen, Victor C. M. Leung
IEEE Internet Things J.3
2024 Connected and Autonomous Vehicles in Web3: An Intelligence-Based Reinforcement Learning Approach
abstract
“Read-write-own” based Web3 has been proposed as a promising user-centric Internet to open the new generation of the World Wide Web, where Web3 users can independently manage data and derive value from creating content without relying on intermediaries. Connected and autonomous vehicles (CAVs) in Web3 can trade models in a self-controlled and decentralized credible way, which is a fundamentally and principally innovation based on novel architecture. Effectively implementing such paradigms involves proper model trading strategies. However, reinforcement learning (RL)-based strategies face challenges of poor generalization ability, low feasibility, and the exploration-exploitation dilemma. It is also difficult to define an explicit and appropriate reward function. Therefore, in this paper, we propose an intelligence-based reinforcement learning (IRL) approach for CAVs in Web3. We present a framework to enable model transactions between CAVs. Also, we provide a decentralized identifier (DID)-based identity management system for resource description and data verification to access Web3, followed by the mechanism and supporting smart contracts. Furthermore, we formulate the model trading issue as an active inference to form higher-level cognition about the environment without rewards. Then we use IRL to solve it. And we use “intelligence”, a high-level indicator, to quantify the efficiency of such cognition. It can evaluate the difference between the predicted state and the real state in policy exploration. The proposed scheme shows good generalization and can auto-balance exploration and exploitation, simultaneously achieving outperforming performance on the model trading issue with no rewards. In simulations, the performance of the proposed scheme is compared with existing methods.
Yuzheng Ren, Renchao Xie, F. Richard Yu, Ran Zhang 0004, Yuhang Wang 0019, Ying He 0006, Tao Huang 0005
IEEE Trans. Intell. Transp. Syst.6
2024 Large Language Models (LLMs) Inference Offloading and Resource Allocation in Cloud-Edge Computing: An Active Inference Approach
abstract
With the increasing popularity and demands for large language model applications on mobile devices, it is difficult for resource-limited mobile terminals to run large-model inference tasks efficiently. Traditional deep reinforcement learning (DRL) based approaches have been used to offload large language models (LLMs) inference tasks to servers. However, existing DRL solutions suffer from data inefficiency, insensitivity to latency requirements, and non-adaptability to task load variations, which will degrade the performance of LLMs. In this paper, we propose a novel approach based on active inference for LLMs inference task offloading and resource allocation in cloud-edge computing. Extensive simulation results show that our proposed method has superior performance over mainstream DRLs, improves in data utilization efficiency, and is more adaptable to changing task load scenarios.
Ying He 0006, Jingcheng Fang, F. Richard Yu, Victor C. M. Leung
IEEE Trans. Mob. Comput.1
2024 A Deep Learning System for Detecting IoT Web Attacks With a Joint Embedded Prediction Architecture (JEPA)
abstract
The advancement of Internet of Things (IoT) technology has significantly transformed the dynamic between humans and devices, as well as device-to-device interactions. This paradigm shift has led to profound changes in human lifestyles and production processes. Through the interconnectedness of numerous sensors and controllers via networks, the IoT facilitates the seamless integration of humans with diverse devices, leading to substantial economic advantages. Nevertheless, the burgeoning IoT industry and the rapid proliferation of various IoT devices have also introduced a multitude of security vulnerabilities. Cyber attackers frequently exploit cyber attacks to compromise IoT devices, jeopardizing user privacy and property security, thereby posing a grave menace to the overall security of the IoT ecosystem. In this paper, we propose a novel IoT Web attack detection system based on a joint embedded prediction architecture (JEPA), which effectively alleviates the security issues faced by IoT. It can obtain high-level semantic features in IoT traffic data through non-generative self-supervised learning. These features can more effectively distinguish normal data from attack data and help improve the overall detection performance of the system. Moreover, we propose a feature interaction module based on a dual-branch network, which effectively fuses low-level features and high-level features, and comprehensively aggregates global features and local features. Simulation results on multiple datasets show that our proposed system has better detection performance and robustness.
Yufei An, F. Richard Yu, Ying He 0006, Jianqiang Li 0001, Jianyong Chen, Victor C. M. Leung
IEEE Trans. Netw. Serv. Manag.3
2023 A Dynamic Selective Parameter Sharing Mechanism Embedded with Multi-Level Reasoning Abstractions
abstract
Cooperative multi-agent reinforcement learning (Co-MARL) commonly employs different parameter sharing mechanisms, such as full and partial sharing. However, imprudent application of these mechanisms can potentially constrain policy diversity and limit cooperation flexibility. Recent methods that group agents into distinct sharing categories often exhibit poor performance due to challenges in precisely differentiating agents and neglecting the issue of promoting cooperation among these categories. To address these issues, we introduce a dynamic selective parameter sharing mechanism embedded with multi-level reasoning abstractions (DSPS-MA). Our approach uses self-comparison sequences to infer agents’ abstract concepts, defining the differences between agents and allowing them to dynamically select partners to share parameters based on these abstract concepts. We also design an intrinsic reward to offer comprehensive collaboration guidance for agents, and introduce a policy cosine similarity regularization term to ensure sufficient policy diversity. Empirical evaluations demonstrate that our approach yields higher returns and faster convergence than state-of-the-art methods.
Yan Liu 0004, Ying He 0006, Zhong Ming 0001, F. Richard Yu
ECAI2
2023 A Novel Intrusion Detection Architecture for the Internet of Things (IoT) with Knowledge Discovery and Sharing
abstract
The super data transmission capability and connectivity of wireless technologies have promoted the arrival of the Internet of Things (IoT) era. However, the distinct characteristics of IoT devices make them vulnerable to malicious attacks such as hackers and viruses. This paper designs a novel IoT intrusion detection architecture that combines knowledge extraction and sharing, which can extract human understandable knowledge from the trained deep learning model and apply it to the training process of the detection model. The obtained knowledge can also be shared with other detection systems based on the blockchain, which will effectively improve the intrusion detection capabilities of the IoT and realize collective learning. In addition, we propose a CNN-based semi-supervised learning method under the constraints of rules, which can effectively alleviate the catas-trophic interference generated during the self-training process and improve detection accuracy. Simulation results confirm the effectiveness of the proposed architecture and method.
Yufei An, F. Richard Yu, Ying He 0006, Jianqiang Li 0001, Jianyong Chen, Victor C. M. Leung
GLOBECOM3
2023 Task Offloading and Resource Allocation for SLAM Back-End Optimization: A Rewardless Active Inference Approach
abstract
With the increasingly sophisticated algorithms of simultaneous localization and mapping (SLAM), it is difficult for mobile terminals with limited resources to exploit the performance of SLAM algorithms fully. Traditional deep reinforcement learning (DRL)-based approaches have offloaded SLAM tasks to servers. However, existing solutions suffer from low data efficiency and poor generalization problems. This paper proposes a novel approach based on recent advances in rewardless active inference for SLAM back-end optimization. Specifically, the reward function is replaced with simple rewardless guidance in active inference. In addition, instead of simply considering the SLAM task as a whole, we delve into the sub-tasks of back-end optimization of SLAM for offloading and resource allocation. Simulation results show the superior performance of the proposed scheme.
Jingcheng Fang, Ying He 0006, F. Richard Yu, Jianqiang Li 0001, Victor C. M. Leung
GLOBECOM2
2023 Resource Allocation for Cognitive Radio Inspired Non-Orthogonal Multiple Access Networks: A Quantum Soft Actor-Critic Method
abstract
With the growth of the communication industry, the demand for spectrum resources has been increasing steadily. However, the spectrum resources are limited and the utilization rate is low. In this case, Cognitive Radio (CR) technologies and Non-Orthogonal Multiple Access (NOMA) technology are proposed to improve the number of user access and spectrum resource utilization. In this paper, we consider a CR-inspired NOMA network and propose a power allocation problem based on a time-varying system. Due to the dynamic nature of CR, traditional methods are often faced with problems such as low training efficiency and high network volatility. To address the problem, we propose a novel Quantum Reinforcement Learning (QRL) algorithm with quantum world model, interacting with the environment in the quantum way. Specifically, we construct a quantum soft actor-critic algorithm using variable quantum circuit (VQC), and the experience data is encoded into a quantum sum tree and transformed into quantum world model. Simulation results show the superior performance of the proposed framework.
Ying He 0006, F. Richard Yu, Peigen Zeng
GLOBECOM2
2023 Quantum Reinforcement Learning with Quantum World Model
abstract
Quantum reinforcement learning (QRL) can outperform classical reinforcement learning (RL) by utilizing quantum parallel theory and quantum phenomena such as superposition and entanglement. Although some excellent work has been done on QRL, most existing works either fail to show the exponential advantage of quantum computation over classical computation in terms of performance or are too demanding on quantum devices. In this paper, we provide a novel perspective on combining quantum computing and RL with faster convergence speed and relatively relaxed demands on quantum devices. Specifically, we propose a method to construct a world model with quantum circuit that allows it to interact in a quantum way. In addition, we use Grover's algorithm to efficiently extract high-value information from the quantum world model. Extensive simulation results show that the proposed method can have superior performance compared to classical RL algorithms.
Peigen Zeng, Ying He 0006, F. Richard Yu, Victor C. M. Leung
GLOBECOM2
2023 Bagging R-CNN: Ensemble for Object Detection in Complex Traffic Scenes
abstract
Generic object detection methods have achieved preferable results, but it is still challenging to detect objects from complicated traffic scenes like extreme illumination and adverse weather. The existing methods are not robust enough to be extended to new complex traffic scenes. To address this issue, we leverage the idea of ensemble learning for strong robustness and propose a novel Bagging R-CNN framework. Specially, we design a bagging classification branch that uses adaptive sampling to train base learners and make them different from each other. The final predictions are the ensemble of the base learners, achieving strong robustness to the challenging objects. For localizing more accurately, a progressive regression branch is proposed in which bounding boxes are continuously optimized for high quality. Extensive experiment results on TJU-DHD-traffic and Pascal VOC datasets show that our Bagging R-CNN achieves superior detection accuracy over state-of-the-art methods. The source code can be found at https://github.com/PungTeng/BaggingRCNN.
Pengteng Li, Ying He 0006, Dongfu Yin, F. Richard Yu, Pinhao Song
ICASSP2
2023 RePaint-NeRF: NeRF Editting via Semantic Masks and Diffusion Models
abstract
The emergence of Neural Radiance Fields (NeRF) has promoted the development of synthesized high-fidelity views of the intricate real world. However, it is still a very demanding task to repaint the content in NeRF. In this paper, we propose a novel framework that can take RGB images as input and alter the 3D content in neural scenes. Our work leverages existing diffusion models to guide changes in the designated 3D content. Specifically, we semantically select the target object and a pre-trained diffusion model will guide the NeRF model to generate new 3D objects, which can improve the editability, diversity, and application range of NeRF. Experiment results show that our algorithm is effective for editing 3D objects in NeRF under different text prompts, including editing appearance, shape, and more. We validate our method on both real-world datasets and synthetic-world datasets for these editing tasks. Please visit https://repaintnerf.github.io for a better view of our results.
Xingchen Zhou, Ying He 0006, F. Richard Yu, Jianqiang Li 0001
IJCAI2
2023 IGG: Improved Graph Generation for Domain Adaptive Object Detection
abstract
Domain Adaptive Object Detection (DAOD) transfers an object detector from a labeled source domain to a novel unlabeled target domain. Recent works bridge the domain gap by aligning cross-domain pixel-pairs in the non-euclidean graphical space and minimizing the domain discrepancy for adapting semantic distribution. Though great successes, these methods model graphs roughly with coarse semantic sampling due to ignoring the non-informative noises and failing to concentrate on precise semantics alignment. Besides, the coarse graph generation inevitably contains abnormal nodes. These challenges result in biased domain adaptation. Therefore, we propose an Improved Graph Generation (IGG) framework which conducts high-quality graph generation for DAOD. Specifically, we design an Intensive Node Refinement (INR) module that reconstructs the noisy sampled nodes with a memory bank, and contrastively regularizes the noisy features. For better semantics alignment, we decouple the domain-specific style and category-invariant content encoded in graph covariance and selectively eliminate only the domain-specific style. Then, a Precision Graph Optimization (PGO) adaptor is proposed which utilizes the variational inference to down-weight abnormal nodes. Comprehensive experiments on three adaptation benchmarks demonstrate that IGG achieves state-of-the-art results in unsupervised domain adaptation.
Pengteng Li, Ying He 0006, F. Richard Yu, Pinhao Song, Dongfu Yin, Guang Zhou
ACM Multimedia2
2023 Large Language Models (LLMs) Inference Offloading and Resource Allocation in Cloud-Edge Networks: An Active Inference Approach
abstract
As the research and applications of large language model (LLM) become increasingly sophisticated, it is difficult for resource-limited mobile terminals to run large-model inference tasks efficiently. Traditional deep reinforcement learning (DRL) based approaches have been used to offload LLM inference tasks to servers. However, existing solutions suffer from data inefficiency, insensitivity to latency requirements, and non-adaptability to task load variations. In this paper, we propose an active inference with rewardless guidance algorithm using expected future free energy for offloading decisions and allocating resources for the LLM inference task offloading and resource allocation problem of cloud-edge networks systems. Experimental results show that our proposed method has superior performance over mainstream DRLs, improves in data utilization efficiency, and is more adaptable to changing task load scenarios.
Jingcheng Fang, Ying He 0006, F. Richard Yu, Jianqiang Li 0001, Victor C. M. Leung
VTC Fall2
2023 A Novel Visual SLAM System for Autonomous Vehicles in Dynamic Environments
abstract
With the development of autonomous vehicles and intelligent robots, visual simultaneous localization and mapping (SLAM) has attracted great attentions. Most existing visual SLAM systems assume that the objects are stationary in static environments. However, in the real world, there are many objects that are non-stationary in dynamic environments, which will cause performance degradation of visual SLAM systems. In this paper, to address this issue, we propose a novel visual SLAM system based on multi-task deep neural networks. Specifically, we apply multi-task deep neural networks to extract oriented keypoints and perceive dynamic semantic regions, which are used to perform outlier rejection in the SLAM system. We evaluate our method on public datasets, and the results show that our method outperforms existing visual SLAM systems. The presentation video url is: https://youtu.be/qGE1OvaJvV0.
Xinyu Zeng, Ying He 0006, F. Richard Yu, Guang Zhou
VTC Fall2
2023 A Hybrid Driving Decision-Making System Integrating Markov Logic Networks and Connectionist AI
abstract
Connectionist artificial intelligence (AI) can power many critical tasks for connected and autonomous vehicles (CAVs). However, connectionist AI lacks interpretability and usually needs large amount of data for learning. A Markov logic network (MLN), which combines first-order logic (FOL) with statistical learning, learns weighted FOL formulas for inference. MLNs can incorporate domain expert knowledge in the form of FOL formulas to achieve data-efficient learning and transparent decision process. In this paper, we propose a hybrid driving decision-making system, which integrates a MLN module and a deep Q-network (DQN) for enhanced driving safety. The MLN module evaluates the safety of ranked actions from DQN to reduce potential collisions. A collective MLN (Co-MLN) learning algorithm is proposed and it enables CAVs collectively learn a global MLN model for safe state transitions, given distributed small amount of noisy data. A hybrid DQN-MLN learning algorithm is also developed for CAVs to collectively learn to drive in new driving environments. Simulations performed using a highway driving simulator show that the proposed Co-MLN algorithm is highly data-efficient and the learned hybrid driving system can effectively reduce collisions. In addition, the learned MLN module provides transparency for safety-critical driving decisions.
F. Richard Yu, Peter Xiaoping Liu, Ying He 0006
IEEE Trans. Intell. Transp. Syst.4
2023 Efficient Resource Allocation in Multi-UAV Assisted Vehicular Networks With Security Constraint and Attention Mechanism
abstract
With the rapid development of intelligent transportation systems, there is an increasingly strong demand for low-latency and high-bandwidth vehicular services. Unmanned aerial vehicles (UAVs) can be used as a supplement to the ground networks, to relieve the communication pressure on ground facilities, such as base stations. In this paper, we use multiple UAVs to provide services for vehicles and model the multi-UAV scenario as a collaborative multi-agent system. In addition, we take vehicle safety as the top priority and the delay requirement as the constraints. Then we exploit the Lagrange multiplier to combine the constraint function and cost function, so as to reduce the resource consumption as much as possible on the premise of ensuring the safety of the vehicles. The influence of spectrum efficiency and computing power should also be taken into account when allocating resources. We adopt the multi-agent reinforcement learning to train the UAVs, and meanwhile introduce the attention mechanism so that each UAV can optimize itself better with the information of other UAVs. Through extensive simulations, the effectiveness of our proposed method is verified. Particularly, the limited resources can allocated efficiently according to the vehicle’s needs under the premise of ensuring vehicle safety.
Yuhang Wang 0019, Ying He 0006, F. Richard Yu, Qiuzhen Lin, Victor C. M. Leung
IEEE Trans. Wirel. Commun.2
2022 DADEs: 5G Dual-Adaptive Delay-aware and Energy-saving System with Tandem Learning
abstract
Nowadays, numerous primary technologies, like ultra-dense networks (UDNs) and Base Stations (BSs) sleeping state, are developed in fifth-generation (5G) networks. Due to the UDNs, the number of BSs in 5G networks is proliferating, along with the energy consumption. Therefore, it is necessary to cut down the energy attrition in 5G networks under the assurance of delay. Till now, some researchers have proved that the association of users and the sleeping states of BSs have a significant effect on energy consumption and latency in 5G networks. However, the traditional solutions associate users and select states nonadaptively without the dual consideration of energy-saving and delay. In view of this, we propose a dual-adaptive delay-aware and energy-saving system (DADEs) in 5G networks. To further optimize the energy and delay of 5G BSs, the model is split into two tandem problems: user association and BS state selection. Meanwhile, a tandem deep reinforcement learning (T-DRL) algorithm is presented to make decisions in these problems for optimizing and balancing performance between delay and energy adaptively. Additionally, the real datasets of 5G users and BSs are used and trained in this paper. Finally, simulation results show that the DADEs saves more than 50% of energy with an adaptive and satisfying latency.
Chao Qiu, Jingchao Tan, Xiaofei Wang 0001, Yajun Yang, Ying He 0006, Jing Jiang 0026
GLOBECOM6
2022 Multi-Constraint Deep Reinforcement Learning for Smooth Action Control
abstract
Deep reinforcement learning (DRL) has been studied in a variety of challenging decision-making tasks, e.g., autonomous driving. \textcolor{black}{However, DRL typically suffers from the action shaking problem, which means that agents can select actions with big difference even though states only slightly differ.} One of the crucial reasons for this issue is the inappropriate design of the reward in DRL. In this paper, to address this issue, we propose a novel way to incorporate the smoothness of actions in the reward. Specifically, we introduce sub-rewards and add multiple constraints related to these sub-rewards. In addition, we propose a multi-constraint proximal policy optimization (MCPPO) method to solve the multi-constraint DRL problem. Extensive simulation results show that the proposed MCPPO method has better action smoothness compared with the traditional proportional-integral-differential (PID) and mainstream DRL algorithms. The video is available at https://youtu.be/F2jpaSm7YOg.
Guangyuan Zou, Ying He 0006, F. Richard Yu, Longquan Chen, Weike Pan, Zhong Ming 0001
IJCAI2
2022 $Q_{C}-DQN$: A Novel Constrained Reinforcement Learning Method for Computation Offloading in Multi-access Edge Computing
abstract
In recent years, multi-access edge computing (MEC) is emerging to provide computation and storage resources to the Internet of things (IoT) devices to assist them in high-performance demanding tasks. Real-time task requests from the IoT devices often have strict delay constraints. However, in practice, the delay requirements of task requests often fail to be satisfied because of the inappropriate computation processing method and inefficient resource allocation in MEC networks. In this article, we present a novel framework for MEC networks with unmanned aerial vehicles (UAVs) and intelligent reflecting surfaces (IRSs) to facilitate computation offloading with delay constraints. In addition, we propose a novel constrained reinforcement learning method with a dynamic balance mechanism named$Q_{c}-DQN$. Finally, we conduct extensive simulations to verify the effectiveness of our proposed method. Compared to the benchmark schemes, our scheme not only improves the overall network performance and reduces the task completion time, but also meets the delay constraints.
Shen Zhuang, Chengxi Gao, Ying He 0006, F. Richard Yu, Yuhang Wang 0019, Weike Pan, Zhong Ming 0001
IJCNN3
2022 When Multi-access Edge Computing Meets Multi-area Intelligent Reflecting Surface: A Multi-agent Reinforcement Learning Approach
abstract
In recent years, multi-access edge computing (MEC) is emerging to provide computation and storage capabilities to the Internet of things (IoT) devices to improve the quality of service (QoS) of IoT applications. In addition, intelligent reflecting surface (IRS) techniques have attracted great interests from both academia and industry to improve the communication efficiency. Although existing works leverage the IRS technique in MEC networks, they mainly focus on the single-IRS single-area scenario. However, in practice, multi-IRS will be deployed in multi-area scenarios in future networks. Consequently, considering the single-IRS single-area scenario will have inferior performance. In this paper, to address the aforementioned issue, we propose an efficient resource provisioning scheme for multi-IRS multi-area scenarios in MEC networks. We first model the problem as a cooperative multi-agent reinforcement learning process, where each agent manages one area and all agents share the network bandwidth and computation resources. Then, we propose a multi-agent actor-critic method with an attention mechanism for resource management with latency guarantee. Finally, we conduct extensive simulations to verify the effectiveness of the proposed scheme. Our scheme can reduce the required computation resources by up to 11.84% when compared with the benchmark works. It is also shown that our proposed scheme can improve the efficiency of resource allocation and scale well with the increasing demand from IoT devices.
Shen Zhuang, Ying He 0006, F. Richard Yu, Chengxi Gao, Weike Pan, Zhong Ming 0001
IWQoS2
2022 Bift: A Blockchain-Based Federated Learning System for Connected and Autonomous Vehicles
abstract
Machine learning (ML) algorithms are essential components in autonomous driving. In most existing connected and autonomous vehicles (CAVs), a large amount of driving data collected from multiple vehicles are sent to a central server for unified training. However, data privacy and security have become crucial during the data-sharing process. Federated learning (FL) for data security has arisen nowadays, and it can improve the data privacy of distribute machine learning. However, the malicious attackers can still be able to attack the training process. Due to the complete reliance on the central server, FL is very fragile. To address the above problem, we propose Bift: 1) a fully decentralized ML system combined with FL and 2) blockchain to provide a privacy-preserving ML process for CAVs. Bift enables distributed CAVs to train ML models locally using their own driving data and then to upload the local models to get a better global model. More importantly, Bift provides a consensus algorithm named Proof of Federated Learning to resist possible adversaries. We evaluate the performance of Bift and demonstrate that Bift is scalable and robust, and can defend against malicious attacks.
Ying He 0006, Guangzheng Zhang, F. Richard Yu, Jianyong Chen, Jianqiang Li 0001
IEEE Internet Things J.1
2022 RTT-Based Rogue UAV Detection in IoV Networks
abstract
Unmanned aerial vehicles (UAVs) are being used in different emerging domains for accomplishing many critical tasks. However, due to the various constraints, such as battery life, computational resources, etc., a UAV under a mission (M-UAV) often needs assistance from an edge/cloud server that is reachable from the M-UAV’s location. A connection between an M-UAV and edge server can be established via an access point or AP. Therefore, before sharing any sensitive information with the edge server, it is essential for an M-UAV to determine the legitimacy of the selected AP. Recently, some works in this direction indicate that a rogue UAV (R-UAV) can successfully mimic a legitimate AP for intercepting the communication channel. Hence, there should be a robust detection mechanism in place for addressing such a threat scenario. In this article, considering one of the emerging domains—the Internet of Vehicle (IoV) networks, at first, we show that communication in the IoV networks can get benefit from the presence of M-UAVs. However, as the link between the M-UAV and edge server can be intercepted by an R-UAV, the adversary may access the sensitive information from the IoV networks. Followed by this, we propose atiming-basedalgorithm for identifying the presence of rogue APs (or R-UAVs) in the channel. The M-UAV executes the timing-based algorithm, and the detection methoddoes notrequire any auxiliary hardware or any modification to the network protocols for meeting the objective. Supported by an extensive evaluation study, we show that without any rigid restriction on the M-UAV’s speed (e.g., by limiting it to almost static) the proposed approach significantly enhances the detection accuracy (at least by a margin of 29.7% and 16.65%) compared to the state-of-the-art methods.
Nilesh Chakraborty, Yao Chao, Jianqiang Li 0001, Sumit Mishra, Chengwen Luo 0001, Ying He 0006, Jie Chen 0027, Yi Pan 0001
IEEE Internet Things J.6
2022 An Efficient Ciphertext-Policy Attribute-Based Encryption Scheme Supporting Collaborative Decryption With Blockchain
abstract
In the last few decades, ciphertext-policy attribute-based encryption (CP-ABE) technology has attracted great interest, since it can provide fine-grained, flexible, and access control for sensitive data to implement a high secure and efficient data-sharing mechanism. In this article, based on the linear secret sharing scheme (LSSS), an efficient scheme is proposed to realize a collaborative decryption function. For any user group, when the user’s attribute set cannot access the ciphertext alone, the private key of other users in the same group can be used for collaborative decryption with the permission of the data owner. Our scheme uses the LSSS matrix that can significantly reduce the computation and storage overhead when comparing with the existing schemes. Then, a multiauthorization model is created based on the Bohen–Lynn–Shacham technology in order to solve the key-management issue. Finally, we implemented the specific functions of the framework through JAVA, and built a private chain to verify the feasibility of data transfer between users.
Ying He 0006, Haiyan Wang 0009, Victor C. M. Leung, F. Richard Yu, Zhong Ming 0001
IEEE Internet Things J.1
2022 Efficient Resource Allocation for Multi-Beam Satellite-Terrestrial Vehicular Networks: A Multi-Agent Actor-Critic Method With Attention Mechanism
abstract
With the rapid development of intelligent transportation systems, there is an increasing demand for a variety of vehicular services, such as automated driving assistance, emergency alert, infotainment, etc. However, in some situations (e.g., remote areas or maritime scenarios), the terrestrial networks alone cannot serve the vehicular applications very well due to the infrastructure deployment and maintenance issues. Satellite networks have become an effective supplement to terrestrial networks, which complement well in terms of coverage, flexibility, reliability, and availability. In this paper, we consider the low orbit multi-beam satellite-terrestrial networks to serve for vehicles. We model this problem as a cooperative multi-agent reinforcement learning process, where each beam acts as an agent, and the global bandwidth is cooperatively shared among all the agents. A multi-agent actor-critic method with attention mechanism is proposed to allocate resources for vehicles with strict delay requirements and minimum bandwidth consumption. When allocating bandwidth, the channel efficiency, the angle of the beams and the priorities of requests in different regions are also considered. Centralized training and distributed execution is performed in the training of the agents. Extensive simulation results verify the effectiveness of our proposed method, where all the agents can well cooperative to achieve efficient resource allocation on-demand for the vehicles under strictly limited bandwidth resources.
Ying He 0006, Yuhang Wang 0019, F. Richard Yu, Qiuzhen Lin, Jianqiang Li 0001, Victor C. M. Leung
IEEE Trans. Intell. Transp. Syst.1
2021 A Fast-adaptive Edge Resource Allocation Strategy for Dynamic Vehicular Networks
abstract
With the rapid development of vehicular networks, there is an increasing demand for extensive networking, computing and caching resources. In fact, vehicular networks are nonstationary, and how to allocate multiple resources effectively and efficiently for dynamic vehicular networks is extremely important, however, really challenging. In this paper, we propose a general framework that can enable fast-adaptive edge resource allocation for dynamic vehicular environment. Specifically, we model the dynamics of the vehicular environment as a series of related Markov Decision Processes (MDPs). We combine hierarchical reinforcement learning with meta learning, which makes our proposed framework available to quickly adapt to a new environment by only fine-tuning the top-level master network, and meanwhile the low-level sub-networks can make the right resource allocation policy. The extensive simulation results show the effectiveness of our proposed framework, which can quickly adapt to different scenarios. This is consistent with the real-world situations and can significantly improve the performance of resource allocation in dynamic vehicular networks.
Ying He 0006, Yuhang Wang 0019, Qiuzhen Lin, Jianqiang Li 0001, Victor C. M. Leung
ICNP1
2021 Blockchain-Based Edge Computing Resource Allocation in IoT: A Deep Reinforcement Learning Approach
abstract
With the exponential growth in the number of Internet-of-Things (IoT) devices, the cloud-centric computing paradigm can hardly meet the increasingly high requirements for low latency, high bandwidth, ease of availability, and more intelligent services. Therefore, a distributed and decentralized computing architecture is imperative, where edge-centric computing, such as fog computing and mist computing, has been recently proposed. Edge-centric computing resources can be managed locally and personally rather than being administered by a remote centralized third party. However, security and privacy issues are the main challenges due to the absence of trust between the IoT devices and edge computing nodes (ECNs). A blockchain, as a decentralized, trustless, and immutable public ledger, can well solve the trust-absence issue. In this article, we first elaborate on the security and privacy issues of edge-computing-enabled IoT, and then present the key characteristics of blockchains, which make blockchains well suited for the edge-centric IoT scenarios. Furthermore, we propose a general framework for blockchain-based edge-computing-enabled IoT scenarios that specifies the step-by-step procedure of a single transaction between an IoT end and an ECN. In addition, we design a smart contract within a private blockchain network that exploits the state-of-the-art machine learning algorithm, asynchronous advantage actor-critic (A3C), to allocate the edge computing resources, which exemplifies how artificial intelligence (AI) can be combined with blockchains. We further discuss the benefits of the convergence of AI and blockchains. Finally, simulation results are presented.
Ying He 0006, Yuhang Wang 0019, Chao Qiu, Qiuzhen Lin, Jianqiang Li 0001, Zhong Ming 0001
IEEE Internet Things J.1
2020 Improving Approximate Logic Neuron Model by Means of a Novel Learning Algorithm
Jiajun Zhao, Minhui Dong, Cheng Tang 0001, Junkai Ji, Ying He 0006
ICIC (1)5
2019 Trust management for secure cognitive radio vehicular ad hoc networks
Ying He 0006, F. Richard Yu, Zhexiong Wei, Victor C. M. Leung
Ad Hoc Networks1
2018 Integrated Computing, Caching, and Communication for Trust-Based Social Networks: A Big Data DRL Approach
abstract
Recent advances of computing, caching, and communication (3C) can have significant impacts on mobile social networks (MSNs). MSNs can leverage these new paradigms to provide a new mechanism for users to share resources (e.g., information, computation-based services). In this paper, we exploit the intrinsic nature of social networks, i.e., the trust formed through social relationships among users, to enable users to share resources under the framework of 3C. Specifically, we consider the mobile edge computing (MEC), in-network caching and device-to-device (D2D) communications. When considering the trust-based MSNs with MEC, caching and D2D, we apply a novel big data deep reinforcement learning (DRL) approach to automatically make a decision for optimally allocating the network resources. The decision is made purely through observing the network's states, rather than any handcrafted or explicit control rules, which makes it adaptive to variable network conditions. Google TensorFlow is used to implement the proposed deep Q-learning approach. Simulation results with different network parameters are presented to show the effectiveness of the proposed scheme.
Ying He 0006, Chengchao Liang, F. Richard Yu, Victor C. M. Leung
GLOBECOM1
2018 Enhancing Video Rate Adaptation With Mobile Edge Computing and Caching in Software-Defined Mobile Networks
abstract
Recent advances in software-defined mobile networks (SDMNs), in-network caching, and mobile edge computing (MEC) can have significant effects on video services in next generation mobile networks. In this paper, we jointly consider SDMNs, in-network caching, and MEC to enhance the video service in next generation mobile networks. We use a new video experience evaluation standard called U-video mean opinion score (vMOS), which is a more advanced measurement of the video quality based on the well-known vMOS. With the objective of maximizing the mean U-vMOS, an optimization problem is formulated. Due to the coupling of video data rate, computing resource, and traffic engineering (bandwidth provisioning and paths selection), the problem becomes intractable in practice. Thus, we utilize a dual-decomposition method to decouple those three sets of variables. By this decoupling, video rate adaptation is performed at users with network assistants. End nodes can schedule computing resource independently. Traffic engineering is performed by the software-defined networking controller and base stations. Furthermore, to address the challenges of dynamic change of network status and the drawbacks caused by the frequent exchange of information, we design a decentralized algorithm based on alternating direction method of multipliers to solve the traffic engineering problem. Extensive simulations are conducted with different system configurations to show the effectiveness of the proposed scheme.
Chengchao Liang, Ying He 0006, F. Richard Yu, Nan Zhao 0001
IEEE Trans. Wirel. Commun.2
2017 A Big Data Deep Reinforcement Learning Approach to Next Generation Green Wireless Networks
abstract
Recent advances in networking, caching and computing technologies can have great impacts on the developments of green heterogeneous wireless networks, where different sizes of cells co-exist. Nevertheless, these important enabling technologies have traditionally been studied separately in the existing works on wireless networks. In this paper, we propose an integrated framework that can enable dynamic orchestration of networking, caching and computing resources to improve the performance of green heterogeneous wireless networks. We use an energy-efficient caching strategy based on storing maximum-distance separable (MDS) encoded packets. The resource allocation strategy in this framework is formulated as a joint optimization problem. The decision on how to allocate the dynamic resources is very complicated when considering networking, caching and computing. Therefore, we propose a novel deep reinforcement learning approach, which can effectively handle systems with large complexity. In addition, we use Google TensorFlow to implement deep reinforcement learning. Simulation results with different system parameters are presented to show the effectiveness of the proposed scheme.
Ying He 0006, Zheng Zhang 0037, Yanhua Zhang
GLOBECOM1
2017 Optimization of cache-enabled opportunistic interference alignment wireless networks: A big data deep reinforcement learning approach
abstract
Both caching and interference alignment (IA) are promising techniques for future wireless networks. Nevertheless, most of existing works on cache-enabled IA wireless networks assume that the channel is invariant, which is unrealistic considering the time-varying nature of practical wireless environments. In this paper, we consider realistic time-varying channels. Specifically, the channel is formulated as a finite-state Markov channel (FSMC). The complexity of the system is very high when we consider realistic FSMC models. Therefore, we propose a novel big data reinforcement learning approach in this paper. Deep reinforcement learning is an advanced reinforcement learning algorithm that uses deep Q network to approximate the Q value-action function. Deep reinforcement learning is used in this paper to obtain the optimal lA user selection policy in cache-enabled opportunistic lA wireless networks. Simulation results are presented to show the effectiveness of the proposed scheme.
Ying He 0006, Chengchao Liang, F. Richard Yu, Nan Zhao 0001, Hongxi Yin
ICC1
2017 Resource Allocation in Software-Defined and Information-Centric Vehicular Networks with Mobile Edge Computing
abstract
Recent advances in networking, caching and computing have significant impacts on the developments of vehicular networks. Nevertheless, these important enabling technologies have traditionally been studied separately in the existing works on vehicular networks. In this paper, we propose an integrated framework that can enable dynamic orchestration of networking, caching and computing resources to improve the performance of next generation vehicular networks. We formulate the resource allocation strategy in this framework as a joint optimization problem. The complexity of the system is very high when we jointly consider these three technologies. Therefore, we propose a novel deep reinforcement learning approach in this paper. Simulation results are presented to show the effectiveness of the proposed scheme.
Ying He 0006, Chengchao Liang, Zheng Zhang 0037, F. Richard Yu, Nan Zhao 0001, Hongxi Yin, Yanhua Zhang
VTC Fall1
2017 Video Rate Adaptation and Traffic Engineering in Mobile Edge Computing and Caching-Enabled Wireless Networks
abstract
Recent advances in software-defined mobile networks (SDMNs), in-network caching, and mobile edge computing (MEC) can have great effects on video services in next generation mobile networks. In this paper, we jointly consider SDMNs, in- network caching, and MEC to enhance the video service in next generation mobile networks. With the objective of maximizing the mean measurement of video quality, an optimization problem is formulated. Due to the coupling of video data rate, computing resource, and traffic engineering (bandwidth provisioning and paths selection), the problem becomes intractable in practice. Thus, we utilize dual-decomposition method to decouple those three sets of variables. Extensive simulations are conducted with different system configurations to show the effectiveness of the proposed scheme.
Chengchao Liang, Ying He 0006, F. Richard Yu, Nan Zhao 0001
VTC Fall2
2017 Enhancing QoE-Aware Wireless Edge Caching With Software-Defined Wireless Networks
abstract
Software-defined networking and in-network caching are promising technologies in the next generation wireless networks. In this paper, we propose enhancing the quality of experience (QoE)-aware wireless edge caching with bandwidth provisioning in software-defined wireless networks (SDWNs). Specifically, we design a novel mechanism to jointly provide proactive caching, bandwidth provisioning, and adaptive video streaming. The caches are requested to retrieve data in advance dynamically according to the behaviors of users, the current traffic, and the resource status. Then, we formulate a novel optimization problem regarding the QoE-aware bandwidth provisioning in SDWNs with jointly considering in-network caching strategy. The caching problem is decoupled from the bandwidth provisioning problem by deploying the dual-decomposition method. Additionally, we relax the binary variables to real numbers so that those two problems are formulated as a linear problem and a convex problem, respectively, which can be solved efficiently. Simulation results are presented to show that the latency is decreased and the utilization of caches is improved in the proposed scheme.
Chengchao Liang, Ying He 0006, F. Richard Yu, Nan Zhao 0001
IEEE Trans. Wirel. Commun.2