EDBT 2026 Demo / reviewers in the wild / expert
Yuning Jiang 0002
dblp:99/7950-2
· DBLP profile ↗
16ranked-venue papers
0as first author
15since 2021 · last 2026
0000-0002-7145-0995ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 11 · 11 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Federated Linear Bandit Learning via UAV Aided Over-the-Air ComputationabstractThis paper investigates federated contextual linear bandit learning in a wireless network with a central server and multiple devices. To reduce communication latency, devices interact with the server via over-the-air computation (AirComp) over noisy, fading channels, where signal distortion can occur due to channel imperfections. Departing from traditional AirComp designs for static networks, we propose a novel federated bandit learning framework that leverages unmanned aerial vehicles (UAVs) as mobile servers to aggregate data from distributed IoT devices. To optimize this system, we employ a block coordinate descent method combined with the alternating direction method of multipliers (BCD-ADMM), jointly optimizing the UAV trajectory, receive normalization factor, and transmission power to minimize the time-averaged mean square error (MSE) of AirComp. Our approach addresses the challenge of decentralized data across multiple devices, enabling secure and efficient collaboration without direct data sharing. Theoretical analysis establishes an upper bound on the algorithm's regret, affirming the framework's scalability and robustness against noise. Simulation results support these findings, highlighting notable performance improvements in federated bandit learning with UAV-assisted AirComp. Junkai Qian, Yuning Jiang 0002, Xin Liu 0049, Ting Wang 0001, Yuanming Shi, Colin N. Jones |
IEEE Trans. Mob. Comput. | 2 |
| 2026 | Microservice Deployment in Space Computing Power Networks Via Robust Reinforcement LearningabstractWith the growing demand for Earth observation, it is important to provide reliable real-time remote sensing inference services to meet the low-latency requirements. The Space Computing Power Network (Space-CPN) offers a promising solution by providing onboard computing and extensive coverage capabilities for real-time inference. This paper presents a remote sensing artificial intelligence applications deployment framework designed for Low Earth Orbit satellite constellations to achieve real-time inference performance. The framework employs the microservice architecture, decomposing monolithic inference tasks into reusable, independent modules to address high latency and resource heterogeneity. This distributed approach enables optimized microservice deployment, minimizing resource utilization while meeting quality of service and functional requirements. We introduce Robust Optimization to the deployment problem to address data uncertainty. Additionally, we model the Robust Optimization problem as a Partially Observable Markov Decision Process and propose a robust reinforcement learning algorithm to handle the semi-infinite Quality of Service constraints. Our approach yields sub-optimal solutions that minimize accuracy loss while maintaining acceptable computational costs. Simulation results demonstrate the effectiveness of our framework. Yuning Jiang 0002, Xin Liu 0049, Yuanming Shi, Chunxiao Jiang, Linling Kuang |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | SAI: Latency-Aware Satellite Edge LAM Inference with Looped TransformerabstractThe rapid advancements in computing and communication capabilities of Low Earth Orbit (LEO) satellites have made it feasible to execute complex and collaborative inorbit computation missions. Transformer-based large AI models (LAMs), known for their exceptional performance in in-context learning (ICL) and prompt-based reasoning, have attracted significant attention, providing powerful intelligence across sectors such as industry and aerospace. However, the significant parameter volume of LAMs poses a substantial challenge for direct deployment on satellites with constrained computing power and energy provision. To address this, the looped Transformer model reduces parameter requirements through layerwise parameter sharing, achieving performance comparable to vanilla Transformer-based LAMs in ICL tasks. Despite this efficiency, the limited and heterogeneous space-borne computing and storage capabilities complicate the orchestration for balanced workload allocation during multi-satellite cooperation. In this paper, we propose SAI, a collaborative multi-satellite space AI system that exploits the memory efficiency of the looped Transformer and the inherent parallelism in batch data processing. SAI enables accelerated on-satellite inference by integrating heterogeneous onboard resources and introducing a novel hybrid approach combining data and pipeline parallelism. This approach supports cross-satellite cooperation with parallelism planning and asynchronous inter-batch overlapping, significantly reducing inference latency and enhancing resource efficiency. Furthermore, SAI optimizes inference latency by formulating it as a shortest-path problem, effectively solved via Dijkstras algorithm. Extensive evaluations demonstrate SAIs superior performance in reducing inference latency and runtime memory usage compared to existing baselines. Honggang Yuan, Yuning Jiang 0002, Xin Liu 0049, Yuanming Shi, Ting Wang 0001 |
ICC | 3 |
| 2024 | Dynamic Communication in Multi-Agent Reinforcement Learning via Information BottleneckabstractEffective information sharing is essential for multi-agent systems to execute cooperative tasks successfully. Typically, agents within such systems are either stationary or possess unrestricted communication ranges. However, in more complex scenarios where agent mobility is introduced, the communication network’s topology becomes dynamic over time. This dynamism can result in partial communication unreachability among certain agents. Consequently, striking a balance between minimizing overall communication overhead and optimizing task performance becomes a formidable challenge. In this paper, we address the issue of dynamic communication in multi-agent systems. We propose a novel approach that leverages the principle of information bottleneck theory to develop a multi-mean field multi-agent reinforcement learning algorithm called MMIB. Through a series of experiments, we demonstrate the effectiveness of our proposed algorithm in reducing communication overhead while maintaining task performance at a level comparable to other state-of-the-art multi-agent reinforcement learning algorithms. Jiawei You, Youlong Wu, Dingzhu Wen, Yong Zhou 0006, Yuning Jiang 0002, Yuanming Shi |
GLOBECOM | 5 |
| 2024 | Principled Preferential Bayesian OptimizationabstractWe study the problem of preferential Bayesian optimization (BO), where we aim to optimize a black-box function with only preference feedback over a pair of candidate solutions. Inspired by the likelihood ratio idea, we construct a confidence set of the black-box function using only the preference feedback. An optimistic algorithm with an efficient computational method is then developed to solve the problem, which enjoys an information-theoretic bound on the total cumulative regret, a first-of-its-kind for preferential BO. This bound further allows us to design a scheme to report an estimated best solution, with a guaranteed convergence rate. Experimental results on sampled instances from Gaussian processes, standard test functions, and a thermal comfort optimization problem all show that our method stably achieves better or competitive performance as compared to the existing state-of-the-art heuristics, which, however, do not have theoretical guarantees on regret bounds or convergence. Yuning Jiang 0002, Bratislav Svetozarevic, Colin N. Jones |
ICML | 3 |
| 2024 | Microservice Deployment for Satellite Edge AI Inference via Deep Reinforcement LearningabstractArtificial intelligence (AI) is critical in evolving 5G and developing 6G networks, running on edge devices, and solving resource management challenges. The burgeoning number of edge devices draws attention to the potential of low-earth orbit (LEO) satellite networks with their onboard computing capabilities for edge inference. This paper explores LEO scenarios where multiple remote sensing edge AI inference tasks concurrently process data from a single source. However, due to there being parts with the same functions between different AI applications, traditional monolithic edge AI architecture must be deployed repeatedly and falls short in efficiently harnessing the heterogeneous resources of LEO satellite networks. To solve this problem, we utilize the microservice architecture to decouple a single AI application into several independent microservices to reuse these same functions. However, due to the high latency caused by multiple microservices’ communication, we need to design a deployment strategy to fully utilize resources to reduce the service latency. We present a microservice deployment model to minimize the total service latency across all AI applications and meet resource constraints with the constraints of hardware, energy, and memory limitations. This latency optimization problem is rewritten as a Markov decision process (MDP) to effectively deal with the challenge posed by the time-varying transmission rate caused by satellite mobility. To increase the training data utilization, we employ a Proximal Policy Optimization (PPO) based reinforcement learning algorithm to meet the dynamic environment challenge. Finally, we obtain a sub-optimal solution with minimal accuracy loss and an acceptable solution time. Hei Victor Cheng, Zhanpeng Yang, Xin Liu 0049, Yuning Jiang 0002, Yong Zhou 0006, Yuanming Shi |
PIMRC | 5 |
| 2024 | Federated Reinforcement Learning for Electric Vehicles Charging Control on Distribution NetworksabstractWith the growing popularity of electric vehicles (EVs), maintaining power grid stability has become a significant challenge. To address this issue, EV charging control strategies have been developed to manage the switch between vehicle-to-grid (V2G) and grid-to-vehicle (G2V) modes for EVs. In this context, multiagent deep reinforcement learning (MADRL) has proven its effectiveness in EV charging control. However, existing MADRL-based approaches fail to consider the natural power flow of EV charging/discharging in the distribution network and ignore driver privacy. To deal with these problems, this article proposes a novel approach that combines multi-EV charging/discharging with a radial distribution network (RDN) operating under optimal power flow (OPF) to distribute power flow in real time. A mathematical model is developed to describe the RDN load. The EV charging control problem is formulated as a Markov decision process (MDP) to find an optimal charging control strategy that balances V2G profits, RDN load, and driver anxiety. To effectively learn the optimal EV charging control strategy, a federated deep reinforcement learning algorithm named FedSAC is further proposed. Comprehensive simulation results demonstrate the effectiveness and superiority of our proposed algorithm in terms of the diversity of the charging control strategy, the power fluctuations on RDN, the convergence efficiency, and the generalization ability. Junkai Qian, Yuning Jiang 0002, Xin Liu 0049, Ting Wang 0001, Yuanming Shi, Wei Chen 0002 |
IEEE Internet Things J. | 2 |
| 2024 | Decentralized Over-the-Air Federated Learning by Second-Order Optimization MethodabstractFederated learning (FL) is an emerging technique that enables privacy-preserving distributed learning. Most related works focus on centralized FL, which leverages the coordination of a parameter server to implement local model aggregation. However, this scheme heavily relies on the parameter server, which could cause scalability, communication, and reliability issues. To tackle these problems, decentralized FL, where information is shared through gossip, starts to attract attention. Nevertheless, current research mainly relies on first-order optimization methods that have a relatively slow convergence rate, which leads to excessive communication rounds in wireless networks. To design communication-efficient decentralized FL, we propose a novel over-the-air decentralized second-order federated algorithm. Benefiting from the fast convergence rate of the second-order method, total communication rounds are significantly reduced. Meanwhile, owing to the low-latency model aggregation enabled by over-the-air computation, the communication overheads in each round can also be greatly decreased. The convergence behavior of our approach is then analyzed. The result reveals an error term, which involves a cumulative noise effect, in each iteration. To mitigate the impact of this error term, we conduct system optimization from the perspective of the accumulative term and the individual term, respectively. Numerical experiments demonstrate the superiority of our proposed approach and the effectiveness of system optimization. Peng Yang 0027, Yuning Jiang 0002, Dingzhu Wen, Ting Wang 0001, Colin N. Jones, Yuanming Shi |
IEEE Trans. Wirel. Commun. | 2 |
| 2024 | One-Bit Byzantine-Tolerant Distributed Learning via Over-the-Air ComputationabstractDistributed learning has become a promising computational parallelism paradigm that enables a wide scope of intelligent applications from the Internet of Things (IoT) to autonomous driving and the healthcare industry. This paper studies distributed learning in wireless data center networks, which contain a central edge server and multiple edge workers to collaboratively train a shared global model and benefit from parallel computing. However, the distributed nature causes the vulnerability of the learning process to faults and adversarial attacks from Byzantine edge workers, as well as the severe communication and computation overhead induced by the periodical information exchange process. To achieve fast and reliable model aggregation in the presence of Byzantine attacks, we develop a signed stochastic gradient descent (SignSGD)-based Hierarchical Vote framework via over-the-air computation (AirComp), where one voting process is performed locally at the wireless edge by taking advantage of Bernoulli coding while the other is operated over-the-air at the central edge server by utilizing the waveform superposition property of the multiple-access channels. We comprehensively analyze the proposed framework on the impacts including Byzantine attacks and the wireless environment (channel fading and receiver noise), followed by characterizing the convergence behavior under non-convex settings. Simulation results validate our theoretical achievements and demonstrate the robustness of our proposed framework in the presence of Byzantine attacks and receiver noise. Youlong Wu, Yuning Jiang 0002, Yuanming Shi |
IEEE Trans. Wirel. Commun. | 3 |
| 2023 | Federated Linear Bandit Learning via Over-the-air ComputationabstractIn this paper, we investigate federated contextual linear bandit learning within a wireless system that comprises a server and multiple devices. Each device interacts with the environment, selects an action based on the received reward, and sends model updates to the server. The primary objective is to minimize cumulative regret across all devices within a finite time horizon. To reduce the communication overhead, devices communicate with the server via over-the-air computation (AirComp) over noisy fading channels, where the channel noise may distort the signals. In this context, we propose a customized federated linear bandits scheme, where each device transmits an analog signal, and the server receives a superposition of these signals distorted by channel noise. A rigorous mathematical analysis is conducted to determine the regret bound of the proposed scheme. Both theoretical analysis and numerical experiments demonstrate the competitive performance of our proposed scheme in terms of regret bounds in various settings. Yuning Jiang 0002, Xin Liu 0049, Ting Wang 0001, Yuanming Shi |
GLOBECOM | 2 |
| 2023 | Constrained Efficient Global Optimization of Expensive Black-box FunctionsabstractWe study the problem of constrained efficient global optimization, where both the objective and constraints are expensive black-box functions that can be learned with Gaussian processes. We propose CONFIG (CONstrained efFIcient Global Optimization), a simple and effective algorithm to solve it. Under certain regularity assumptions, we show that our algorithm enjoys the same cumulative regret bound as that in the unconstrained case and similar cumulative constraint violation upper bounds. For commonly used Matern and Squared Exponential kernels, our bounds are sublinear and allow us to derive a convergence rate to the optimal solution of the original constrained problem. In addition, our method naturally provides a scheme to declare infeasibility when the original black-box optimization problem is infeasible. Numerical experiments on sampled instances from the Gaussian process, artificial numerical problems, and a black-box building controller tuning problem all demonstrate the competitive performance of our algorithm. Compared to the other state-of-the-art methods, our algorithm significantly improves the theoretical guarantees while achieving competitive empirical performance. Yuning Jiang 0002, Bratislav Svetozarevic, Colin N. Jones |
ICML | 2 |
| 2023 | Self-triggered Control with Energy Harvesting Sensor NodesabstractDistributed embedded systems are pervasive components jointly operating in a wide range of applications. Moving toward energy harvesting powered systems enables their long-term, sustainable, scalable, and maintenance-free operation. When these systems are used as components of an automatic control system to sense a control plant, energy availability limits when and how often sensed data are obtainable and therefore when and how often control updates can be performed. The time-varying and non-deterministic availability of harvested energy and the necessity to plan the energy usage of the energy harvesting sensor nodes ahead of time, on the one hand, have to be balanced with the dynamically changing and complex demand for control updates from the automatic control plant and thus energy usage, on the other hand. We propose a hierarchical approach with which the resources of the energy harvesting sensor nodes are managed on a long time horizon and on a faster timescale, self-triggered model predictive control controls the plant. The controller of the harvesting-based nodes’ resources schedules the future energy usage ahead of time and the self-triggered model predictive control incorporates these time-varying energy constraints. For this novel combination of energy harvesting and automatic control systems, we derive provable properties in terms of correctness, feasibility, and performance. We evaluate the approach on a double integrator and demonstrate its usability and performance in a room temperature and air quality control case study. Naomi Stricker, Yingzhao Lian, Yuning Jiang 0002, Colin N. Jones, Lothar Thiele |
ACM Trans. Cyber Phys. Syst. | 3 |
| 2022 | Over-the-Air Federated Learning via Second-Order OptimizationabstractFederated learning (FL) is a promising learning paradigm that can tackle the increasingly prominent isolated data islands problem while keeping users’ data locally with privacy and security guarantees. However, FL could result in task-oriented data traffic flows over wireless networks with limited radio resources. To design communication-efficient FL, most of the existing studies employ the first-order federated optimization approach that has a slow convergence rate. This however results in excessive communication rounds for local model updates between the edge devices and edge server. To address this issue, in this paper, we instead propose a novel over-the-air second-order federated optimization algorithm to simultaneously reduce the communication rounds and enable low-latency global model aggregation. This is achieved by exploiting the waveform superposition property of a multi-access channel to implement the distributed second-order optimization algorithm over wireless networks. The convergence behavior of the proposed algorithm is further characterized, which reveals a linear-quadratic convergence rate with an accumulative error term in each iteration. We thus propose a system optimization approach to minimize the accumulated error gap by joint device selection and beamforming design. Numerical results demonstrate the system and communication efficiency compared with the state-of-the-art approaches. Peng Yang 0027, Yuning Jiang 0002, Ting Wang 0001, Yong Zhou 0006, Yuanming Shi, Colin N. Jones |
IEEE Trans. Wirel. Commun. | 2 |
| 2021 | Joint Energy Management for Distributed Energy Harvesting SystemsabstractEmploying energy harvesting to power the Internet of Things supports their long-term, self-sustainable, and maintenance-free operation. These energy harvesting systems have an energy management subsystem to orchestrate the flow of energy and optimize their achievable system performance. Numerous such algorithms for a single harvesting-based system have been proposed. When envisioning the joint use of multiple distributed energy harvesting nodes in a single application, the performance and behavior of the distributed system depends on the mutual energy availability and therefore energy management of all nodes. We propose to perform the energy management of multiple distributed energy harvesting nodes jointly and thus, optimize the distributed system's performance as opposed to the performance of each energy harvesting node individually. We demonstrate the novel joint optimization in a scenario with multiple energy harvesting nodes and observe that the distributed system's performance improves by 28 % compared to when each node's energy is managed individually. Naomi Stricker, Yingzhao Lian, Yuning Jiang 0002, Colin N. Jones, Lothar Thiele |
SenSys | 3 |
| 2021 | Over-the-Air Computation via Reconfigurable Intelligent SurfaceabstractOver-the-air computation (AirComp) is a disruptive technique for fast wireless data aggregation in Internet of Things (IoT) networks via exploiting the waveform superposition property of multiple-access channels. However, the performance of AirComp is bottlenecked by the worst channel condition among all links between the IoT devices and the access point. In this paper, a reconfigurable intelligent surface (RIS) assisted AirComp system is proposed to boost the received signal power and thus mitigate the performance bottleneck by reconfiguring the propagation channels. With an objective to minimize the AirComp distortion, we propose a joint design of AirComp transceivers and RIS phase-shifts, which however turns out to be a highly intractable non-convex programming problem. To this end, we develop a novel alternating minimization framework in conjunction with the successive convex approximation technique, which is proved to converge monotonically. To reduce the computational complexity, we transform the subproblem in each alternation as a smooth convex-concave saddle point problem, which is then tackled by proposing a Mirror-Prox method that only involves a sequence of closed-form updates. Simulations show that the computation time of the proposed algorithm can be two orders of magnitude smaller than that of the state-of-the-art algorithms, while achieving a similar distortion performance. Wenzhi Fang, Yuning Jiang 0002, Yuanming Shi, Yong Zhou 0006, Wei Chen 0002, Khaled Ben Letaief |
IEEE Trans. Commun. | 2 |
| 2020 | An Optical Flow Based Multi-Object Tracking Approach Using Sequential Convex ProgrammingabstractObject tracking, as one of the open topics in the past few decades, has been widely applied in the field of video processing. Although there has been intensive research on this topic, some challenges still exist, such as occlusions, clutter and dynamic scenarios. In this paper, we propose a new multi-layer formulation for multi-object tracking, which incorporates the traditional optical flow constraints with a new product term. Then, a sequential convex programming (SCP) based method is presented to solve the resulting non-convex optimization problem. The proposed method can achieve a linear convergence rate, which is demonstrated by our numerical results. In order to illustrate the performance of our approach, it is implemented for composite videos, where two objects move in opposite directions. Different experimental settings based on noise-free, occlusions, clutter and dynamic scenarios show that the method can robustly recover multiple layers and the velocity of the objects. Moreover, the experiments indicate considerable potential for handling more complex scenarios. Qingwen Xu, Zhengpeng He, Yuning Jiang 0002 |
ICARCV | 4 |