Di Wu 0044

dblp:52/328-44 · DBLP profile ↗
← Back
41ranked-venue papers
8as first author
39since 2021 · last 2026
0000-0001-7419-9903ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 28 · 6 first-author · 28 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Decoding RWA Tokenized U.S. Treasuries: Functional Dissection and Address Role Inference
abstract
Tokenized U.S. Treasuries have emerged as a prominent subclass of real-world assets (RWAs), offering cryptographically secured, yield-bearing instruments issued across multi-chain Web3 infrastructures, with growing significance for transparency, accessibility, and financial inclusion. While the market has expanded rapidly, empirical analyses of transaction-level behaviours remain limited. This paper conducts a quantitative, function-level dissection of U.S. Treasury-backed RWA tokens, including BUIDL, BENJI, and USDY across multi-chain: mostly Ethereum and Layer-2s. Decoded contract calls expose core financial primitives such as issuance, redemption, transfer, and bridging, revealing patterns that distinguish institutional participants from smaller or retail users for the extent and limits of inclusivity in current RWA adoption. To infer address-level economic roles, we introduce a curvature-aware representation learning model. Our method outperforms baseline models in role inference on our collected U.S. Treasury transaction dataset and generalizes to address classification across broader public blockchain transaction datasets. The decoded transaction-level patterns in tokenized U.S. Treasuries across chains surface the degree of retail participation, and the role inference model enables the distinction between institutional treasuries, arbitrage bots, and retail traders based on behavioral patterns, facilitating future more transparent, inclusive, and accountable Web3 finance.
Junliang Luo, Katrin Tinn, Samuel Ferreira Duran, Di Wu 0044, Xue (Steve) Liu
ICBC4
2026 STEMS: Spatial-Temporal-Enhanced Safe Multiagent Coordination for Building Energy Management
abstract
Building energy management is essential for achieving carbon reduction goals, improving occupant comfort and reducing energy costs. Coordinated building energy management faces critical challenges in exploiting spatial-temporal dependencies while ensuring operational safety across multi-building systems. Current multi-building energy systems face three key challenges: insufficient spatial-temporal information exploitation, lack of rigorous safety guarantees, and system complexity. This paper proposes Spatial-Temporal Enhanced Safe Multi-Agent Coordination (STEMS), a novel safety-constrained multi-agent reinforcement learning framework for coordinated building energy management. STEMS integrates two core components: (1) a spatial-temporal graph representation learning framework using GCN-Transformer fusion architecture to capture inter-building relationships and temporal patterns, and (2) a safety-constrained multi-agent RL algorithm incorporating Control Barrier Functions to provide mathematical safety guarantees. Extensive experiments on real-world building datasets demonstrate STEMS’s strong performance over existing methods, showing that STEMS achieves 21% cost reduction and 18% emission reduction over traditional methods, 65% safety violation reduction compared with learning-based baselines, and maintains the lowest discomfort rate. The framework also demonstrates strong robustness during extreme weather conditions and maintains effectiveness across different building types.
Huiliang Zhang, Di Wu 0044, Arnaud Zinflou, Benoit Boulet
IEEE Internet Things J.2
2025 Probabilistic Electrical Load Forecasting via Prior-Guided Meta Diffusion Models
abstract
Accurate electric load forecasting is of critical importance for modern power grids. It can help optimize energy management, reduce operational costs, and enhance grid stability. Existing load forecasting tools typically perform well when modelling long-term trends with substantive data upon which to build a model, but can perform poorly for short-term load forecasting or when data is sparse or incomplete. Diffusion models have recently emerged as powerful generative tools that excel in modelling complex distributions, making them a promising approach for electric load forecasting. In this paper, we build upon recent diffusion models for time series forecasting and explore the potential of combining diffusion models with prior models to improve performance. Additionally, we propose a metric-based meta-learning approach for fast data adaptation. Experimental results with this metric-based meta-learning approach on real-world load forecasting datasets outperform state-of-the-art baselines, showcasing the potential of diffusion-based refinement in practical forecasting applications.
Zhiqi Zhuang, Di Wu 0044, Michael R. M. Jenkin, Ekram Hossain 0001, Arnaud Zinflou, Alexia Marchand, Benoit Boulet
GLOBECOM2
2025 Time Series Forecasting via Reinforcement-Learning-Based Model Combination
abstract
Time series data permeates our daily existence and has been recognized as of significant importance for many sectors, such as energy, transportation, telecommunication, and health care. Ensemble learning stands as a prevalent technique in time series forecasting, adept at handling dynamic data distributions and enhancing robustness. Nevertheless, the determination of model combination weights for computing ensemble outputs remains an unresolved issue. In this endeavour, we introduce a comprehensive reinforcement learning-based approach to tackle this challenge. Moreover, our method remains agnostic to the base models within the ensemble and exhibits versatility across a spectrum of time series forecasting tasks, thereby holding promise for application across diverse real-world scenarios. In this work, we extensively test the performance of our method on real-world datasets. Experimental results show that our proposed method achieves an average 6.2% forecasting accuracy improvement over all other single baseline methods.
Yuwei Fu, Di Wu 0044, Benoit Boulet, Arnaud Zinflou
IEEE Internet Things J.2
2024 Surrogate Data Source Transfer (SDST): An Efficient Transfer Learning Approach for Time Series Forecasting
abstract
Time series prediction plays a crucial role in optimizing the operation of communication networks. Applications of time series prediction include traffic prediction, channel state prediction, handover prediction, etc. However, training high-quality models for these tasks requires large volumes of historical data. This requirement may not be available in some scenarios. In this case, instance-based Transfer Learning (TL) comes as a prominent solution for this problem. However, a few concerns could be raised such as: 1) the time and bandwidth resources consumed in the transfer, 2) it will be hard to specify the amount of data to be transferred, and 3) in case of transferring a subset of the data, which subset is better to transfer. To address these challenges, we propose a novel approach for TL, which is similar to, but different than, instance-based TL based on generative models. We coined the new approach as Surrogate Data Source Transfer (SDST), in which a generative model is trained on the source task. We then transfer the model to the target task (with limited historical data). Extensive experiments confirm the superior performance of the proposed approach in terms of prediction accuracy and consumed resources (time and bandwidth). Our TL approach reduced the mean absolute percentage error (MAPE) by a margin that hits 81% in some datasets. For the source code and data, we refer to the repository https://github.com/MoeR3za/Korsahy_TGAN.
Mostafa Hussien, Mohamed Shoaib, Di Wu 0044, Kim Khoa Nguyen, Mohamed Cheriet
ICC3
2024 Accelerating Digital Twin Calibration with Warm-Start Bayesian Optimization
abstract
Digital twins are expected to play an important role in the widespread adaptation of AI-based networking solutions in the real world. The calibration of these virtual replicas is critical to ensure a trustworthy replication of the real environment. This work focuses on the input parameter calibration of radio access network (RAN) simulators using real network performance metrics as supervision signals. Usually, the RAN digital twin is considered a black-box function and each calibration problem is viewed as a standalone search problem. RAN simulators are slow and non-differentiable, often posing as the bottleneck in the execution time for these search problems. In this work, we aim to accelerate the search process by reducing the number of interactions with the simulator by leveraging RAN interactions from previous problems. We present a sequential Bayesian optimization framework that uses information from the past to warm-start the calibration process. Assuming that the network performance exhibits gradual and periodic changes, the stored information can be reused in future calibrations. We test our method across multiple physical sites over one week and show that using the proposed framework, we can obtain better calibration with a smaller number of interactions with the simulator during the search phase.
Abhisek Konar, Amal Feriani, Di Wu 0044, Seowoo Jang, Xue Liu 0004, Gregory Dudek
ICC3
2024 Optimizing Energy Saving for Wireless Networks Via Offline Decision Transformer
abstract
With the global aim of reducing carbon emissions, energy saving for communication systems has gained tremendous attention. Efficient energy-saving solutions are not only required to accommodate the fast growth in communication demand but solutions are also challenged by the complex nature of the load dynamics. Recent reinforcement learning (RL)-based methods have shown promising performance for network optimization problems, such as base station energy saving. However, a major limitation of these methods is the requirement of online exploration of potential solutions using a high-fidelity simulator or the need to perform exploration in a real-world environment. We circumvent this issue by proposing an offline reinforcement learning energy saving (ORES) framework that allows us to learn an efficient control policy using previously collected data. We first deploy a behavior energy-saving policy on base stations and generate a set of interaction experiences. Then, using a robust deep offline reinforcement learning algorithm, we learn an energy-saving control policy based on the collected experiences. Results from experiments conducted on a diverse collection of communication scenarios with different behavior policies showcase the effectiveness of the proposed energy-saving algorithms.
Yi Tian Xu, Di Wu 0044, Michael R. M. Jenkin, Seowoo Jang, Xue Liu 0004, Gregory Dudek
ICC2
2024 FuRL: Visual-Language Models as Fuzzy Rewards for Reinforcement Learning
abstract
In this work, we investigate how to leverage pre-trained visual-language models (VLM) for online Reinforcement Learning (RL). In particular, we focus on sparse reward tasks with pre-defined textual task descriptions. We first identify the problem of reward misalignment when applying VLM as a reward in RL tasks. To address this issue, we introduce a lightweight fine-tuning method, named Fuzzy VLM reward-aided RL (FuRL), based on reward alignment and relay RL. Specifically, we enhance the performance of SAC/DrQ baseline agents on sparse reward tasks by fine-tuning VLM representations and using relay RL to avoid local minima. Extensive experiments on the Meta-world benchmark tasks demonstrate the efficacy of the proposed method. Code is available at: https://github.com/fuyw/FuRL.
Yuwei Fu, Di Wu 0044, Benoit Boulet
ICML3
2024 Towards Enhanced Fairness and Sample Efficiency in Traffic Signal Control
abstract
Traffic signal control (TSC) has seen substantial advancements through the application of reinforcement learning (RL) algorithms, which have shown remarkable potential in enhancing traffic flow efficiency. These RL-based approaches often surpass traditional rule-based methods, particularly in dynamic traffic environments. However, current RL solutions for TSC predominantly rely on model-free methods, necessitating extensive environmental interactions during training. This requirement can be prohibitively expensive or unfeasible in real-world implementations. Furthermore, existing methods have frequently neglected the issue of fairness in multi-intersection control, resulting in unbalanced congestion across different intersections. To address these challenges, we present FM2Light, a fairness-aware model-based multi-agent RL framework for TSC. Our approach leverages an ensemble of global world models for generating synthetic samples to enhance sample efficiency, thereby mitigating the data-intensive nature of the training process. Additionally, FM2Light incorporates a refined reward structure to promote fairness and improve coordination across multiple intersections. Extensive evaluations conducted in diverse real-world scenarios demonstrate that FM2Light achieves performance comparable to or exceeding that of model-free RL (MFRL) methods, while significantly reducing sample requirements and ensuring more equitable control among multiple agents.
Xingshuai Huang, Di Wu 0044, Michael R. M. Jenkin, Benoit Boulet
IROS2
2024 Robot Policy Learning with Temporal Optimal Transport Reward
abstract
Reward specification is one of the most tricky problems in Reinforcement Learning, which usually requires tedious hand engineering in practice. One promising approach to tackle this challenge is to adopt existing expert video demonstrations for policy learning. Some recent work investigates how to learn robot policies from only a single/few expert video demonstrations. For example, reward labeling via Optimal Transport (OT) has been shown to be an effective strategy to generate a proxy reward by measuring the alignment between the robot trajectory and the expert demonstrations. However, previous work mostly overlooks that the OT reward is invariant to temporal order information, which could bring extra noise to the reward signal. To address this issue, in this paper, we introduce the Temporal Optimal Transport (TemporalOT) reward to incorporate temporal order information for learning a more accurate OT-based proxy reward. Extensive experiments on the Meta-world benchmark tasks validate the efficacy of the proposed method. Our code is available at: https://github.com/fuyw/TemporalOT.
Yuwei Fu, Di Wu 0044, Benoit Boulet
NeurIPS3
2023 AdaTeacher: Adaptive Multi-Teacher Weighting for Communication Load Forecasting
abstract
To deal with notorious delays in communication systems, it is crucial to forecast key system characteristics, such as the communication load. Most existing studies aggregate data from multiple edge nodes for improving the forecasting accuracy. However, the bandwidth cost of such data aggregation could be unacceptably high from the perspective of system operators. To achieve both the high forecasting accuracy and bandwidth efficiency, this paper proposes an Adaptive Multi-Teacher Weighting in Teacher-Student Learning approach, namely AdaTeacher, for communication load forecasting of multiple edge nodes. Each edge node trains a local model on its own data. A target node collects multiple models from its neighbor nodes and treats these models as teachers. Then, the target node trains a student model from teachers via Teacher-Student (T-S) learning. Unlike most existing T-S learning approaches that treat teachers evenly, resulting in a limited performance, AdaTeacher introduces a bilevel optimization algorithm to dynamically learn an importance weight for each teacher toward a more effective and accurate T-S learning process. Compared to the state-of-the-art methods, Ada Teacher not only reduces the bandwidth cost by 53.85%, but also improves the load forecasting accuracy by 21.56% and 24.24% on two real-world datasets.
Chengming Hu, Ju Wang 0003, Di Wu 0044, Jianzhong Zhang 0002, Xue Liu 0004, Gregory Dudek
GLOBECOM3
2023 Energy Saving in Cellular Wireless Networks via Transfer Deep Reinforcement Learning
abstract
With the increasing use of data-intensive mobile applications and the number of mobile users, the demand for wireless data services has been increasing exponentially in recent years. In order to address this demand, a large number of new cellular base stations are being deployed around the world, leading to a significant increase in energy consumption and greenhouse gas emission. Consequently, energy consumption has emerged as a key concern in the fifth-generation (5G) network era and beyond. Reinforcement learning (RL), which aims to learn a control policy via interacting with the environment, has been shown to be effective in addressing network optimization problems. However, for reinforcement learning, especially deep reinforcement learning, a large number of interactions with the environment are required. This often limits its applicability in the real world. In this work, to better deal with dynamic traffic scenarios and improve real-world applicability, we propose a transfer deep reinforcement learning framework for energy optimization in cellular communication networks. Specifically, we first pre-train a set of RL-based energy-saving policies on source base stations and then transfer the most suitable policy to the given target base station in an unsupervised learning manner. Experimental results demonstrate that base station energy consumption can be reduced significantly using this approach.
Di Wu 0044, Yi Tian Xu, Michael R. M. Jenkin, Seowoo Jang, Ekram Hossain 0001, Xue Liu 0004, Gregory Dudek
GLOBECOM1
2023 Learning to Adapt: Communication Load Balancing via Adaptive Deep Reinforcement Learning
abstract
The association of mobile devices with network resources (e.g., base stations, frequency bands/channels), known as load balancing, is critical to reduce communication traffic congestion and network performance. Reinforcement learning (RL) has shown to be effective for communication load balancing and achieves better performance than currently used rule-based methods, especially when the traffic load changes quickly. However, RL-based methods usually need to interact with the environment for a large number of time steps to learn an effective policy and can be difficult to tune. In this work, we aim to improve the data efficiency of RL-based solutions to make them more suitable and applicable for real-world applications. Specifically, we propose a simple, yet efficient and effective deep RL-based wireless network load balancing framework. In this solution, a set of good initialization values for control actions are selected with some cost-efficient approach to center the training of the RL agent. Then, a deep RL-based agent is trained to find offsets from the initialization values that optimize the load balancing problem. Experimental evaluation on a set of dynamic traffic scenarios demonstrates the effectiveness and efficiency of the proposed method.
Di Wu 0044, Yi Tian Xu, Jimmy Li 0001, Michael R. M. Jenkin, Ekram Hossain 0001, Seowoo Jang, Jianzhong Zhang 0002, Xue Liu 0004, Gregory Dudek
GLOBECOM1
2023 Multi-Agent Attention Actor-Critic Algorithm for Load Balancing in Cellular Networks
abstract
In cellular networks, User Equipment (UE) handoff from one Base Station (BS) to another, giving rise to the load balancing problem among the BSs. To address this problem, BSs can work collaboratively to deliver a smooth migration (or handoff) and satisfy the UEs' service requirements. This paper formulates the load balancing problem as a Markov game and proposes a Robust Multi-agent Attention Actor-Critic (Robust-MA3C) algorithm that can facilitate collaboration among the BSs (i.e., agents). In particular, to solve the Markov game and find a Nash equilibrium policy, we embrace the idea of adopting a nature agent to model the system uncertainty. Moreover, we utilize the self-attention mechanism, which encourages high-performance BSs to assist low-performance BSs. In addition, we consider two types of schemes, which can facilitate load balancing for both active UEs and idle UEs. We carry out extensive evaluations by simulations, and simulation results illustrate that, compared to the state-of-the-art MARL methods, Robust-MA3C scheme can improve the overall performance by up to 45%.
Jikun Kang, Di Wu 0044, Ju Wang 0003, Ekram Hossain 0001, Xue Liu 0004, Gregory Dudek
ICC2
2023 Communication Load Balancing via Efficient Inverse Reinforcement Learning
abstract
Communication load balancing aims to balance the load between different available resources, and thus improve the quality of service for network systems. After formulating the load balancing (LB) as a Markov decision process problem, reinforcement learning (RL) has recently proven effective in addressing the LB problem. To leverage the benefits of classical RL for load balancing, however, we need an explicit reward definition. Engineering this reward function is challenging, because it involves the need for expert knowledge and there lacks a general consensus on the form of an optimal reward function. In this work, we tackle the communication load balancing problem from an inverse reinforcement learning (IRL) approach. To the best of our knowledge, this is the first time IRL has been successfully applied in the field of communication load balancing. Specifically, first, we infer a reward function from a set of demonstrations, and then learn a reinforcement learning load balancing policy with the inferred reward function. Compared to classical RL-based solution, the proposed solution can be more general and more suitable for real-world scenarios. Experimental evaluations implemented on different simulated traffic scenarios have shown our method to be effective and better than other baselines by a considerable margin.
Abhisek Konar, Di Wu 0044, Yi Tian Xu, Seowoo Jang, Steve Liu, Gregory Dudek
ICC2
2023 Self-Supervised Transformer Architecture for Change Detection in Radio Access Networks
abstract
Radio Access Networks (RANs) for telecommunications represent large agglomerations of interconnected hardware consisting of hundreds of thousands of transmitting devices (cells). Such networks undergo frequent and often heterogeneous changes caused by network operators, who are seeking to tune their system parameters for optimal performance. The effects of such changes are challenging to predict and will become even more so with the adoption of fifth-generation/sixth-generation (5G/6G) networks. Therefore, RAN monitoring is vital for network operators. We propose a self-supervised learning framework that leverages self-attention and self-distillation for this task. It works by detecting changes in Performance Measurement data, a collection of time-varying metrics which reflect a set of diverse measurements of the network performance at the cell level. Experimental results show that our approach outperforms the state of the art by 4% on a real-world based dataset consisting of about hundred thousands time series. It also has the merits of being scalable and generalizable. This allows it to provide deep insight into the specifics of mode of operation changes while relying minimally on expert knowledge.
Igor Kozlov, Dmitriy Rivkin, Wei-Di Chang, Di Wu 0044, Xue Liu 0004, Gregory Dudek
ICC4
2023 Policy Reuse for Communication Load Balancing in Unseen Traffic Scenarios
abstract
With the continuous growth in communication network complexity and traffic volume, communication load balancing solutions are receiving increasing attention. Specifically, reinforcement learning (RL)-based methods have shown impressive performance compared with traditional rule-based methods. However, standard RL methods generally require an enormous amount of data to train, and generalize poorly to scenarios that are not encountered during training. We propose a policy reuse framework in which a policy selector chooses the most suitable pre-trained RL policy to execute based on the current traffic condition. Our method hinges on a policy bank composed of policies trained on a diverse set of traffic scenarios. When deploying to an unknown traffic scenario, we select a policy from the policy bank based on the similarity between the previous-day traffic of the current scenario and the traffic observed during training. Experiments demonstrate that this framework can outperform classical and adaptive rule-based methods by a large margin.
Jimmy Li 0001, Di Wu 0044, Michael R. M. Jenkin, Seowoo Jang, Xue Liu 0004, Gregory Dudek
ICC3
2023 Mixed-Variable PSO with Fairness on Multi-Objective Field Data Replication in Wireless Networks
abstract
Digital twins have shown a great potential in supporting the development of wireless networks. They are virtual representations of 5G/6G systems enabling the design of machine learning and optimization-based techniques. Field data replication is one of the critical aspects of building a simulation-based twin, where the objective is to calibrate the simulation to match field performance measurements. Since wireless networks involve a variety of key performance indicators (KPIs), the replication process becomes a multi-objective optimization problem in which the purpose is to minimize the error between the simulated and field data KPIs. Unlike previous works, we focus on designing a data-driven search method to calibrate the simulator and achieve accurate and reliable reproduction of field performance. This work proposes a search-based algorithm based on mixed-variable particle swarm optimization (PSO) to find the optimal simulation parameters. Furthermore, we extend this solution to account for potential conflicts between the KPIs using a-fairness concept to adjust the importance attributed to each KPI during the search. Experiments on field data showcase the effectiveness of our approach to (i) improve the accuracy of the replication, (ii) enhance the fairness between the different KPIs, and (iii) guarantee faster convergence compared to other methods.
Dun Yuan, Yujin Nam, Amal Feriani, Abhisek Konar, Di Wu 0044, Seowoo Jang, Xue Liu 0004, Gregory Dudek
ICC5
2023 Gap Minimization for Knowledge Sharing and Transfer
abstract
Learning from multiple related tasks by knowledge sharing and transfer has become increasingly relevant over the last two decades. In order to successfully transfer information from one task to another, it is critical to understand the similarities and differences between the domains. In this paper, we introduce the notion of performance gap, an intuitive and novel measure of the distance between learning tasks. Unlike existing measures which are used as tools to bound the difference of expected risks between tasks (e.g., $\mathcal{H}$-divergence or discrepancy distance), we theoretically show that the performance gap can be viewed as a data- and algorithm-dependent regularizer, which controls the model complexity and leads to finer guarantees. More importantly, it also provides new insights and motivates a novel principle for designing strategies for knowledge sharing and transfer: gap minimization. We instantiate this principle with two algorithms: 1. gapBoost, a novel and principled boosting algorithm that explicitly minimizes the performance gap between source and target domains for transfer learning; and 2. gapMTNN, a representation learning algorithm that reformulates gap minimization as semantic conditional matching for multitask learning. Our extensive evaluation on both transfer learning and multitask learning benchmark data sets shows that our methods outperform existing baselines.
Boyu Wang 0004, Jorge A. Mendez, Changjian Shui, Fan Zhou 0006, Di Wu 0044, Gezheng Xu, Christian Gagné 0001, Eric Eaton
J. Mach. Learn. Res.5
2023 MetaProbformer for Charging Load Probabilistic Forecasting of Electric Vehicle Charging Stations
abstract
The penetration of electric vehicles (EV) has been increasing rapidly in recent years. Electric vehicle charging load poses a huge demand on the power grids. The forecasting for electric vehicle charging load, especially for the charging load of EV charging stations, is of significant importance for the safe operation of power grids. However, most of the existing forecasting methods fail to capture the long-term dependencies efficiently and assume the availability of a large amount of training data. Hence, they cannot address newly built charging stations with scarce historical charging load data. Meanwhile, most of the methods focus on point forecasting, which lacks risk consideration. In this work, we aim to leverage the benefits of Transformer-based models for EV charging forecasting. Specifically, we propos Probformer, a Transformer-based forecasting model for charging load forecasting. To enable Probformer to adapt fast to unseen environments, we further extend it to MetaProbformer, a meta-learning-based forecasting framework. Extensive experiments have been done on real-world datasets for both point forecasting and probabilistic forecasting. Experimental results show that our methods can consistently outperform baseline methods by a large margin.
Xingshuai Huang, Di Wu 0044, Benoit Boulet
IEEE Trans. Intell. Transp. Syst.2
2023 On the Benefits of Two Dimensional Metric Learning
abstract
In this paper, we study two dimensional metric learning (2DML) for matrix data from both theoretical and algorithmic perspectives. We first investigate the generalization bounds of 2DML based on the notion of Rademacher complexity, which theoretically justifies the benefits of learning from matrices directly. Furthermore, we present a novel boosting-based algorithm that scales well with the feature dimension. Finally, we introduce an efficient rank-one correction algorithm, which is tailored to our boosting learning procedure to produce a low-rank solution to 2DML. As our algorithm works directly on the data in matrix representation, it scales well with the feature dimension, keeps the structure and dependence in the data, and has a more compact structure and much fewer parameters to optimize. Extensive evaluations on several benchmark data sets also empirically verify the effectiveness and efficiency of our algorithm.
Di Wu 0044, Fan Zhou 0006, Boyu Wang 0004, Qicheng Lao, Chiman Wong, Changjian Shui, Yuan Zhou 0006, Feng Wan 0003
IEEE Trans. Knowl. Data Eng.1
2022 Reinforcement Learning Based Dynamic Model Combination for Time Series Forecasting
abstract
Time series data appears in many real-world fields such as energy, transportation, communication systems. Accurate modelling and forecasting of time series data can be of significant importance to improve the efficiency of these systems. Extensive research efforts have been taken for time series problems. Different types of approaches, including both statistical-based methods and machine learning-based methods, have been investigated. Among these methods, ensemble learning has shown to be effective and robust. However, it is still an open question that how we should determine weights for base models in the ensemble. Sub-optimal weights may prevent the final model from reaching its full potential. To deal with this challenge, we propose a reinforcement learning (RL) based model combination (RLMC) framework for determining model weights in an ensemble for time series forecasting tasks. By formulating model selection as a sequential decision-making problem, RLMC learns a deterministic policy to output dynamic model weights for non-stationary time series data. RLMC further leverages deep learning to learn hidden features from raw time series data to adapt fast to the changing data distribution. Extensive experiments on multiple real-world datasets have been implemented to showcase the effectiveness of the proposed method.
Yuwei Fu, Di Wu 0044, Benoit Boulet
AAAI2
2022 Accurate Communication Traffic Forecasting with Multi-Source Adaptive Feature Boosting
abstract
Advanced communication network functions, such as resource allocation and dynamic spectrum management, heavily rely on the accurate forecasting of traffic. Data-driven solutions, e.g., Neural Network (NN) based forecasting methods, have been proven to be effective only when sufficient data is available. However, Base Stations (BSs) have limited data in the real world, since big data for communication networks could be extremely expensive to collect, store, and migrate. Therefore, most existing traffic forecasting methods have limited accuracy in reality due to the lack of big data. To tackle this problem, our key observation is that, despite the data “amount” in a BS is limited, the data “source” is rich and diverse, i.e., in addition to Internet traffic logs, there are logs of Call and SMS. More importantly, our analysis shows a high correlation between different sources, which can be utilized to improve the forecasting accuracy. Motivated by this, we introduce AdaSource, a Multi-Source Adaptive Feature Boosting approach, which utilizes data source correlations for accurate traffic forecasting even on data-limited BSs. The core idea of AdaSource is a novel two-branch NN structure that adaptively trains multiple Encoder-Decoders for refining different data sources and multiple Encoder-Predictors for utilizing data source correlations to improve the accuracy. The experiments on a real-world dataset show that AdaSource improves the forecasting accuracy by up to 30.14%, compared to the state-of-the-art methods.
Chengming Hu, Ju Wang 0003, Di Wu 0044, Xue Liu 0004, Gregory Dudek
GLOBECOM3
2022 Efficient Neural Data Compression for Machine Type Communications via Knowledge Distillation
abstract
The anticipated huge number of devices and large traffic volumes impose new challenges on the communication system requirements and design. One of the main requirements of massive machine-type communication (mMTC) is to support network energy efficiency. Data compression is a widely adopted technique that enables higher energy efficiency, lower latency, and better bandwidth utilization. Unfortunately, the current compression techniques are mainly designed for human-type communications (HTC). Therefore, they consider the reconstruction fidelity, rather than the accuracy of inferred decisions, as the sole performance metric. In this work, we propose a novel encoder for data compression in mMTC communications, which is termed Distillation Encoder (DE). Unlike prior work, the design of the proposed DE aims to achieve high compression ratios while preserving the accuracy of the inferred decisions. DE inherits the knowledge of a large teacher model (trained on the raw data) through knowledge distillation. Evaluating the proposed framework on several public datasets shows a clear performance advantage compared with baseline models in terms of the inferred decision accuracy and generalizing to yet-unseen data. Moreover, the DE can be applied to learn efficient quantizers, as shown in the results.
Mostafa Hussien, Yi Tian Xu, Di Wu 0044, Xue Liu 0004, Gregory Dudek
GLOBECOM3
2022 Attentive Knowledge Transfer for Short-term Load Forecasting
abstract
The modern power system is transitioning towards increasing penetration of renewable energy generation and demand from different types of electrical appliances. With this transition, residential load forecasting, especially short-term load forecasting (STLF), is becoming more and more challenging and important. Accurate short-term load forecasting can help improve energy dispatching efficiency and, as a consequence, reduce overall power system operation cost. Most current load forecasting algorithms assume that there is a large amount of training data available upon which to learn a reliable load forecasting model. However, this assumption can be challenging for real-world applications. In this work, we first propose the use of transfer learning and an attention mechanism to improve short-term load forecasting for a target domain with only a limited amount of available data. Furthermore, we extend the proposed method to utilize heterogeneous features which enables the approach to deal with more complex scenarios in the real world. Experimental results using real-world data sets show that the proposed methods can improve forecasting accuracy by a large margin over several existing baselines.
Di Wu 0044, Michael R. M. Jenkin, Yi Tian Xu, Xue Liu 0004, Gregory Dudek
GLOBECOM1
2022 Traffic Scenario Clustering and Load Balancing with Distilled Reinforcement Learning Policies
abstract
Due to the rapid increase in wireless communication traffic in recent years, load balancing is becoming increasingly important for ensuring the quality of service. However, variations in traffic patterns near different serving base stations make this task challenging. On one hand, crafting a single control policy that performs well across all base station sectors is often difficult. On the other hand, maintaining separate controllers for every sector introduces overhead, and leads to redundancy if some of the sectors experience similar traffic patterns. In this paper, we propose to construct a concise set of controllers that cover a wide range of traffic scenarios, allowing the operator to select a suitable controller for each sector based on local traffic conditions. To construct these controllers, we present a method that clusters similar scenarios and learns a general control policy for each cluster. We use deep reinforcement learning (RL) to first train separate control policies on diverse traffic scenarios, and then incrementally merge together similar RL policies via knowledge distillation. Experimental results show that our concise policy set reduces redundancy with very minor performance degradation compared to policies trained separately on each traffic scenario. Our method also outperforms handcrafted control parameters, joint learning on all tasks, and two popular clustering methods.
Jimmy Li 0001, Di Wu 0044, Yi Tian Xu, Tianyu Li 0008, Seowoo Jang, Xue Liu 0004, Gregory Dudek
ICC2
2022 Coordinated Load Balancing in Mobile Edge Computing Network: a Multi-Agent DRL Approach
abstract
Mobile edge computing (MEC) networks have been recently adopted to accommodate the fast-growing number of mobile devices performing complicated tasks with limited hardware capability. Recently, edge nodes with communication, computation, and caching capacities are starting to be deployed in MEC networks. Due to the physical separation of these resources, efficient coordination and scheduling are important for efficient resource utilization and optimal network performance. In this paper, we study mobility load balancing for communication, computation, and caching-enabled heterogeneous MEC networks. Specifically, we propose to tackle this problem via a multi-agent deep reinforcement learning-based framework. Users served by overloaded edge nodes are handed over to less loaded ones, to minimize the load in the most loaded base station in the network. In this framework, the handover decision for each user is made based on the user’s own observation which comprises the user’s task at hand and the load status of the MEC network. Simulation results show that our proposed multi-agent deep reinforcement learning-based approach can reduce the time-average maximum load by up to 30% and the end-to-end delay by 50% compared to baseline algorithms.
Manyou Ma, Di Wu 0044, Yi Tian Xu, Jimmy Li 0001, Seowoo Jang, Xue Liu 0004, Gregory Dudek
ICC2
2022 Active Deep Multi-task Learning for Forecasting Short-Term Loads
abstract
With the increasing adoption of renewable energy generation and electric devices, electric load forecasting, especially short-term load forecasting (STLF), is becoming more and more important. The widespread adoption of smart meters makes it possible to utilize complex machine learning models for both aggregated load and single-home residential load forecasting. Similar homes in nearby locations are likely to have similar load consumption patterns and this similarity can be used to improve the overall forecasting performance. However, most current work on load forecasting focuses on single learning task without exploiting the benefit of joint learning. In this paper, we propose the use of the multi-task learning (MTL) framework with long short-term memory (LSTM) recurrent neural networks for both aggregated and single home STLF. We propose a MTL-based forecasting algorithm for aggregated load forecasting in which single home forecasting is formulated as a single learning task within the MTL framework. This algorithm is extended for single home load forecasting in which load forecasting for a particular home becomes the primary learning task. Experimental results on real-world data sets demonstrate that residential load forecasting for both aggregated load and a single home can be improved within the MTL framework.
Di Wu 0044, Michael R. M. Jenkin, Xue Liu 0004, Gregory Dudek
ICC1
2022 Short-term Load Forecasting with Deep Boosting Transfer Regression
abstract
With the increasing popularity of electric vehicles and the growing trend of working from home, electricity consumption in the residential sector is expected to continue to grow rapidly over the next few years. As a consequence, short-term residential load forecasting is becoming even more vital for the reliability and sustainability of the smart grid. Although deep learning models have shown impressive success in different areas including short-term electric load forecasting, such models require a large amount of training data. For many real-world load forecasting cases, we may not have enough training data to learn a reliable forecasting model. In this paper, we address this challenge through the use of boosting-based transfer learning with multiple sources. We first train a set of deep regression models on source houses that can provide relatively abundant data. We then transfer these learned models via the boosting framework to support data-scarce target houses. The transfer process is selective and customized for each target house to minimize the potential for negative transfer. Experimental results, based on real-world residential data sets, show that the proposed method can significantly improve forecasting accuracy.
Di Wu 0044, Yi Tian Xu, Michael R. M. Jenkin, Ju Wang 0003, Xue Liu 0004, Gregory Dudek
ICC1
2022 A Closer Look at Offline RL Agents
abstract
Despite recent advances in the field of Offline Reinforcement Learning (RL), less attention has been paid to understanding the behaviors of learned RL agents. As a result, there remain some gaps in our understandings, i.e., why is one offline RL agent more performant than another? In this work, we first introduce a set of experiments to evaluate offline RL agents, focusing on three fundamental aspects: representations, value functions and policies. Counterintuitively, we show that a more performant offline RL agent can learn relatively low-quality representations and inaccurate value functions. Furthermore, we showcase that the proposed experiment setups can be effectively used to diagnose the bottleneck of offline RL agents. Inspired by the evaluation results, a novel offline RL algorithm is proposed by a simple modification of IQL and achieves SOTA performance. Finally, we investigate when a learned dynamics model is helpful to model-free offline RL agents, and introduce an uncertainty-based sample selection method to mitigate the problem of model noises. Code is available at: https://github.com/fuyw/RIQL.
Yuwei Fu, Di Wu 0044, Benoit Boulet
NeurIPS2
2022 Fidora: Robust WiFi-Based Indoor Localization via Unsupervised Domain Adaptation
abstract
Emerging Internet of Things (IoT) applications, such as cashier-less shopping, mobile ads targeting, and geo-based augmented reality (AR), are expected to bring us much more convenience and infotainment. To realize this amazing future, we need to feed these applications with user locations of (sub)meter-level resolution anytime and anywhere. Unfortunately, many widely used location sources are either unavailable indoor (e.g., global positioning system) or coarse grained (e.g., user check-ins). In order to provide ubiquitous localization services, the widespread WiFi signals are being leveraged to establish (sub)meter-level localization systems. Fine-grained WiFi propagation characteristics, which are sensitive to human body locations, have been employed to create location fingerprints. However, these WiFi characteristics are also sensitive to: 1) the body shapes of different users and 2) the objects in the background environment. Consequently, systems based on WiFi fingerprints are vulnerable in the presence of: 1) new users with different body shapes and 2) daily changes of the environment, e.g., opening/closing doors. To tackle this issue, this article proposes a WiFi-based localization system based on domain-adaptation with cluster assumption, named Fidora. Fidora is able to: 1) localize different users with labeled data from only one or two example users and 2) localize the same user in a changed environment without labeling any new data. To achieve these, Fidora integrates two major modules. It first adopts a data augmenter that introduces data diversity using a variational autoencoder (VAE). It then trains a domain-adaptive classifier that adjusts itself to newly collected unlabeled data using a joint classification-reconstruction structure. We conducted real-world experiments to evaluate Fidora against the state of the art. It is demonstrated that when tested on an unlabeled user, Fidora increases the average$F1$score by 17.8% and improves the worst case accuracy by 20.2%. Moreover, when applied in a varied environment, Fidora outperforms the state of the art by 23.1%.
Xi Chen 0009, Chenyi Zhou, Xue Liu 0004, Di Wu 0044, Gregory Dudek
IEEE Internet Things J.5
2022 Multiobjective Load Balancing for Multiband Downlink Cellular Networks: A Meta- Reinforcement Learning Approach
abstract
Load balancing has become a key technique to handle the increasing traffic demand and improve the user experience. It evenly distributes the traffic across network resources by offloading users from overloaded base stations or channels to less crowded ones. Load balancing is a multi-objective optimization problem involving the automatic adjustment of several parameters to simultaneously maximize multiple network performance indicators. However, the existing methods mostly rely on single-objective approaches which lead to sub-optimal solutions. In this paper, we introduce the first multi-objective reinforcement learning (MORL) framework for load balancing. Specifically, we propose a solution based on meta-reinforcement learning (meta-RL) to learn a general policy capable of quickly adapting to new trade-offs between the objectives. We further enhance the generalization of our proposed solution using policy distillation techniques. To showcase the effectiveness of our framework, experiments are conducted based on real-world traffic scenarios. Our results show that our load balancing framework can (i) significantly outperform the existing rule-based and single-objective solutions, (ii) compute better Pareto front approximations compared to MORL baselines, and (iii) quickly adapt to new objective trade-offs.
Amal Feriani, Di Wu 0044, Yi Tian Xu, Jimmy Li 0001, Seowoo Jang, Ekram Hossain 0001, Xue Liu 0004, Gregory Dudek
IEEE J. Sel. Areas Commun.2
2021 One for All: Traffic Prediction at Heterogeneous 5G Edge with Data-Efficient Transfer Learning
abstract
By placing the computing, storage and networking resources close to the end users, distributed edge computing greatly benefits the performance of 5G communication systems. However, as a tradeoff, resources on the edge are usually limited and imbalanced among the heterogeneous edge nodes. To overcome this drawback, this paper proposes a Transfer Learning based Prediction (TLP) framework that allows the edge nodes to share their resources and data in an efficient manner. In particular, the TLP framework focuses on the prediction of the future traffic load, which is a key reference for many automated network functions. To enhance the efficiency of data and bandwidth, TLP first learns a base model on a data-abundant edge node (the source), and then transfers this model (instead of data) to other data-limited nodes (the targets). To achieve a delicate balance between maintaining common features and learning target-specific features, we develop a new transfer learning technique named Similarity-based Elastic Weight Con-solidation (SEWC), and integrate it into TLP. Experiments on real-world data illustrate that, compared to the state-of-the-art methods, TLP-SEWC reduces the Mean Absolute Error (MAE) of traffic prediction by up to 57.9%.
Xi Chen 0009, Ju Wang 0003, Yi Tian Xu, Di Wu 0044, Xue Liu 0004, Gregory Dudek, Taeseop Lee, Intaik Park
GLOBECOM5
2021 AFB: Improving Communication Load Forecasting Accuracy with Adaptive Feature Boosting
abstract
Prediction of key system characteristics, such as the communication load, is required to overcome the delays in wireless communication systems. State-of-The-Art (SOTA) approaches mostly apply existing Neural Network (NN) structures, and extract latent features purely based on their sensitivity to the forecasting accuracy. This way of feature extraction may neglect some non-obvious yet informative dimensions in the model input, leading to inaccurate forecasting results. In this paper, we present an Adaptive Feature Boosting (AFB) approach, which integrates multiple AutoEncoders (AEs) to automatically extract robust and comprehensive latent features for communication load forecasting. The recurrent and residual connections among the AEs make sure that the extracted latent features are representative for all input dimensions. With more comprehensive information extracted from the history, the forecasting accuracy is thus improved. We evaluate AFB against existing approaches on a real-world dataset that contains Call Detail Records (CDRs) of the Milan city over a period of two months. The evaluation shows that our AFB-based approach achieves 35.2% more accurate load forecasting results than the SOTA deep approaches.
Chengming Hu, Xi Chen 0009, Ju Wang 0003, Jikun Kang, Yi Tian Xu, Xue Liu 0004, Di Wu 0044, Seowoo Jang, Intaik Park, Gregory Dudek
GLOBECOM8
2021 Learning Assisted Identification of Scenarios Where Network Optimization Algorithms Under-Perform
abstract
We present a generative adversarial method that uses deep learning to identify network load traffic conditions in which network optimization algorithms under-perform other known algorithms: the Deep Convolutional Failure Generator (DCFG). The spatial distribution of network load presents challenges for network operators for tasks such as load balancing, in which a network optimizer attempts to maintain high quality communication while at the same time abiding capacity constraints. Testing a network optimizer for all possible load distributions is challenging if not impossible. We propose a novel method that searches for load situations where a target network optimization method underperforms baseline, which are key test cases that can be used for future refinement and performance optimization. By modeling a realistic network simulator's quality assessments with a deep network and, in parallel, optimizing a load generation network, our method efficiently searches the high dimensional space of load patterns and reliably finds cases in which a target network optimization method under-performs a baseline by a significant margin.
Dmitriy Rivkin, David Meger, Di Wu 0044, Xi Chen 0009, Xue Liu 0004, Gregory Dudek
GLOBECOM3
2021 Load Balancing for Communication Networks via Data-Efficient Deep Reinforcement Learning
abstract
Within a cellular network, load balancing between different cells is of critical importance to network performance and quality of service. Most existing load balancing algorithms are manually designed and tuned rule-based methods where near-optimality is almost impossible to achieve. These rule-based meth-ods are difficult to adapt quickly to traffic changes in real-world environments. Given the success of Reinforcement Learning (RL) algorithms in many application domains, there have been a number of efforts to tackle load balancing for communication systems using RL-based methods. To our knowledge, none of these efforts have addressed the need for data efficiency within the RL framework, which is one of the main obstacles in applying RL to wireless network load balancing. In this paper, we formulate the communication load balancing problem as a Markov Decision Process and propose a data-efficient transfer deep reinforcement learning algorithm to address it. Experimental results show that the proposed method can significantly improve the system performance over other baselines and is more robust to environmental changes.
Di Wu 0044, Jikun Kang, Yi Tian Xu, Jimmy Li 0001, Xi Chen 0009, Dmitriy Rivkin, Michael R. M. Jenkin, Taeseop Lee, Intaik Park, Xue Liu 0004, Gregory Dudek
GLOBECOM1
2021 Hierarchical Policy Learning for Hybrid Communication Load Balancing
abstract
Due to the uneven demographic distribution and people’s daily activities, communication systems usually experience highly imbalanced load across different cells. This imbalance leads to unsatisfied users in the congested cells and under-utilized resources in the less-loaded cells. To deal with this issue, existing work migrates the load from heavily loaded cells to lightly loaded cells, by either handing over active mode User Equipment (UEs) to other serving cells, or re-selecting the camping cells for idle mode UEs. In this paper, we further advance the research on Load Balancing (LB) with a hybrid control of both active and idle UEs. This task is challenging, due to the conflicts between Active-UE LB (AULB) and Idle-UE LB (IULB) policies. To overcome this challenge, we propose a Hierarchical Policy Learning (HPL) framework, which coordinates the actions between LB policies with a two-level learning structure. In this way, HPL produces AULB and IULB policies that are better aligned with each other. Extensive simulation results illustrate the efficiency and efficacy of the proposed HPL.
Jikun Kang, Xi Chen 0009, Di Wu 0044, Yi Tian Xu, Xue Liu 0004, Gregory Dudek, Taeseop Lee, Intaik Park
ICC3
2021 Optimizing Cellular Networks via Continuously Moving Base Stations on Road Networks
abstract
Although existing cellular network base stations are typically immobile, the recent development of small form factor base stations and self driving cars has enabled the possibility of deploying a team of continuously moving base stations that can reorganize the network infrastructure to adapt to changing network traffic usage patterns. Given such a system of mobile base stations (MBSes) that can freely move on the road, how should their path be planned in an effort to optimize the experience of the users? This paper addresses this question by modeling the problem as a Markov Decision Process where the actions correspond to the MBSes deciding which direction to go at traffic intersections; states corresponds to the position of MBSes; and rewards correspond to minimization of packet loss in the network. A Monte Carlo Tree Search (MCTS)-based anytime algorithm that produces path plans for multiple base stations while optimizing expected packet loss is proposed. Simulated experiments in the city of Verdun, QC, Canada with varying user equipment (UE) densities and random initial conditions show that the proposed approach consistently outperforms myopic planners, and is able to achieve near-optimal performance.
Yogesh A. Girdhar, Dmitriy Rivkin, Di Wu 0044, Michael R. M. Jenkin, Xue Liu 0004, Gregory Dudek
ICRA3
2021 Residential Electric Load Forecasting via Attentive Transfer of Graph Neural Networks
abstract
An accurate short-term electric load forecasting is critical for modern electric power systems' safe and economical operation. Electric load forecasting can be formulated as a multi-variate time series problem. Residential houses in the same neighborhood may be affected by similar factors and share some latent spatial dependencies. However, most of the existing works on electric load forecasting fail to explore such dependencies. In recent years, graph neural networks (GNNs) have shown impressive success in modeling such dependencies. However, such GNN based models usually would require a large amount of training data. We may have a minimal amount of data available to train a reliable forecasting model for houses in a new neighborhood area. At the same time, we may have a large amount of historical data collected from other houses that can be leveraged to improve the new neighborhood's prediction performance. In this paper, we propose an attentive transfer learning-based GNN model that can utilize the learned prior knowledge to improve the learning process in a new area. The transfer process is achieved by an attention network, which generically avoids negative transfer by leveraging knowledge from multiple sources. Extensive experiments have been conducted on real-world data sets. Results have shown that the proposed framework can consistently outperform baseline models in different areas.
Weixuan Lin, Di Wu 0044
IJCAI2
2020 FiDo: Ubiquitous Fine-Grained WiFi-based Localization for Unlabelled Users via Domain Adaptation
abstract
To fully support the emerging location-aware applications, location information with meter-level resolution (or even higher) is required anytime and anywhere. Unfortunately, most of the current location sources (e.g., GPS and check-in data) either are unavailable indoor or provide only house-level resolutions. To fill the gap, this paper utilizes the ubiquitous WiFi signals to establish a (sub)meter-level localization system, which employs WiFi propagation characteristics as location fingerprints. However, an unsolved issue of these WiFi fingerprints lies in their inconsistency across different users. In other words, WiFi fingerprints collected from one user may not be used to localize another user. To address this issue, we propose a WiFi-based Domain-adaptive system FiDo, which is able to localize many different users with labelled data from only one or two example users. FiDo contains two modules: 1) a data augmenter that introduces data diversity using a Variational Autoencoder (VAE); and 2) a domain-adaptive classifier that adjusts itself to newly collected unlabelled data using a joint classification-reconstruction structure. Compared to the state of the art, FiDo increases average F1 score by 11.8% and improves the worst-case accuracy by 20.2%.
Xi Chen 0009, Chenyi Zhou, Xue (Steve) Liu, Di Wu 0044, Gregory Dudek
WWW5
2017 Boosting Based Multiple Kernel Learning and Transfer Regression for Electricity Load Forecasting
Di Wu 0044, Boyu Wang 0004, Doina Precup, Benoit Boulet
ECML/PKDD (3)1