EDBT 2026 Demo / reviewers in the wild / expert
Ting Wang 0001
dblp:12/2633-1
· DBLP profile ↗
74ranked-venue papers
22as first author
55since 2021 · last 2026
0000-0002-7223-8849ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 52 · 18 first-author · 34 since 2021Systems, architecture and hardware · 15 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-objective reinforcement learning for adaptive load balancing in cloud data centers
Xindi He, Ting Wang 0001 |
Comput. Networks | 2 |
| 2026 | DMT-PPO: Weight-Adaptive Multiobjective Task Offloading With Dynamic Preference Learning in Heterogeneous Edge Computing
Honggang Yuan, Yuxiang Deng, Qin Li 0002, Ting Wang 0001, Yuanming Shi |
IEEE Internet Things J. | 5 |
| 2026 | Throughput-Optimized Service Routing for Microservice Flows in LEO Satellite NetworksabstractSatellite-based microservice systems have emerged as a promising architecture for enabling scalable and flexible service deployment in Low Earth Orbit (LEO) satellite networks, which are increasingly recognized as a key solution to meet the rising demand for global communication and computation, especially in remote and underserved regions. However, routing microservices efficiently in such systems presents major challenges due to the dynamic topology, intermittent connectivity, and unstable link conditions inherent to satellite constellations. These issues become even more severe under high traffic loads. To address these challenges, we propose Service Pressure, a novel routing algorithm specifically designed for satellite-based microservice systems. Service Pressure comprises two key components: first, the construction of an augmented subgraph to model the complex execution and data transmission dependencies of microservices; second, a distributed service routing algorithm that utilizes queue backlogs. This combination enables the algorithm to effectively handle high throughput and adapt to the dynamic network conditions of satellite constellations. By optimizing resource utilization, minimizing latency, and balancing load across satellite nodes, Service Pressure ensures efficient and stable service orchestration even under fluctuating traffic conditions. Extensive simulations demonstrate that our approach significantly outperforms existing routing methods, particularly in terms of throughput, latency, and stability. Service Pressure offers a significant advancement in satellite microservice routing, making it ideal for next-generation space-ground integrated networks. Xindi He, Ting Wang 0001, Yuanming Shi, Xin Liu 0049 |
IEEE Trans. Mob. Comput. | 2 |
| 2026 | Federated Linear Bandit Learning via UAV Aided Over-the-Air ComputationabstractThis paper investigates federated contextual linear bandit learning in a wireless network with a central server and multiple devices. To reduce communication latency, devices interact with the server via over-the-air computation (AirComp) over noisy, fading channels, where signal distortion can occur due to channel imperfections. Departing from traditional AirComp designs for static networks, we propose a novel federated bandit learning framework that leverages unmanned aerial vehicles (UAVs) as mobile servers to aggregate data from distributed IoT devices. To optimize this system, we employ a block coordinate descent method combined with the alternating direction method of multipliers (BCD-ADMM), jointly optimizing the UAV trajectory, receive normalization factor, and transmission power to minimize the time-averaged mean square error (MSE) of AirComp. Our approach addresses the challenge of decentralized data across multiple devices, enabling secure and efficient collaboration without direct data sharing. Theoretical analysis establishes an upper bound on the algorithm's regret, affirming the framework's scalability and robustness against noise. Simulation results support these findings, highlighting notable performance improvements in federated bandit learning with UAV-assisted AirComp. Junkai Qian, Yuning Jiang 0002, Xin Liu 0049, Ting Wang 0001, Yuanming Shi, Colin N. Jones |
IEEE Trans. Mob. Comput. | 5 |
| 2026 | Robust Information Bottleneck for Satellite Edge Inference Over MIMO Channel
Jielin Zhu, Jingyang Zhu, Youlong Wu, Ting Wang 0001, Yuanming Shi, Wei Chen 0002, Khaled Ben Letaief |
IEEE Trans. Wirel. Commun. | 5 |
| 2026 | Satellite Federated Fine-Tuning for Foundation Models in Space Computing Power NetworksabstractAdvancements in artificial intelligence and low-earth orbit satellites have promoted the application of large remote sensing foundation models (FMs) for various downstream tasks. However, direct downloading of these models for fine-tuning on the ground is impeded by privacy concerns and limited bandwidth. Satellite federated learning (FL) offers a solution by enabling model fine-tuning directly on-board satellites and aggregating model updates without data downloading. Nevertheless, for large FMs, the computational capacity of satellites is insufficient to support effective on-board fine-tuning in traditional satellite FL frameworks. To address these challenges, we propose a satellite-ground collaborative federated fine-tuning framework. The key of the framework lies in how to reasonably decompose and allocate model components to alleviate insufficient on-board computation capabilities. During fine-tuning, satellites exchange intermediate results with ground stations or other satellites for forward propagation and back propagation, which brings communication challenges due to the special communication topology of space transmission networks, such as intermittent satellite-ground communication, short duration of satellite-ground communication windows, and unstable inter-orbit inter-satellite links. To reduce transmission delays, we further introduce tailored communication strategies that integrate both communication and computing resources. Specifically, we propose a parallel intra-orbit communication strategy, a topology-aware satellite-ground communication strategy, and a latency-minimization inter-orbit communication strategy to reduce space communication costs. Simulation results demonstrate significant reductions in training time to 33% of on-board training time. Jingyang Zhu, Ting Wang 0001, Yuanming Shi, Chunxiao Jiang, Khaled Ben Letaief |
IEEE Trans. Wirel. Commun. | 3 |
| 2025 | Multi-agent Independent PPO-based Automatic ECN Tuning for High-Speed Data Center NetworksabstractExplicit Congestion Notification (ECN)-based congestion control schemes have been widely adopted in high-speed data center networks (DCNs), where the ECN marking threshold plays a determinant role in guaranteeing a packet lossless DCN. However, existing approaches either employ static settings with immutable thresholds that cannot be dynamically self-adjusted to adapt to network dynamics, or fail to take into account many-to-one traffic patterns and different requirements of different types of traffic, resulting in relatively poor performance. To address these problems, this paper proposes a novel learningbased automatic ECN tuning scheme, named PET, based on the multi-agent Independent Proximal Policy Optimization (IPPO) algorithm. PET dynamically adjusts ECN thresholds by fully considering pivotal congestion-contributing factors, including queue length, output data rate, output rate of ECN-marked packets, current ECN threshold, the extent of incast, and the ratio of mice and elephant flows. PET adopts the Decentralized Training and Decentralized Execution (DTDE) paradigm and combines offline and online training to accommodate network dynamics. PET is also fair and readily deployable with commodity hardware. Comprehensive experimental results demonstrate that, compared with state-of-the-art static schemes and the learningbased automatic scheme, our PET achieves better performance in terms of flow completion time, convergence rate, queue length variance, and system robustness. Ting Wang 0001 |
CLUSTER | 1 |
| 2025 | Latency-Minimal Decentralized LAM Training with Looped Transformers in Heterogeneous LEO Satellite ConstellationsabstractThe traditional approach of transmitting sensing data to ground stations for training large AI models (LAMs) is becoming increasingly impractical due to unstable ground-to-satellite links (GSLs) and growing privacy concerns. Fortunately, the rapid expansion of Low Earth Orbit (LEO) satellite constellations offers a new solution. The increasing number of LEO satellites constitutes a distributed computing network, enabling on-orbit computation without the need for data transmission to ground stations. However, the limited computational resources on satellites result in substantial latency when training edge LAMs directly on a single satellite, and the memory capacity is insufficient to accommodate large models for training. To address these challenges, we propose a latency-optimized decentralized training framework that enables collaborative in-orbit LAM pretraining across heterogeneous LEO satellites without the need for downloading data to the ground. In the proposed architecture, we adopt a looped Transformer architecture that reduces memory usage through cross-layer parameter sharing. To improve training efficiency, we introduce a hybrid parallel training strategy that combines intra-group pipeline parallelism and inter-group data parallelism through satellite grouping. Furthermore, to accommodate heterogeneous satellite capabilities, we formulate an optimization problem for workload balancing and develop a dynamic programming–based strategy to minimize overall system latency. Extensive experiments on a heterogeneous satellite testbed demonstrate that our method significantly outperforms existing baselines in both training speed and memory efficiency. He Xian, Honggang Yuan, Ting Wang 0001, Yuanming Shi |
GLOBECOM | 4 |
| 2025 | MIMO Over-The-Air Federated Learning With Spiking Neural Network Via Lattice CodeabstractSpiking neural networks (SNNs) have emerged as an energy-efficient alternative to the traditional artificial neural networks (ANNs) which are compute-intensive. This paper proposes a novel MIMO over-the-air federated learning scheme trained on SNNs using lattice code. Based on the lattice structure, we design a reliable transceiver with lattice quantizer that can combat the noise and interference from the devices. We further derive a convergence analysis of the proposed method considering the nondifferentiable spikes of SNNs. The experimental results verify that the proposed method is effective by showing that the proposed method can achieve comparable accuracy to the ideal benchmarks and outperform the existing approach by employing a small number of antennas at the server and devices. We also show that SNNs are$23.08 \times$more energy-efficient than ANNs. Chenye Wang, Youlong Wu, Ting Wang 0001, Yuanming Shi |
ICC | 3 |
| 2025 | SAI: Latency-Aware Satellite Edge LAM Inference with Looped TransformerabstractThe rapid advancements in computing and communication capabilities of Low Earth Orbit (LEO) satellites have made it feasible to execute complex and collaborative inorbit computation missions. Transformer-based large AI models (LAMs), known for their exceptional performance in in-context learning (ICL) and prompt-based reasoning, have attracted significant attention, providing powerful intelligence across sectors such as industry and aerospace. However, the significant parameter volume of LAMs poses a substantial challenge for direct deployment on satellites with constrained computing power and energy provision. To address this, the looped Transformer model reduces parameter requirements through layerwise parameter sharing, achieving performance comparable to vanilla Transformer-based LAMs in ICL tasks. Despite this efficiency, the limited and heterogeneous space-borne computing and storage capabilities complicate the orchestration for balanced workload allocation during multi-satellite cooperation. In this paper, we propose SAI, a collaborative multi-satellite space AI system that exploits the memory efficiency of the looped Transformer and the inherent parallelism in batch data processing. SAI enables accelerated on-satellite inference by integrating heterogeneous onboard resources and introducing a novel hybrid approach combining data and pipeline parallelism. This approach supports cross-satellite cooperation with parallelism planning and asynchronous inter-batch overlapping, significantly reducing inference latency and enhancing resource efficiency. Furthermore, SAI optimizes inference latency by formulating it as a shortest-path problem, effectively solved via Dijkstras algorithm. Extensive evaluations demonstrate SAIs superior performance in reducing inference latency and runtime memory usage compared to existing baselines. Honggang Yuan, Yuning Jiang 0002, Xin Liu 0049, Yuanming Shi, Ting Wang 0001 |
ICC | 6 |
| 2025 | Multi-Objective Deep Reinforcement Learning for Adaptive Virtual Machine Allocation in CloudsabstractWith the exponential growth of data and demands for computing capabilities, optimizing resource utilization has become increasingly critical in cloud data centers. Employing virtual machine (VM) allocation technology to maintain hosts within an appropriate workload range holds substantial promise for improving workload balance, energy efficiency, and quality of service (QoS). Existing multi-objective VM allocation strategies based on greedy heuristics and reinforcement learning with predefined fixed objective weights lack generalizability and quick adaptability in dynamic workload scenarios. In this paper, we present MOVMA, based on a novel Multi-Objective Reinforcement Learning (MORL) algorithm to coordinately optimize three objectives, i.e., energy consumption, load balancing in multidimensional resource utilization, and service level agreement (SLA) violations. MOVMA adopts our proposed Sliding Time Window-based Dynamic Weight (STWDW) method to adaptively calculate the weights instantly based on the current system condition, ensuring the actual impact of the parameters. Furthermore, it integrates our proposed Priority-based Selection and Adjustment (PSA) scheme and a Near on-policy Experience Replay (NER) strategy in model training to accelerate convergence and avoid catastrophic forgetting. The experiments conducted on a real-world dataset demonstrate the superior performance of our MOVMA against state-of-the-art multi-objective optimization algorithms. Yuzi Chen, Puyu Cai, Ting Wang 0001 |
ICPADS | 5 |
| 2025 | SAFL: Structure-Aware Personalized Federated Learning via Client-Specific Clustering and SCSI-Guided Model PruningabstractFederated Learning (FL) enables collaborative model training across distributed clients while preserving data privacy. However, conventional FL approaches often struggle to deliver accurate and personalized models in the presence of non-IID data. Although model pruning has been proposed to improve model adaptability, existing methods relying solely on local data often yield sub-optimal sub-models due to limited task-specific information. To address this, we propose SAFL (Structure-Aware Federated Learning), a novel framework that enhances personalization by integrating client clustering with Similar Client Structure Information (SCSI)-guided pruning. SAFL adopts a two-stage process: it first clusters clients based on data similarity and uses aggregated structural insights to guide pruning; then, clients train the resulting sub-models and participate in heterogeneous model aggregation. Extensive experiments on benchmark datasets demonstrate that SAFL achieves superior accuracy and model compactness compared to existing methods, particularly under non-IID settings. These results highlight the effectiveness of structure-aware pruning and collaboration in advancing personalized federated learning. Puyu Cai, Ting Wang 0001 |
ICPADS | 6 |
| 2025 | ECMSA: Dual-Agent Learning-Based Edge Caching with Multi-Strategy Adaptation in Dynamic EnvironmentsabstractWith the proliferation of mobile devices and IoT applications, edge caching has become vital for mitigating network congestion and enhancing user Quality of Experience (QoE). However, traditional caching policies, such as Least Frequently Used (LFU), First-In-First-Out (FIFO), and Least Recently Used (LRU), often struggle to perform effectively in highly dynamic and heterogeneous environments, particularly when content sizes vary significantly. Moreover, existing approaches, whether AI-driven or heuristic-based, typically adopt a single caching strategy, which inherently limits their flexibility and adaptability. To address these limitations, we propose ECMSA, a learning-based multi-strategy edge caching algorithm that integrates a reinforcement learning-driven proactive caching strategy with three conventional reactive caching strategies. Specifically, ECMSA operates in two stages: First, it generates four candidate cache lists—three derived from traditional caching policies (LFU, FIFO, LRU) and one produced by our self-attention-enhanced Deep Deterministic Policy Gradient (Atten-Actor DDPG)-based proactive caching strategy. Next, it employs another Atten-Actor DDPG agent to dynamically select the optimal strategy in real time, leveraging current state features. This dual-agent framework enables continuous learning and adaptation of caching decisions, effectively optimizing content placement and update policies in response to evolving user demands. Extensive experiments conducted on both synthetic and real-world datasets demonstrate that ECMSA achieves 15-17% higher cache-hit ratios and 16-22% lower latency than baseline methods under constrained cache capacities and diverse content sizes. Furthermore, ECMSA exhibits strong robustness and generalization ability, allowing it to rapidly adapt to unseen environments. Ting Wang 0001, Lu Yang 0003, Yuanming Shi, Haibin Cai |
ICPADS | 2 |
| 2025 | HAPFL: Heterogeneity-Aware Personalized Federated Learning via Hierarchical RL and Model DistillationabstractFederated Learning (FL) enables multiple clients to collaboratively train models without sharing raw data, making it well-suited for privacy-preserving applications in heterogeneous IoT environments. However, disparities in client model architectures and computational resources often lead to accuracy degradation and the straggler problem, undermining training efficiency. To address these challenges, we propose HAPFL, a novel Heterogeneity-aware Personalized Federated Learning framework based on multi-level Reinforcement Learning (RL). HAPFL integrates three key components: 1) An RL-based model allocation mechanism that employs a PPO agent to assign appropriately sized models to clients based on their computing capabilities; 2) An RL-based training intensity adjustment scheme that dynamically controls local training epochs per client to reduce straggling latency; 3) A mutual learning scheme using knowledge distillation between each client's local model and a homogeneous lightweight model (LiteModel), which also serves as the global aggregation model to tackle model heterogeneity. Experiments on MNIST, CIFAR-10, and ImageNet-10 demonstrate that HAPFL achieves superior accuracy while reducing overall training time by$20.9 \%-40.4 \%$and straggling latency by$\mathbf{1 9. 0 \% - 4 8. 0 \%}$compared to existing approaches. Ting Wang 0001, Qin Li 0002, Haibin Cai |
ICWS | 2 |
| 2025 | Multi-Task Reinforcement Learning for Collaborative Network Optimization in Data Centers
Ting Wang 0001 |
INFOCOM | 1 |
| 2025 | Robust Multimodal Information Bottleneck for Satellite-to-Ground Task-Oriented CommunicationabstractIn this paper, we study satellite-to-ground taskoriented communication for edge inference tasks, where a satellite extracts, fuses and encodes multimodal feature vectors and then sends them to a ground server under inevitable channel noise conditions for downstream processing. However, the multispectral and multi-resolution characteristics of multimodal satellite remote sensing data render traditional multimodal methods inapplicable. To reduce the data redundancy caused by the high-dimensional and complex multimodal vectors generated onboard while retaining key information and enhancing robustness against channel noise. We propose a Robust Multimodal Information Bottleneck (RMIB) framework which considers channel noise and communication bandwidth and introduces a new information bottleneck optimization objective. By applying this objective through end-to-end training, we optimize the feature extraction, fusion and encodes multimodal data into robust and effective feature vector in noisy communication environments by reducing redundancy and enhancing feature discrimination. To tackle the RMIB objective function, we derive a tractable variational upper bound using the Variational Information Bottleneck technique to overcome the computational intractability of mutual information. Experimental results demonstrate that our method not only outperforms baseline techniques in classification accuracy on three datasets but also enhances robustness against channel noise and reduces communication overhead. Dingzhu Wen, Youlong Wu, Yuanming Shi, Ting Wang 0001 |
ISCC | 5 |
| 2025 | Topology-Aware Routing for Federated Learning Over Multi-Layer Satellite NetworksabstractRecent advancements in space computing power networks, particularly the integration of onboard computing capabilities in Low Earth Orbit (LEO) satellites, have paved the way for federated learning (FL) in satellite networks. Despite its potential, satellite FL faces unique challenges, such as the dynamic nature of satellite networks and the instability of inter-orbit communication links, which complicate global model aggregation. To address these challenges, we explore FL over multi-layer satellite networks, incorporating LEO, Medium Earth Orbit (MEO), and Geostationary Earth Orbit (GEO) satellites. Specifically, by modeling the dynamic network as a series of time-varying graph snapshots, we propose a novel topology-aware FL framework. To optimize the aggregation routing in the multi-layer satellite network, we leverage the directed minimum spanning tree (DMST) problem in graph theory and introduce a communication-efficient satellite aggregation routing algorithm (CESAR), which effectively reduces communication overhead and aggregation delays, ensuring efficient training and model updates across the satellite network. Extensive experimental results validate the efficacy of the proposed framework, demonstrating its potential to overcome the inherent challenges of satellite FL and significantly advance the capabilities of multi-layer satellite networks. Ruanjun Li, Jingyang Zhu, Yijie Mao, Yuanming Shi, Ting Wang 0001, Chunxiao Jiang |
WCNC | 5 |
| 2025 | A Framework for Runtime Safety of Industrial Control Systems Through Runtime VerificationabstractEnsuring the safety of complex industrial control systems (ICS) cannot be fully achieved during the design and development phases. Many uncertainties and unknowns only become apparent during real-world operation, especially in the context of Industry 4.0, where ICS integrate increasing characteristics of cyber-physical systems (CPS), such as openness and connectivity. Runtime verification (RV) is extensively employed to guarantee the runtime safety of systems. However, current RV methods face substantial challenges in ICS, particularly due to extensive device heterogeneity, intricate real-time constraints, and the need for coordinating multiple controllers. In this article, we propose a novel framework that incorporates stream-based RV to ensure the runtime safety of ICS. By leveraging a communication bridge based on the open platform communications unified architecture (OPC UA) standard, our framework achieves platform compatibility. This framework, coupled with its nonintrusive verification feature, is well-suited for scenarios involving heterogeneous devices and collaborative controllers. Additionally, stream-based formal specification captures complex time-sensitive constraints, such as real-time synchronizations involving various signals, including triggering, duration, and timeout. To further enhance safety, the framework offers online correction strategies for addressing runtime violations, aiming to preserve or restore system safety. Experimental results from general case studies demonstrate that our approach surpasses existing methods in managing device heterogeneity, complex real-time constraints, and multicontroller cooperation scenarios. Qin Li 0002, Xia Mao, Ting Wang 0001, Tengfei Li 0002 |
IEEE Internet Things J. | 4 |
| 2025 | Brain-Inspired Decentralized Satellite Learning in Space Computing Power Networks
Peng Yang 0027, Ting Wang 0001, Haibin Cai, Yuanming Shi, Chunxiao Jiang, Linling Kuang |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | GIRP: Energy-Efficient QoS-Oriented Microservice Resource Provisioning via Multi-Objective Multi-Task Reinforcement LearningabstractMicroservice architecture has revolutionized web service development by facilitating loosely coupled and independently developable components distributed as containers or virtual machines. While existing studies emphasize end-to-end latency, this paper investigates energy-efficient quality-of-service (QoS)-oriented microservice provisioning, focusing on both QoS satisfaction and power consumption (PC) conservation. We propose the Green and Intelligent Resource Provision (GIRP) architecture, integrating a data-driven energy-latency-aware resource allocation and scheduling manager to balance latency and PC. To reconcile the trade-offs involved, a dual-objective optimization problem is formulated to minimize latency and energy use by selecting proper servers, allocating CPU cores, and determining service replicas. To address challenges with discrete variables, dual objectives, and implicit mappings, we leverage a model-free deep deterministic policy gradient-based reinforcement learning algorithm. Specifically, we develop a multi-task agent via the Multi-gate Mixture-of-Experts model to simultaneously make two separate actions regarding CPU core numbers and service replica numbers, followed by a single-task agent to determine service scheduling. Extensive experiments on the DeathStarBenchmark testbed validate GIRP’s effectiveness, demonstrating approximately 52% resource savings and a 43% reduction in PC compared to leading methods like Sinan, Firm, and heuristic-based algorithms. These results highlight GIRP’s capability to optimize microservice orchestration by balancing end-to-end latency and power efficiency. Honggang Yuan, Ting Wang 0001, Min Fu 0003, Yuanming Shi |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | Multi-Objective Deep Reinforcement Learning for Function Offloading in Serverless Edge ComputingabstractFunction offloading problems play a crucial role in optimizing the performance of applications in serverless edge computing (SEC). Existing research has extensively explored function offloading strategies based on optimizing a single objective. However, a significant challenge arises when users expect to optimize multiple objectives according to the relative importance of these objectives. This challenge becomes particularly pronounced when the relative importance of the objectives dynamically shifts. Consequently, there is an urgent need for research into multi-objective function offloading methods. In this paper, we redefine the SEC function offloading problem as a dynamic multi-objective optimization issue and propose a novel approach based on Multi-objective Reinforcement Learning (MORL) called MOSEC. MOSEC can coordinately optimize three objectives, i.e., application completion time, User Device (UD) energy consumption, and user cost. To reduce the impact of extrapolation errors, MOSEC integrates a Near-on Experience Replay (NER) strategy during the model training. Furthermore, MOSEC adopts our proposed Earliest First (EF) scheme to maintain the policies learned previously, which can efficiently mitigate the catastrophic policy forgetting problem. Extensive experiments conducted on various generated applications demonstrate the superiority of MOSEC over state-of-the-art multi-objective optimization algorithms. Yaning Yang, Yutong Ye 0001, Jiepin Ding, Ting Wang 0001, Mingsong Chen 0001, Keqin Li 0001 |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | Situation-Dependent Causal Influence-Based Cooperative Multi-Agent Reinforcement LearningabstractLearning to collaborate has witnessed significant progress in multi-agent reinforcement learning (MARL). However, promoting coordination among agents and enhancing exploration capabilities remain challenges. In multi-agent environments, interactions between agents are limited in specific situations. Effective collaboration between agents thus requires a nuanced understanding of when and how agents' actions influence others.To this end, in this paper, we propose a novel MARL algorithm named Situation-Dependent Causal Influence-Based Cooperative Multi-agent Reinforcement Learning (SCIC), which incorporates a novel Intrinsic reward mechanism based on a new cooperation criterion measured by situation-dependent causal influence among agents.Our approach aims to detect inter-agent causal influences in specific situations based on the criterion using causal intervention and conditional mutual information. This effectively assists agents in exploring states that can positively impact other agents, thus promoting cooperation between agents.The resulting update links coordinated exploration and intrinsic reward distribution, which enhance overall collaboration and performance.Experimental results on various MARL benchmarks demonstrate the superiority of our method compared to state-of-the-art approaches. Yutong Ye 0001, Yaning Yang, Mingsong Chen 0001, Ting Wang 0001 |
AAAI | 6 |
| 2024 | Satellite Federated Fine-Tuning for Foundation Models: Architecture Design and System OptimizationabstractWith the surge in the number of low earth orbit (LEO) satellites, continuous research has emerged on using satellite data to train artificial intelligence models. On one hand, traditional centralized training on the ground is not feasible due to privacy concerns and limited bandwidth for downloading raw satellite data. On the other hand, due to the limited energy and computational capability of satellites, training directly on satellites suffers from prolonged latency, especially for large models. To alleviate these issues, we propose a novel satellite-ground collaborative federated fine-tuning architecture, where ground stations (GSs) and satellites collaboratively train a global model without the need for data downloads. In this proposed architecture, satellites serve as edge devices and the ground server serves as a coordinator. However, the short satellite-ground communication windows caused by the high mobility of satellites and the substantial intra-orbit data transmission bring special challenges to the transmission process of federated edge learning. To tackle these challenges, we carefully design the satellite-ground collaborative fine-tuning architecture and utilize an optimized ring all-reduce algorithm and network flow algorithm to enhance the intra-orbit and ground-satellite transmissions, respectively. Experimental results demonstrate that our proposed architecture significantly reduces the training time by 40% compared to training solely on satellite. Peng Yang 0027, Jingyang Zhu, Dingzhu Wen, Ting Wang 0001, Yong Zhou 0006, Yuanming Shi, Chunxiao Jiang |
GLOBECOM | 5 |
| 2024 | Federated Multi-Objective Meta-Reinforcement Learning for Adaptive Edge Task OffloadingabstractWith the proliferation of the Internet of Things (IoT) and mobile network technologies, efficient task offloading in edge computing has become pivotal for optimizing network resource allocation and enhancing data processing speed. However, edge task offloading for diverse applications of different users is typically multi-objective, where the complexity of multi-objective optimization presents significant challenges as wireless channel state and idle resources as well as the interference can change rapidly and the importance attached to different objectives by users may vary depending on the situation. Particularly in cases where the preference weights of these objectives fluctuate over time, traditional optimization techniques are typically unable to provide effective solutions. Moreover, the centralized training utilized by most Artificial Intelligence (AI)-based optimization algorithms raises concerns regarding the potential leakage of local private data to third parties. To address these challenges, we propose a novel federated multi-objective reinforcement learning (FMORL) algorithm, which employs a federated learning framework to perform collaborative learning on distributed nodes working in parallel, allowing for the fast and flexible acquisition of the optimal offloading strategy from dynamic environments, and introduces a meta-learning mechanism to enhance the fast adaptation of the model. Simulation experiments demonstrate that compared with the traditional MORL algorithm, the FMORL algorithm, embedded with meta-learning, improves the overall performance while preserving data privacy. Ting Wang 0001 |
HPCC | 2 |
| 2024 | A lightweight group-based SDN-driven encryption protocol for smart home IoT devices
Arif Raza, Salabat Khan, Shivanshu Shrivastava, Muhammad Wasim Abbas Ashraf, Ting Wang 0001, Kaishun Wu, Lu Wang 0002 |
Comput. Networks | 5 |
| 2024 | Federated Reinforcement Learning for Electric Vehicles Charging Control on Distribution NetworksabstractWith the growing popularity of electric vehicles (EVs), maintaining power grid stability has become a significant challenge. To address this issue, EV charging control strategies have been developed to manage the switch between vehicle-to-grid (V2G) and grid-to-vehicle (G2V) modes for EVs. In this context, multiagent deep reinforcement learning (MADRL) has proven its effectiveness in EV charging control. However, existing MADRL-based approaches fail to consider the natural power flow of EV charging/discharging in the distribution network and ignore driver privacy. To deal with these problems, this article proposes a novel approach that combines multi-EV charging/discharging with a radial distribution network (RDN) operating under optimal power flow (OPF) to distribute power flow in real time. A mathematical model is developed to describe the RDN load. The EV charging control problem is formulated as a Markov decision process (MDP) to find an optimal charging control strategy that balances V2G profits, RDN load, and driver anxiety. To effectively learn the optimal EV charging control strategy, a federated deep reinforcement learning algorithm named FedSAC is further proposed. Comprehensive simulation results demonstrate the effectiveness and superiority of our proposed algorithm in terms of the diversity of the charging control strategy, the power fluctuations on RDN, the convergence efficiency, and the generalization ability. Junkai Qian, Yuning Jiang 0002, Xin Liu 0049, Ting Wang 0001, Yuanming Shi, Wei Chen 0002 |
IEEE Internet Things J. | 5 |
| 2024 | Parameterized Deep Reinforcement Learning With Hybrid Action Space for Edge Task OffloadingabstractMultiaccess edge computing (MEC) has emerged as a promising solution that can enable low-end terminal devices to run large complex applications by offloading their tasks to edge servers. The task offloading strategy, determining how to offload tasks, remains the most critical issue of MEC. Traditional offloading approaches either suffer from high computational complexity or poor self-adjustability to dynamic changes in the edge environment. Deep reinforcement learning (DRL) provides an effective way to tackle these issues. However, most existing DRL-based methods solely consider either a continuous or a discrete action space, where the limited action space results in accuracy loss and restricts the optimality of offloading decisions. Nevertheless, the edge task offloading problem in practice often confronts both discrete and continuous actions. In this article, we propose a tailored proximal policy optimization (PPO)-based method, named Hybrid-PPO, enhanced by the parameterized discrete-continuous hybrid action space. Assisted with Hybrid-PPO, we further design a novel DRL-based multiserver multitask collaborative partial task offloading scheme adhering to a series of specifically built formal models. Experimental results prove that our approach achieves high offloading efficiency and outperforms the existing state-of-the-art offloading schemes in terms of convergence rate, energy cost, time cost, and generalizability under various network conditions. Ting Wang 0001, Yuxiang Deng, Yang Wang 0019, Haibin Cai |
IEEE Internet Things J. | 1 |
| 2024 | GreedW: A Flexible and Efficient Decentralized Framework for Distributed Machine LearningabstractWith the ever-increasing demand for computing power in deep learning, distributed training techniques have proven to be effective in meeting these demands. However, current existing state-of-the-art distributed training frameworks, such as Parameter Server (PS), Ring-All-Reduce, and their varieties, still face significant challenges. In particular, the existence of communication bottlenecks can severely limit the efficiency and scalability of distributed training frameworks, making it difficult to fully and effectively exert the computing power of large-scale clusters, especially in the presence of dynamic and ever-changing network environments. To address these issues and further maximize the utilization of the computing power of clusters, in this paper we propose an efficient and dynamic distributed training framework, named GreedW. GreedW can greatly improve the training efficiency of workers by dynamically constructing an adaptive customized communication network and adaptively scheduling the workload. Specifically, GreedW employs a greedy strategy to dynamically construct the communication network tree in each iteration for gradient transmission with minimum communication cost and applies a heterogeneity-aware workload allocation scheme to adaptively balance the heavy traffic across heterogeneous workers in the cluster taking into account the available computing capabilities of each node, which effectively alleviates the network bottleneck. It is worth noting that GreedW is enabled to dynamically adjust the assigned job on each worker node based on their completion time during each round of model aggregation to ensure that each worker node completes its assignments around the same time, thus mitigating the intractable straggler issue and minimizing their idle waiting time. Comprehensive experimental evaluations on three different-scaled training models (i.e., Mnist-2NN, Mnist-CNN, and TextCNN) for image recognition and natural language processing tasks demonstrate that GreedW outperforms the existing state-of-the-art frameworks in terms of training efficiency, system flexibility, and robustness. Ting Wang 0001, Xin Jiang 0027, Qin Li 0002, Haibin Cai |
IEEE Trans. Computers | 1 |
| 2024 | Towards Intelligent Adaptive Edge Caching Using Deep Reinforcement LearningabstractThe tremendous expansion of edge data traffic poses great challenges to network bandwidth and service responsiveness for mobile computing. Edge caching has emerged as a promising method to alleviate these issues by storing a portion of data at the network edge. However, existing caching approaches suffer from either poor caching efficiency with low content-hit ratio or unintelligence of caching policies lacking self-adjustability. In this paper, we propose ICE, a novel Intelligent Edge Caching scheme using a deep reinforcement learning (DRL) method to capture specific valuable information from the requested data. With the benefit of our proposed popularity model based on Newton's law of cooling, ICE fully takes into account the popularity of the contents to be cached and leverages the formulated Markov decision model to decide whether or not the contents should be cached. Moreover, to further improve the caching efficiency, we propose a novel distributed multi-node caching framework, named DCCC, assisted by a multi-tiered caching hierarchy. Comprehensive experiments show that the single-node ICE scheme greatly improves the cache hit rate and contents exchanging time in comparison with both DRL-based and legacy approaches, and our distributed multi-node caching scheme DCCC further significantly improves the overall utilization of caching space. Ting Wang 0001, Yuxiang Deng, Mingsong Chen 0001, Gang Liu 0038, Jieming Di, Keqin Li 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | Green Federated Learning Over Cloud-RAN With Limited Fronthaul Capacity and Quantized Neural NetworksabstractIn this paper, we propose an energy-efficient federated learning (FL) framework for the energy-constrained devices over cloud radio access network (Cloud-RAN), where each device adopts quantized neural networks (QNNs) to train a local FL model and transmits the quantized model parameter to the remote radio heads (RRHs). Each RRH receives the signals from devices over the wireless link and forwards the signals to the server via the fronthaul link. We rigorously develop an energy consumption model for the local training at devices through the use of QNNs and communication models over Cloud-RAN. Based on the proposed energy consumption model, we formulate an energy minimization problem that optimizes the fronthaul rate allocation, device transmit power allocation, and QNN precision levels while satisfying the limited fronthaul capacity constraint and ensuring the convergence of the proposed FL model to a target accuracy. To solve this problem, we analyze the convergence rate and propose efficient algorithms based on the alternative optimization technique. Simulation results show that the proposed FL framework can significantly reduce energy consumption compared to other conventional approaches. We draw the conclusion that the proposed framework holds great potential for achieving a sustainable and environmentally-friendly FL in Cloud-RAN. Yijie Mao, Ting Wang 0001, Yuanming Shi |
IEEE Trans. Wirel. Commun. | 3 |
| 2024 | Decentralized Over-the-Air Federated Learning by Second-Order Optimization MethodabstractFederated learning (FL) is an emerging technique that enables privacy-preserving distributed learning. Most related works focus on centralized FL, which leverages the coordination of a parameter server to implement local model aggregation. However, this scheme heavily relies on the parameter server, which could cause scalability, communication, and reliability issues. To tackle these problems, decentralized FL, where information is shared through gossip, starts to attract attention. Nevertheless, current research mainly relies on first-order optimization methods that have a relatively slow convergence rate, which leads to excessive communication rounds in wireless networks. To design communication-efficient decentralized FL, we propose a novel over-the-air decentralized second-order federated algorithm. Benefiting from the fast convergence rate of the second-order method, total communication rounds are significantly reduced. Meanwhile, owing to the low-latency model aggregation enabled by over-the-air computation, the communication overheads in each round can also be greatly decreased. The convergence behavior of our approach is then analyzed. The result reveals an error term, which involves a cumulative noise effect, in each iteration. To mitigate the impact of this error term, we conduct system optimization from the perspective of the accumulative term and the individual term, respectively. Numerical experiments demonstrate the superiority of our proposed approach and the effectiveness of system optimization. Peng Yang 0027, Yuning Jiang 0002, Dingzhu Wen, Ting Wang 0001, Colin N. Jones, Yuanming Shi |
IEEE Trans. Wirel. Commun. | 4 |
| 2024 | Over-the-Air Computation Empowered Vertically Split InferenceabstractTo tackle the issue of heterogeneous input raw data samples obtained by different devices and enhance the feature extraction capability of edge devices, we propose a vertically split neural network based edge-device collaborative artificial intelligence (AI) inference framework. The local results calculated by various light-size sub-networks at edge devices are transmitted and aggregated at the server for the downstream inference task. Nevertheless, the transmission of such high-dimensional local results involves severe communication overhead. To resolve this issue, the technique of over-the-air computation (AirComp) is adopted to enable low-latency aggregation. The same entry of all devices’ local results is transmitted over a same wireless resource block and aggregated via the waveform superposition property. Furthermore, to simultaneously support the aggregation of all dimensions of the local results, we consider a broadband channel and leverage orthogonal frequency division multiplexing (OFDM) to divide the system bandwidth into multiple subcarriers which are then assigned for different dimensions. Consequently, an extra degree of freedom is introduced to design the aggregation of all dimensions. We then propose a scheme of joint subcarrier allocation, power allocation, and receiver beamforming to minimize the aggregation distortion and enhance inference performance. Extensive experiments are conducted to verify the superiority of the proposed design over benchmarks. Peng Yang 0027, Dingzhu Wen, Qunsong Zeng, Yong Zhou 0006, Ting Wang 0001, Haibin Cai, Yuanming Shi |
IEEE Trans. Wirel. Commun. | 5 |
| 2023 | Federated Linear Bandit Learning via Over-the-air ComputationabstractIn this paper, we investigate federated contextual linear bandit learning within a wireless system that comprises a server and multiple devices. Each device interacts with the environment, selects an action based on the received reward, and sends model updates to the server. The primary objective is to minimize cumulative regret across all devices within a finite time horizon. To reduce the communication overhead, devices communicate with the server via over-the-air computation (AirComp) over noisy fading channels, where the channel noise may distort the signals. In this context, we propose a customized federated linear bandits scheme, where each device transmits an analog signal, and the server receives a superposition of these signals distorted by channel noise. A rigorous mathematical analysis is conducted to determine the regret bound of the proposed scheme. Both theoretical analysis and numerical experiments demonstrate the competitive performance of our proposed scheme in terms of regret bounds in various settings. Yuning Jiang 0002, Xin Liu 0049, Ting Wang 0001, Yuanming Shi |
GLOBECOM | 4 |
| 2023 | Towards Efficient Workflow Scheduling Over Yarn Cluster Using Deep Reinforcement LearningabstractHadoop Yarn is an open-source cluster manager responsible for resource management and job scheduling. However, data-driven applications are typically organized into workflows that consist of a series of jobs with dependencies. Yarn does not manage users' workflows and only considers the current job rather than the entire workflow when scheduling. In practice, multiple workflows share the same Yarn cluster and are pre-assigned separate Yarn resource queues to avoid mutual interference. However, this coarse-grained resource division can sometimes result in low resource utilization and increased pending time of jobs on the Yarn queue. For instance, one resource queue may have exhausted its quota while still having pending jobs, while other queues may have available resources but cannot begin executing any jobs due to unfulfilled data dependencies. To address this problem, we propose a deep reinforcement learning-based workflow scheduling scheme that takes into account job dependencies, job priorities, and dynamic resource usage. The proposed approach can intelligently identify and utilize free windows of different resource queues. Our simulation results demonstrate that the proposed DRL-based workflow scheduling scheme can significantly reduce the average job latency compared to existing approaches. Jianguo Xue, Ting Wang 0001, Puyu Cai |
GLOBECOM | 2 |
| 2023 | LWSA: A Learning-Based Workflow Scheduling Algorithm for Energy-Efficient UAV Delivery SystemabstractDue to their fast speed and easy deployment, Unmanned Aerial Vehicles (UAVs) have been widely used across various sectors, such as earthquake rescue, medical assistance, and smart agriculture. However, UAVs in delivery networks face significant challenges due to limited battery life and computational capabilities, particularly for tasks that entail intensive computing workflows. In this context, Multi-access Edge Computing (MEC), which provides computing resources in close proximity to mobile terminal devices, has emerged as a promising solution. UAVs can offload computing tasks to MEC resources across diverse Internet of Things (IoT) environments. Although task offloading can enhance their task processing capability, it simultaneously brings additional costs, encompassing data transmission time and energy consumption. To address these issues, this paper proposes a novel workflow scheduling method based on the Proximal Policy Optimization (PPO) algorithm, aimed at optimizing UAV energy consumption within MEC environments. The proposed approach establishes a learning-based workflow scheduling strategy harnessing the adaptability of the PPO algorithm to manage dynamic and intricate scenarios, which facilitates efficient task allocation to optimal computational resources while accounting for flight time constraints. Extensive experiments conducted on various well-known scientific workflow benchmarks in real-world UAV delivery networks validate the effectiveness of our method. Compared with state-of-the-art methods, our approach significantly reduces UAV energy consumption and task completion time, simultaneously increasing UAV’s effective payload capacity. Yutong Ye 0001, Ting Wang 0001, Mingsong Chen 0001 |
ICPADS | 3 |
| 2023 | InitLight: Initial Model Generation for Traffic Signal Control Using Adversarial Inverse Reinforcement LearningabstractDue to repetitive trial-and-error style interactions between agents and a fixed traffic environment during the policy learning, existing Reinforcement Learning (RL)-based Traffic Signal Control (TSC) methods greatly suffer from long RL training time and poor adaptability of RL agents to other complex traffic environments. To address these problems, we propose a novel Adversarial Inverse Reinforcement Learning (AIRL)-based pre-training method named InitLight, which enables effective initial model generation for TSC agents. Unlike traditional RL-based TSC approaches that train a large number of agents simultaneously for a specific multi-intersection environment, InitLight pre-trains only one single initial model based on multiple single-intersection environments together with their expert trajectories. Since the reward function learned by InitLight can recover ground-truth TSC rewards for different intersections at optimality, the pre-trained agent can be deployed at intersections of any traffic environments as initial models to accelerate subsequent overall global RL training. Comprehensive experimental results show that, the initial model generated by InitLight can not only significantly accelerate the convergence with much fewer episodes, but also own superior generalization ability to accommodate various kinds of complex traffic environments. Yutong Ye 0001, Yingbo Zhou 0001, Jiepin Ding, Ting Wang 0001, Mingsong Chen 0001, Xiang Lian 0001 |
IJCAI | 4 |
| 2023 | Model-Contrastive Learning for Backdoor EliminationabstractDue to the popularity of Artificial Intelligence (AI) techniques, we are witnessing an increasing number of backdoor injection attacks that are designed to maliciously threaten Deep Neural Networks (DNNs) causing misclassification. Although there exist various defense methods that can effectively erase backdoors from DNNs, they greatly suffer from both high Attack Success Rate (ASR) and a non-negligible loss in Benign Accuracy (BA). Inspired by the observation that a backdoored DNN tends to form a new cluster in its feature spaces for poisoned data, in this paper, we propose a novel two-stage backdoor defense method, named MCLDef, based on Model-Contrastive Learning (MCL). MCLDef can purify the backdoored model by pulling the feature representations of poisoned data towards those of their clean data counterparts. Due to the shrunken cluster of poisoned data, the backdoor formed by end-to-end supervised learning can be effectively eliminated. Comprehensive experimental results show that, with only 5% of clean data, MCLDef significantly outperforms state-of-the-art defense methods by up to 95.79% reduction in ASR, while in most cases, the BA degradation can be controlled within less than 2%. Our code is available at https://github.com/Zhihao151/MCL. Zhihao Yue, Jun Xia 0003, Zhiwei Ling, Ming Hu 0003, Ting Wang 0001, Xian Wei, Mingsong Chen 0001 |
ACM Multimedia | 5 |
| 2023 | Parameterized deep reinforcement learning with hybrid action space for energy efficient data center networksabstractTo ensure the delivery of high-performance and reliable services, data center networks (DCNs) are often over-provisioned for peak workload and traffic bursts. However, in real-world data centers, network traffic seldom reaches peak capacity of the network, resulting in significant energy waste. Traditional energy conservation approaches either suffer from high computational complexity and low solution quality, or their strategies cannot be dynamically adjusted to accommodate changes in data center network traffic. Deep reinforcement learning (DRL) provides an effective way to deal with these issues. However, most of the existing DRL-based schemes only consider either a continuous action space or a discrete action space, which greatly restricts the optimality of decisions. To solve these problems, this paper proposes a novel DRL-based DCN energy optimization framework, named SmartDCN. Specifically, SmartDCN consists of a traffic prediction module (TPM) and an energy optimization module (EOM). TPM incorporates an improved LSTM model JANET with an attention mechanism providing a high prediction accuracy, while EOM integrates our newly proposed parameterized DRL algorithm , named PAS-DQN, combining with the discrete-continuous hybrid action space. PAS-DQN implements a two-level control mechanism for the network, using TPM to predict future traffic in the data center as input. It is devoted to dynamically aggregating current traffic and makes tradeoffs between energy efficiency, performance, and robustness to optimize the network’s power consumption by dynamically calculating the minimum required network subset and turning off the non-involved network devices to achieve power savings. Experimental results show that SmartDCN significantly outperforms the existing state-of-the-art schemes in terms of energy savings under various network conditions. Ting Wang 0001, Xi Fan, Haibin Cai, Yang Wang 0019 |
Comput. Networks | 1 |
| 2023 | JointPS: Joint Parameter Server Placement and Flow Scheduling for Machine Learning ClustersabstractTo distill more information from training data, more parameters are introduced into machine learning models. As a result, communication becomes the bottleneck of Distributed Machine Learning (DML) systems. To alleviate the communication resource contention among DML jobs, which prolongs the time to train machine learning models, in machine learning clusters, JointPS is proposed in this paper. JointPS first minimizes the completion time of a single training epoch for each DML job via jointly optimizing the parameter server placement and flow scheduling, and predicts the number of remaining training epochs for each DML job by leveraging a dynamic model fitting method. Then, JointPS can estimate the remaining time to complete each DML job. According to such estimation, JointPS schedules DML jobs following the Minimum Remaining Time First (MRTF) principle to minimize the average job completion time. To the best of our knowledge, JointPS should be the first work that minimizes the average completion time of network-intensive DML training jobs by jointly optimizing the parameter server placement and flow scheduling without modifying the DML models and training procedures. Through both testbed experiments and extensive simulations, we demonstrate that JointPS can reduce the average completion time of DML jobs by up to 88% compared with state-of-the-art technology. Yangming Zhao, Gongming Zhao, Yunfei Hou, Ting Wang 0001, Chunming Qiao |
IEEE Trans. Computers | 5 |
| 2023 | Hierarchical Relational Graph Learning for Autonomous Multirobot Cooperative Navigation in Dynamic EnvironmentsabstractAs a specific kind of cyber–physical systems (CPSs), autonomous robot clusters play an important role in various intelligent manufacturing fields. However, due to the increasing design complexity of robot clusters, it is becoming more and more challenging to guarantee the safety and efficiency for multirobot cooperative navigation in dynamic and complex environments. Although deep reinforcement learning (DRL) shows great potential in learning multirobot cooperative navigation policies, existing DRL-based approaches suffer from scalability issues and rarely consider the transferability of trained policies to new tasks. To address these problems, this article presents a novel DRL-based multirobot cooperative navigation approach named HRMR-Navi that equips each robot with both a two-layered hierarchical graph network model and an attention-based communication model. In our approach, the hierarchical graph network model can efficiently figure out hierarchical relations among all agents that either cooperate for efficiency or avoid obstacles for safety to derive more advanced strategies, and the communication model can accurately form a global view of the environment for a specific robot, thus, the multirobot cooperation efficiency can be further strengthened. Meanwhile, we propose an improved proximal policy optimization (PPO) algorithm based on the Maximum Entropy Reinforcement Learning, named MEPPO, to enhance the robot exploration ability. Comprehensive experimental results demonstrate that, compared with state-of-the-art approaches, HRMR-Navi can achieve more efficient cooperative navigation with less time cost, lower collision rate, higher scalability, and better knowledge transferability. Ting Wang 0001, Mingsong Chen 0001, Keqin Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | FairLight: Fairness-Aware Autonomous Traffic Signal Control With Hierarchical Action SpaceabstractAlthough reinforcement learning (RL) approaches are promising in autonomous traffic signal control (TSC), they often suffer from the unfairness problem that causes extremely long waiting time at intersections for partial vehicles. This is mainly because the traditional RL methods focus on optimizing the overall traffic performance, while the fairness of individual vehicles is neglected. To address this problem, we propose a novel RL-based method named FairLight for the fair and efficient control of traffic with variable phase duration. Inspired by the concept of user satisfaction index (USI) proposed in the transportation field, we introduce a fairness index in the design of key RL elements, which specially considers the travel quality (e.g., fairness). Based on our proposed hierarchical action space method, FairLight can accurately allocate the duration of traffic lights for selected phases. Experimental results obtained from various well-known traffic benchmarks show that, compared with the state-of-the-art RL-based TSC methods, FairLight can not only achieve better fairness performance but also improve the control quality from the perspectives of the average travel time of vehicles and RL convergence speed. Yutong Ye 0001, Jiepin Ding, Ting Wang 0001, Junlong Zhou, Xian Wei, Mingsong Chen 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Reconfigurable Intelligent Surfaces Empowered Green Wireless Networks With User Admission ControlabstractReconfigurable intelligent surface (RIS) has emerged as a cost-effective and energy-efficient technique for 6G. By adjusting the phase shifts of passive reflecting elements, RIS is capable of suppressing the interference and combining the desired signals constructively at receivers, thereby significantly enhancing the performance of communication system. In this paper, we consider a green multi-user multi-antenna cellular network, where multiple RISs are deployed to provide energy-efficient communication service to end users. We jointly optimize the phase shifts of RISs, beamforming of the base stations, and the active RIS set with the aim of minimizing the power consumption of the base station (BS) and RISs subject to the quality of service (QoS) constraints of users and the transmit power constraint of the BS. However, the problem is mixed combinatorial and non-convex, and there is a potential infeasibility issue when the QoS constraints cannot be guaranteed by all users. To deal with the infeasibility issue, we further investigate a user admission control problem to jointly optimize the transmit beamforming, RIS phase shifts, and the admitted user set. A unified alternating optimization (AO) framework is then proposed to solve both the power minimization and user admission control problems. Specifically, we first decompose the original non-convex problem into several rank-one constrained optimization subproblems via matrix lifting. A difference-of-convex (DC) algorithm is then developed to solve each decomposed subproblem. The proposed AO framework efficiently minimizes the power consumption of wireless networks as well as user admission control when the QoS constraints cannot be guaranteed by all users. To further address the complexity-sensitive issue for practical implementation, we propose an alternative low-complexity beamforming and RISs phase shifts design algorithm based on zero-forcing (ZF) to enable the green cellular networks. Jinglian He, Yijie Mao, Yong Zhou 0006, Ting Wang 0001, Yuanming Shi |
IEEE Trans. Commun. | 4 |
| 2023 | CERT-DF: A Computing-Efficient and Robust Distributed Deep Forest Framework With Low Communication OverheadabstractAs an alternative to the deep learning model, deep forest outperforms deep neural networks in many aspects with fewer hyperparameters and better robustness. To improve the computing performance of deep forest, ForestLayer proposes an efficient task-parallel algorithm S-FTA at a fine sub-forest granularity, but the granularity of the sub-forest cannot be adaptively adjusted. BLB-gcForest further proposes an adaptive sub-forest splitting algorithm to dynamically adjust the sub-forest granularity. However, with distributed storage, its BLB method needs to scan the whole dataset when sampling, which generates considerable communication overhead. Moreover, BLB-gcForest's tree-based vector aggregation produces extensive redundant transfers and significantly degrades the system's performance in vector aggregation stage. To deal with these existing issues and further improve the computing efficiency and scalability of the distributed deep forest, in this paper, we propose a novel Computing-Efficient and RobusT distributed Deep Forest framework, named CERT-DF. CERT-DF integrates three customized schemes, namely, block-level pre-sampling, two-stage pre-aggregation, and system-level backup. Specifically, CERT-DF adopts the block-level pre-sampling method to implement data blocks' local sampling eliminating frequent data remote access and maximizing parallel efficiency, applies the two-stage pre-aggregation method to adjust the class vector aggregation granularity to greatly decrease the communication overhead, and leverages the system-level backup method to enhance the system's disaster tolerance and immensely accelerate task recovery with minimal system resource overhead. Comprehensive experimental evaluations on multiple datasets show that our CERT-DF significantly outperforms the state-of-the-art approaches with higher computing efficiency, lower system resource overhead, and better system robustness while ensuring good accuracy. Li'an Xie, Ting Wang 0001, Shuyi Du, Haibin Cai |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | Multi-Agent Reinforcement Learning for Dynamic Resource Management in 6G in-X SubnetworksabstractThe 6G network enables a subnetwork-wide evolution, resulting in a “network of subnetworks”. However, due to the dynamic mobility of wireless subnetworks, the data transmission of intra-subnetwork and inter-subnetwork will inevitably interfere with each other, which poses a great challenge to radio resource management. Moreover, most existing approaches require the instantaneous channel gain between subnetworks, which are usually difficult to be collected. To tackle these issues, in this paper we propose a novel effective intelligent radio resource management method using multi-agent deep reinforcement learning (MARL), which only needs the sum of received power, named received signal strength indicator (RSSI), on each channel instead of channel gains. However, to directly separate individual interference from RSSI is an almost impossible thing. To this end, we further propose a novel MARL architecture, named GA-Net, which integrates a hard attention layer to model the importance distribution of inter-subnetwork relationships based on RSSI and excludes the impact of unrelated subnetworks, and employs a graph attention network with a multi-head attention layer to exact the features and calculate their weights that will impact individual throughput. Experimental results prove that our proposed framework significantly outperforms both traditional and MARL-based methods in various aspects. Ting Wang 0001, Qiang Feng 0004, Chenhui Ye, Tao Tao 0004, Lu Wang 0002, Yuanming Shi, Mingsong Chen 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2022 | MonitorLight: Reinforcement Learning-based Traffic Signal Control Using Mixed Pressure MonitoringabstractAlthough Reinforcement Learning (RL) has achieved significant success in the Traffic Signal Control (TSC), most of them focus on the design of RL elements while the impact of the phase duration is neglected. Due to the lack of exploring dynamic phase duration, the overall performance and convergence rate of RL-based TSC approaches cannot be guaranteed, which may result in poor adaptability of RL methods to different traffic conditions. To address these issues, in this paper, we formulate a novel phase-duration-aware TSC (PDA-TSC) problem and propose an effective RL-based TSC approach, named MonitorLight. Our approach adopts a new traffic indicator, mixed pressure, which enables RL agents to simultaneously analyze the impacts of stationary and moving vehicles on intersections. Based on the observed mixed pressure of intersections, RL agents can autonomously determine whether or not to change the current signals in real-time. In addition, MonitorLight can adjust the control method for scenarios with different real-time requirements and achieve excellent results in different situations. Extensive experiments on both real-world and synthetic datasets demonstrate that MonitorLight outperforms the current state-of-the-art IPDALight by up to 2.84% and 5.71% in average vehicle travel time, respectively. Moreover, our method significantly speeds up the convergence, leading IPDALight by 36.87% and 34.58% in the start to converge episode and jumpstart performance, respectively. Zekuan Fang, Ting Wang 0001, Xiang Lian 0001, Mingsong Chen 0001 |
CIKM | 3 |
| 2022 | Eliminating Backdoor Triggers for Deep Neural Networks Using Attention Relation Graph DistillationabstractDue to the prosperity of Artificial Intelligence (AI) techniques, more and more backdoors are designed by adversaries to attack Deep Neural Networks (DNNs). Although the state-of-the-art method Neural Attention Distillation (NAD) can effectively erase backdoor triggers from DNNs, it still suffers from non-negligible Attack Success Rate (ASR) together with lowered classification ACCuracy (ACC), since NAD focuses on backdoor defense using attention features (i.e., attention maps) of the same order. In this paper, we introduce a novel backdoor defense framework named Attention Relation Graph Distillation (ARGD), which fully explores the correlation among attention features with different orders using our proposed Attention Relation Graphs (ARGs). Based on the alignment of ARGs between teacher and student models during knowledge distillation, ARGD can more effectively eradicate backdoors than NAD. Comprehensive experimental results show that, against six latest backdoor attacks, ARGD outperforms NAD by up to 94.85% reduction in ASR, while ACC can be improved by up to 3.23%. Jun Xia 0003, Ting Wang 0001, Jiepin Ding, Xian Wei, Mingsong Chen 0001 |
IJCAI | 2 |
| 2022 | Blocking Island Paradigm Enhanced Intelligent Coordinated Virtual Network Embedding Based on Deep Reinforcement LearningabstractAs an efficient technique for resource sharing in data centers, network virtualization enables resource multiplexing by allowing multiple heterogeneous virtual networks (VNs) to simultaneously coexist on the shared substrate infrastructure. How to effectively embed the VNs onto the substrate network is known as the virtual network embedding (VNE) problem. However, as an NP-hard problem, the VNE problem-solving suffers a high computation complexity. Artificial Intelligence (AI) provides a promising way to alleviate these issues. However, the existing AI-based works still cannot fully and efficiently leverage substrate network information to formulate embedding policies. To this end, in this paper we propose a novel deep reinforcement learning (DRL) based coordinated VNE algorithm, called Intelligent Coordinated Embedding (ICE). To reduce the computation complexity, ICE adopts an efficient resource abstraction model, Blocking Island (BI), which greatly reduces the search space. With the benefit of DRL and BI, ICE can efficiently adjust embedding strategies according to the environment states, aiming to maximize resource utilization and overall revenue while minimizing the embedding cost. The experimental results prove that ICE outperforms both the traditional non-DRL-based approach and the state-of-the-art DRL-based approach. Ting Wang 0001, Peng Yang 0027, Haibin Cai |
SECON | 1 |
| 2022 | Towards an energy-efficient Data Center Network based on deep reinforcement learning
Yang Wang 0019, Ting Wang 0001, Gang Liu 0038 |
Comput. Networks | 3 |
| 2022 | IPDALight: Intensity- and phase duration-aware traffic signal control based on Reinforcement Learning
Wupan Zhao, Yutong Ye 0001, Jiepin Ding, Ting Wang 0001, Tongquan Wei, Mingsong Chen 0001 |
J. Syst. Archit. | 4 |
| 2022 | PervasiveFL: Pervasive Federated Learning for Heterogeneous IoT SystemsabstractFederated learning (FL) has been recognized as a promising collaborative on-device machine learning method in the design of Internet of Things (IoT) systems. However, most existing FL methods fail to deal with IoT applications that contain a variety of IoT devices equipped with different types of neural network (NN) models. This is because traditional FL methods assume that local models on devices should have the same architecture as the global model on cloud. To address this problem, we propose a novel framework named PervasiveFL that enables efficient and effective FL among heterogeneous IoT devices. Without modifying original local models, PervasiveFL installs one lightweight NN model named modellet on each device. By using the deep mutual learning (DML) and our entropy-based decision gating (EDG) method, modellets and local models can selectively learn from each other through soft labels using locally captured data. Meanwhile, since modellets are of the same architecture, the learned knowledge by modellets can be shared among devices in a traditional FL manner. In this way, PervasiveFL can be pervasively applied to any heterogeneous IoT system. Comprehensive experimental results on four well-known datasets show that PervasiveFL can not only pervasively enable FL among heterogeneous devices within a large-scale IoT system, but also significantly enhance the inference accuracy of heterogeneous IoT devices with low communication overhead. Jun Xia 0003, Tian Liu 0005, Zhiwei Ling, Ting Wang 0001, Xin Fu 0001, Mingsong Chen 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | BLB-gcForest: A High-Performance Distributed Deep Forest With Adaptive Sub-Forest SplittingabstractAs an emulous alternative to deep neural networks, Deep Forest emerges with features like low complexity, fewer hyper-parameters, and good robustness, which are predominantly desired in distributed computing applications and ecosystems. Recently, an efficient distributed Deep Forest system, named ForestLayer, was proposed, designing a fine-grained sub-Forest-based task-parallel algorithm to improve the parallel computing efficiency of Deep Forest. However, the sub-Forest splitting of ForestLayer is static and one-off without adaptability to the computing environment, nevertheless, the size of splitting granularity has a significant impact on the system performance. To further improve the computing efficiency and scalability of the distributed Deep Forest, in this paper, we propose a novel distributed Deep Forest algorithm, named BLB-gcForest (Bag of Little Bootstraps-gcForest), which augments the gcForest (multi-Grained Cascade Forest) approach for constructing Deep Forest. BLB-gcForest carries out parallel computation for each tree in sub-Forests at a finer parallel granularity and integrates with the Bag of Little Bootstraps (BLB) mechanism to reduce massive transmitted feature instances for Cascade Forest Layers, utterly improving both computation efficiency and communication efficiency. Moreover, to solve the problem of the forest splitting granularity, we further design an adaptive sub-Forest splitting algorithm to ensure the maximum resource utilization for parallel computation of each sub-Forest. Experimental results on four well-known large-scale datasets, namely YEAST, LETTER, MNIST, CIFAR10, show that the training efficiency of BLB-gcForest achieves up to 20.3x and 1.64x speedups compared with the state-of-the-art gcForest and ForestLayer, respectively while guaranteeing higher accuracy and better robustness Zexi Chen, Ting Wang 0001, Haibin Cai, Subrota K. Mondal, Jyoti Prakash Sahoo |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Over-the-Air Federated Learning via Second-Order OptimizationabstractFederated learning (FL) is a promising learning paradigm that can tackle the increasingly prominent isolated data islands problem while keeping users’ data locally with privacy and security guarantees. However, FL could result in task-oriented data traffic flows over wireless networks with limited radio resources. To design communication-efficient FL, most of the existing studies employ the first-order federated optimization approach that has a slow convergence rate. This however results in excessive communication rounds for local model updates between the edge devices and edge server. To address this issue, in this paper, we instead propose a novel over-the-air second-order federated optimization algorithm to simultaneously reduce the communication rounds and enable low-latency global model aggregation. This is achieved by exploiting the waveform superposition property of a multi-access channel to implement the distributed second-order optimization algorithm over wireless networks. The convergence behavior of the proposed algorithm is further characterized, which reveals a linear-quadratic convergence rate with an accumulative error term in each iteration. We thus propose a system optimization approach to minimize the accumulated error gap by joint device selection and beamforming design. Numerical results demonstrate the system and communication efficiency compared with the state-of-the-art approaches. Peng Yang 0027, Yuning Jiang 0002, Ting Wang 0001, Yong Zhou 0006, Yuanming Shi, Colin N. Jones |
IEEE Trans. Wirel. Commun. | 3 |
| 2021 | ICE: Intelligent Caching at the EdgeabstractThe unprecedented growth of mobile data traffic brings unique challenges for network bandwidth and server resources to meet the diverse QoE (Quality of Experience). Caching becomes a promising way to alleviate these issues by storing a subset of data at the network edge, for which caching policy becomes critical. To this end, various caching schemes have been put forward, however, these schemes are either not intelligent lacking the ability of self-learning and self-decision-making, or inefficient with low data hit rate. Based on these observations, in this paper, we propose a novel Intelligent Caching framework at the Edge, named ICE, via deep reinforcement learning to capture certain valued information of the requested data. Notably, in our approach, the popularity of the data to be cached will be explored and considered. A Markov decision model is further developed to determine whether the data should be cached. The evaluation shows that ICE greatly improves the hit rate in comparison with the state-of-the-art approaches, and reduces the energy consumption for data transmission. Furthermore, based on ICE, the users' QoE is greatly improved. In conclusion, both theoretical analysis and experimental results prove the effectiveness and high performance of ICE compared with conventional strategies. Ting Wang 0001, Mingsong Chen 0001, Gang Liu 0038, Jieming Di, Shui Yu 0001 |
GLOBECOM | 1 |
| 2021 | UAV-Assisted Over-the-Air ComputationabstractOver-the-air computation (AirComp) provides a promising way to support ultrafast aggregation of distributed data. However, its performance cannot be guaranteed in long-distance transmission due to the distortion induced by the channel fading and noise. To unleash the full potential of AirComp, this paper proposes to use a low-cost unmanned aerial vehicle (UAV) acting as a mobile base station to assist AirComp systems. Specifically, due to its controllable high-mobility and high-altitude, the UAV can move sufficiently close to the sensors to enable line-of-sight transmission and adaptively adjust all the links' distances, thereby enhancing the signal magnitude alignment and noise suppression. Our goal is to minimize the time-averaging mean-square error for AirComp by jointly optimizing the UAV trajectory, the scaling factor at the UAV, and the transmit power at the sensors, under constraints on the UAV’s predetermined locations and flying speed, sensors’ average and peak power limits. However, due to the highly coupled optimization variables and time-dependent constraints, the resulting problem is non-convex and challenging. We thus propose an efficient iterative algorithm by applying the block coordinate descent and successive convex optimization techniques. Simulation results verify the convergence of the proposed algorithm and demonstrate the performance gains and robustness of the proposed design compared with benchmarks. Min Fu 0003, Yong Zhou 0006, Yuanming Shi, Ting Wang 0001, Wei Chen 0002 |
ICC | 4 |
| 2021 | Multipath-aware TCP for Data Center Traffic Load-balancingabstractTraffic load-balancing is important to data center performance. However, existing data center load-balancing solutions are either limited to simple topologies or cannot provide satisfactory performance. In this paper, we propose a multipath-aware TCP (MA-TCP) which can sense the path migration of TCP flows. With this new mechanism, the reduction in TCP congestion window due to packet reordering during the path migration can be avoided. This, in turn, makes the path migration more timely as soon as the original path is congested. Furthermore, if the new path is congested (again), the flow can securely continue to migrate without worrying about transmitting rate reduction. Through NS-3 simulations, we show that MA-TCP achieves better flow completion time (FCT) than existing data center load-balancing solutions. Yu Xia 0001, Jinsong Wu 0001, Jingwen Xia, Ting Wang 0001, Sun Mao |
IWQoS | 4 |
| 2018 | eMPTCP: Towards High Performance Multipath Data Transmission by Leveraging SDNabstractMotivated by the poor performance of MPTCP when used for bulk transfers in multipathed networks, in this paper we propose an efficient MPTCP protocol variant named eMPTCP aiming to achieve high throughput, low latency and good bottleneck fairness. The novel MPTCP variant eMPTCP can prevent in-network buffer overflow and packet loss using a distributed and reactive approach for bandwidth allocation and adjusting the congestion window of subflows on multiple paths in a coordinated fashion. It employs ECN feedback and latency (in terms of RTT) to modulate the congestion window via a gamma correction function. This novel congestion- and latency-aware window adjustment mechanism behaves very helpful in handling the traffic bursts, easing the buffer pressure on switches, and greatly mitigates the incast issue. Besides, in this solution SDN is leveraged to compute a set of optimal available routes for subflows and actively adjust the number of subflows of each MPTCP flow according to the instantaneous traffic condition. Moreover, this novel MPTCP protocol variant eMPTCP works well with existing switch hardware and is able to coexist with legacy TCP. It ensures that a multipath flow will not take up more capacity on any shared paths than if it was a single path TCP flow using only one of those paths, which guarantees it will not unduly harm other flows. Simulation results show that eMPTCP achieves smaller MPTCP convergence time, higher aggregate throughput and lower flow completion time than existing legacy competitors, and largely improves the application performance and user experience, making the network more robust and faster. Ting Wang 0001, Mounir Hamdi |
GLOBECOM | 1 |
| 2018 | A cost-effective low-latency overlaid torus-based data center network architecture
Ting Wang 0001, Lu Wang 0002, Mounir Hamdi |
Comput. Commun. | 1 |
| 2018 | Achieving Energy Efficiency in Data Centers Using an Artificial Intelligence Abstraction ModelabstractToday's data center networks are usually over-provisioned for peak workloads. This leads to a great waste of energy since in practice traffic rarely ever hits peak capacity resulting in the links being under-utilized most of the time. Furthermore, the traditional non-traffic-aware routing mechanisms worsen the situation. From the perspective of resource allocation and routing, this paper aims to implement a green data center network and save as much energy as possible. With the benefit of blocking island paradigm, we present a general framework trying to maximize the network power conservation and minimize sacrifices of network performance and reliability. The bandwidth allocation mechanism together with power-aware routing algorithm achieve a bandwidth guaranteed green tighter network. Moreover, our fast efficient heuristics for allocating bandwidth enable the system to scale to large sized data centers. The evaluation result shows that achieving up to more than 50 percent power savings are feasible while guaranteeing network performance and reliability. Ting Wang 0001, Yu Xia 0001, Jogesh K. Muppala, Mounir Hamdi |
IEEE Trans. Cloud Comput. | 1 |
| 2017 | Enforcing timely network policies installation in OpenFlow-based software defined networksabstractAs an efficient network innovation enabler, software defined network is designed to address the networking needs that are poorly addressed by existing networks, and makes it easier to create and introduce new abstractions in networking, simplifying network management and facilitating network evolution. The OpenFlow API enables secure communication between controllers and switches, and standardizes the communications. However, there exist critical issues during the procedure of distributing network policies among the switches (especially for in-band communication schemes), which impose various implicit negative impacts on network reliability and efficiency, such as additional computation overhead on both controllers and switches, communication overhead on secure channel, and waste of storage resources on switches. Based on these observations, this paper proposes four practical solutions to deal with these issues. The evaluation results reveal that the proposed solutions reduce the network latency by 50%, improve the goodput by 10-15%, and decrease the hardware cost by 25% at most, which convinces the effectiveness of proposed solutions. Ting Wang 0001, Mounir Hamdi, Jie Chen 0076 |
ICC | 1 |
| 2017 | JOTA: Joint optimization for the task assignment of sketch-based measurement
Zhiyang Su, Ting Wang 0001, Mounir Hamdi |
Comput. Commun. | 2 |
| 2016 | Presto: Towards efficient online virtual network embedding in virtualized cloud data centers
Ting Wang 0001, Mounir Hamdi |
Comput. Networks | 1 |
| 2016 | Towards cost-effective and low latency data center network architecture
Ting Wang 0001, Zhiyang Su, Yu Xia 0001, Mounir Hamdi |
Comput. Commun. | 1 |
| 2015 | JieLin: A Scalable and Fault Tolerant Server-Centric Data Center Network ArchitectureabstractTo support the fast growing cloud computing services and provide a core infrastructure to meet the increasing computing and storage requirements, the number of servers in today's data centers is expanding exponentially, which leads to enormous challenges in designing an efficient and cost-effective data center network. Traditional proposals either suffer from poor reliability, endure performance bottleneck, scale too slowly, or are expensive to construct. Motivated by these challenges, this paper presents the design, analysis, and implementation of JieLin, a novel server- centric network architecture that has many desirable features for data center networking. Besides the excellent scalability and good fault tolerance, JieLin also achieves high performance in bisection bandwidth, average path length, aggregate bottleneck throughput, and cost-effectiveness. Moreover, in order to maximize the theoretical performance of JieLin a congestion- aware fault-tolerant adaptive routing algorithm has been specially designed. In addition to theoretical analysis, extensive simulations are conducted to further prove the feasibility and good performance of JieLin. Ting Wang 0001, Mounir Hamdi |
GLOBECOM | 1 |
| 2015 | CLOT: A cost-effective low-latency overlaid torus-based network architecture for data centersabstractIn this paper, we present the design, analysis, and implementation of a novel data center network architecture named CLOT, which delivers significant reduction in the network diameter, network latency, and infrastructure cost. CLOT is built based on a switchless torus topology by adding only a number of most beneficial low-end switches in a proper way. Forming the servers in close proximity of each other in torus topology well implements the network locality. The extra layer of switches largely shortens the average routing path length of torus network, which increases the communication efficiency. We show that CLOT can achieve lower latency, smaller routing path length, higher bisection bandwidth and throughput, and better fault tolerance compared to both conventional hierarchical data center networks as well as the recently proposed CamCube network. Coupled with the coordinate based translated IP addresses, the carefully designed POW routing algorithm helps CLOT achieve its maximum theoretical performance. The sufficient mathematical analysis and theoretical derivation prove both guaranteed and ideal performance of CLOT. Ting Wang 0001, Zhiyang Su, Yu Xia 0001, Mounir Hamdi |
ICC | 1 |
| 2015 | COSTA: Cross-layer optimization for sketch-based software defined measurement task assignmentabstractSketch-based measurement provides traffic data summary in a memory-efficient way with provable accuracy bound. Recent advances in software defined networking (SDN) facilitate the development and implementation of sketch-based measurement applications. However, sketch-based measurement usually requires TCAMs which are precious resource in switch to match packet fields. The key challenge for sketch-based measurement is how to accept more concurrent measurement tasks with minimum resource usage. Existing proposals attempt to achieve this goal by exploring different task assignment algorithms. We argue that by sacrificing a small amount of accuracy, the resource usage can be decreased dramatically. In this paper, we propose COSTA, a novel system to improve the performance of the task assignment for sketch-based measurement. We utilize the cross-layer information between the application and the task assignment layers to formulate the problem as a mixed integer nonlinear programming problem. Due to its high computational complexity, we divide the initial problem into two stages and develop a two-stage heuristic to produce task assignment efficiently. In particular, we present an algorithm which guarantees (1 + α) approximation ratio to solve the second stage task assignment. Extensive experiments with three different measurement tasks and real packet traces demonstrate that COSTA significantly reduces the resource usage by up to 40% and accepts 30% more tasks. Zhiyang Su, Ting Wang 0001, Mounir Hamdi |
IWQoS | 2 |
| 2015 | CeMon: A cost-effective flow monitoring system in software defined networks
Zhiyang Su, Ting Wang 0001, Yu Xia 0001, Mounir Hamdi |
Comput. Networks | 2 |
| 2015 | Designing efficient high performance server-centric data center network architecture
Ting Wang 0001, Zhiyang Su, Yu Xia 0001, Jogesh K. Muppala, Mounir Hamdi |
Comput. Networks | 1 |
| 2014 | FlowCover: Low-cost flow monitoring scheme in software defined networksabstractNetwork monitoring and measurement are crucial in network management to facilitate quality of service routing and performance evaluation. Software Defined Networking (SDN) makes network management easier by separating the control plane and data plane. Network monitoring in SDN is lightweight as operators only need to install a monitoring module into the controller. Active monitoring techniques usually introduce too many overheads into the network. The state-of-the-art approaches utilize sampling method, aggregation flow statistics and passive measurement techniques to reduce overheads. However, little work in literature has focus on reducing the communication cost of network monitoring. Moreover, most of the existing approaches select the polling switch nodes by sub-optimal local heuristics. Inspired by the visibility and central control of SDN, we propose FlowCover, a low-cost high-accuracy monitoring scheme to support various network management tasks. We leverage the global view of the network topology and active flows to minimize the communication cost by formulating the problem as a weighted set cover, which is proved to be NP-hard. Heuristics are presented to obtain the polling scheme efficiently and handle flow changes practically. We build a simulator to evaluate the performance of FlowCover. Extensive experiment results show that FlowCover reduces roughly 50% communication cost without loss of accuracy in most cases. Zhiyang Su, Ting Wang 0001, Yu Xia 0001, Mounir Hamdi |
GLOBECOM | 2 |
| 2014 | NovaCube: A low latency Torus-based network architecture for data centersabstractThis paper presents the design, analysis, and implementation of a novel data center network architecture, named NovaCube. Based on regular Torus topology, NovaCube is constructed by adding a number of most beneficial jump-over links, which offers many distinct advantages and practical benefits. Moreover, in order to enable NovaCube to achieve its maximum theoretical performance, a probabilistic oblivious routing algorithm PORA is carefully designed. PORA is a both deadlock and livelock free routing algorithm, which achieves near-optimal performance in terms of average routing path length with better load balancing thus leading to higher throughput. Theoretical derivation and mathematical analysis further prove the good performance of NovaCube and PORA. Ting Wang 0001, Zhiyang Su, Yu Xia 0001, Mounir Hamdi |
GLOBECOM | 1 |
| 2014 | A general framework for performance guaranteed green data center networkingabstractFrom the perspective of resource allocation and routing, this paper aims to save as much energy as possible in data center networks. We present a general framework, based on the blocking island paradigm, to try to maximize the network power conservation and minimize sacrifices of network performance and reliability. The bandwidth allocation mechanism together with power-aware routing algorithm achieve a bandwidth guaranteed tighter network. Besides, our fast efficient heuristics for allocating bandwidth enable the system to scale to large sized data centers. The evaluation result shows that up to more than 50% power savings are feasible while guaranteeing network performance and reliability. Ting Wang 0001, Yu Xia 0001, Jogesh K. Muppala, Mounir Hamdi, Sebti Foufou |
GLOBECOM | 1 |
| 2014 | Fine-grained power control for combined input-crosspoint queued switchesabstractReducing the power consumption of packet switches is becoming increasingly significant to future networks. However, previous research all focused on reducing power in crossbar-based switches, which is either complex or not effective, especially in some extreme cases. This paper proposes to leverage the dynamic voltage and frequency scaling (DVFS) technique in the buffered crossbar-based switches, which is more flexible and simple. The basic idea is to decrease the working frequencies of the crosspoint buffers while still preserving the maximum throughput and the satisfactory delay. Traffic estimators are used at the input and output ports to estimate the traffic arrival rates, based on which the power controller can adjust the working frequencies of the crosspoint buffers at a fine-grained level. Simulation results show that the scheme is effective. Yu Xia 0001, Ting Wang 0001, Zhiyang Su, Mounir Hamdi |
GLOBECOM | 2 |
| 2014 | CheetahFlow: Towards low latency software-defined networkabstractSoftware defined networking (SDN), which enables programmability, has the advantage of global visibility and high flexibility. However, when forwarding new flows in SDN, the interaction between the switch and the controller imposes extra latency such as round-trip time and routing path search time. Even though such latency is acceptable for elephant flows since it only takes limited ratio of total transmission time of elephant flows, it is an overkill to pay certain overheads for mice flows due to the their short transmission time. Moreover, the controller is frequently invoked by the mice flows since the number of mice flows accounts for a large portion of the total number of flows. Hence, the frequent controller invocation is mainly responsible for the controller performance degradation, and thus increasing the flow setup latency significantly. To solve this problem, we propose CheetahFlow, a novel scheme to predict frequent communication pairs via support vector machine and proactively setup wildcard rules to reduce flow setup latency. Particularly, in order to avoid congestion along a fixed path, elephant flows are detected and rerouted to the non-congestion path efficiently by applying blocking island paradigm. Extensive experiments show that CheetahFlow prominently reduces latency without any loss of flexibility of SDN. Zhiyang Su, Ting Wang 0001, Yu Xia 0001, Mounir Hamdi |
ICC | 2 |
| 2014 | SprintNet: A high performance server-centric network architecture for data centersabstractThis paper presents the design, implementation and evaluation of SprintNet, a novel network architecture for data centers. SprintNet achieves high performance in network capacity, fault tolerance, and network latency. SprintNet is also a scalable, yet low-diameter network architecture where the maximum shortest distance between any pair of servers can be limited by no more than four and is independent of the number of layers. The specially designed routing schemes for SprintNet strengthen its merits. Both theoretical analysis and simulation experiments are conducted to evaluate its overall performance with respect to average path length, aggregate bottleneck throughput, and fault tolerance. Ting Wang 0001, Zhiyang Su, Yu Xia 0001, Yang Liu 0081, Jogesh K. Muppala, Mounir Hamdi |
ICC | 1 |
| 2014 | Improving the efficiency of server-centric data center network architecturesabstractData center network architecture is regarded as one of the most important determinants of network performance. As the most typical representatives of architecture design, the server-centric scheme stands out due to its good performance in various aspects. However, there still exist some critical shortcomings in these server-centric architectures. In order to provide an efficient solution to these shortcomings and improve the efficiency of server-centric architectures, in this paper, we propose a hardware based approach, named “Forwarding Unit”. Furthermore, we put forward a traffic aware routing scheme for FlatNet to further evaluate the feasibility and efficiency of our approach. Both theoretical analysis and simulation experiments are conducted to measure its overall performance with respect to cost-effectiveness, fault-tolerance, system latency, packet loss ratio, aggregate bottleneck throughput, and average path length. Ting Wang 0001, Yu Xia 0001, Dong Lin, Mounir Hamdi |
ICC | 1 |