EDBT 2026 Demo / reviewers in the wild / expert
Huaxi Gu
dblp:49/1366
· DBLP profile ↗
69ranked-venue papers
5as first author
43since 2021 · last 2026
0000-0002-6409-2229ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 3 first-author · 20 since 2021Computer networks · 22 · 17 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pulse: Training Acceleration for Large Diffusion Models with Automatic Pipeline Parallelism
Boran Sun, Guoyong Jiang, Yuechen Tao, Zhishu Che, Jieling Yu, Shan Chang, Huaxi Gu, Fangming Liu |
ICDCS | 9 |
| 2026 | FlexLoop: A Distributed Scheduling Strategy for AWGR-Based Optical NetworksabstractOptical networks based on arrayed waveguide grating routers (AWGRs) offer significant advantages, including low latency, high bandwidth, and low power consumption. However, addressing the contention issues associated with AWGRs has become a primary concern in practical deployment. Previous research has focused on centralized control methods to address this challenge, but their scalability is often limited. Additionally, distributed time-slot strategies have been implemented, but fixed time-slot configurations can result in performance loss. To effectively overcome these challenges, we propose a distributed scheduling strategy called FlexLoop. Authorizations are issued by receiving nodes and passed between sending nodes via the software-defined networking (SDN) switches. The authorization holder is permitted to transmit a batch of data to the designated node. Each receiving node has the autonomy to define the authorization passing rules through SDN switch configurations. Similarly, each sending node retains the flexibility to either use the authorization to transmit data or forward it to other nodes. Therefore, FlexLoop operates in a fully distributed manner and shows adaptability to various traffic patterns. Experimental results demonstrate that FlexLoop reduces packet latency by up to 92.01% and improves throughput by up to 77.09% compared to NegotiaToR. Additionally, a prototype system was developed, which validates the practical feasibility of implementing FlexLoop. Zhaoxing Zhou, Huaxi Gu, Xiaoshan Yu 0001, Yunhao Wang 0001 |
IEEE Trans. Computers | 2 |
| 2026 | RCC: Rate-Based Congestion Control for the Lossless Network
Xiaoshan Yu 0001, Huaxi Gu |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2026 | DATCP: A Dynamic and Adaptive Congestion Control Protocol for Reconfigurable Data Center NetworksabstractWith the rapid growth of cloud computing and distributed machine learning, traditional static networks increasingly struggle to meet dynamic communication demands. Optical Circuit Switching brings high bandwidth, low latency, and reconfigurability to data center networks, but the resulting reconfigurable data center networks (RDCNs) introduce dynamic path changes and heterogeneous paths, challenging conventional congestion control protocols. To address this, we propose DATCP, a dynamic and adaptive congestion control protocol designed for RDCNs. DATCP features three key mechanisms: (1) aTTL-based path classification method for accurate and lightweight identification of network types; (2) a receiver-side flow tracking model that detects incast congestion and applies weighted rate control; and (3) an adaptive congestion window algorithm that enables fast convergence and efficient bandwidth utilization. Experiments under synthetic and real-world workloads show that DATCP outperforms existing TCP variants, achieving higher throughput, reducing flow completion time and 99th percentile tail latency by up to 37.8% and 21.1%, respectively, and improving flow fairness. Hong Zou, Huaxi Gu, Xiaoshan Yu 0001, Boxuan Liu |
IEEE Trans. Netw. | 2 |
| 2025 | A survey on vertical interconnection and topology of three-dimensional network-on-chip
Zewei Jing, Qinghai Yang, Nan Cheng 0001, Huaxi Gu, Kyung Sup Kwak |
Integr. | 5 |
| 2025 | Asy-MSFL: Communication-Efficient Federated Learning With Multiserver Adaptive UpdatesabstractFederated Learning (FL) facilitates collaborative model training across distributed data sources while preserving data privacy, yet traditional FL frameworks often encounter severe challenges in communication efficiency and stable model convergence, especially in large-scale, heterogeneous settings. To address these limitations, we propose Asynchronous Multi-Server Federated Learning (Asy-MSFL), a novel FL framework designed to address communication bottlenecks and improve model convergence in large-scale heterogeneous environments. Unlike traditional client-based layered FL architectures, Asy-MSFL adopts layer-specific management, where each server processes updates from specific layers of client models, reducing communication latency and bandwidth consumption. To effectively manage asynchronous updates and client heterogeneity, Asy-MSFL incorporates a dynamic Elastic Averaging Stochastic Gradient Descent (EASGD) mechanism. By dynamically adjusting the elastic force coefficient, it aligns local updates with global objectives, ensuring consistent and stable convergence despite varying client conditions. Experimental results demonstrate that Asy-MSFL achieves substantial communication cost reductions and high model accuracy, making it an advanced solution for scalable, efficient federated learning with reliable convergence. Huaxi Gu |
IEEE Internet Things J. | 2 |
| 2025 | A survey on routing algorithm and router microarchitecture of three-dimensional Network-on-Chip
Zewei Jing, Qinghai Yang, Nan Cheng 0001, Huaxi Gu, Kyung Sup Kwak |
J. Syst. Archit. | 5 |
| 2025 | SROdcn: Scalable and Reconfigurable Optical DCN Architecture for High-Performance ComputingabstractData Center Network (DCN) flexibility is critical for providing adaptive and dynamic bandwidth while optimizing network resources to manage variable traffic patterns generated by heterogeneous applications. To provide flexible bandwidth, this work proposes a machine learning approach with a new Scalable and Reconfigurable Optical DCN (SROdcn) architecture that maintains dynamic and non-uniform network traffic according to the scale of the high-performance optical interconnected DCN. Our main device is the Fiber Optical Switch (FOS), which offers competitive wavelength resolution. We propose a new top-of-rack (ToR) switch that utilizes Wavelength Selective Switches (WSS) to investigate Software-Defined Networking (SDN) with machine learning-enabled flow prediction for reconfigurable optical Data Center Networks (DCNs). Our architecture provides highly scalable and flexible bandwidth allocation. Results from Mininet experimental simulations demonstrate that under the management of an SDN controller, machine learning traffic flow prediction and graph connectivity allow each optical bandwidth to be automatically reconfigured according to variable traffic patterns. The average server-to-server packet delay performance of the reconfigurable SROdcn improves by 42.33% compared to inflexible interconnects. Furthermore, the network performance of flexible SROdcn servers shows up to a 49.67% latency improvement over the Passive Optical Data Center Architecture (PODCA), a 16.87% latency improvement over the optical OPSquare DCN, and up to a 71.13% latency improvement over the fat-tree network. Additionally, our optimized Unsupervised Machine Learning (ML-UnS) method for SROdcn outperforms Supervised Machine Learning (ML-S) and Deep Learning (DL). Kassahun Geresu, Huaxi Gu, Xiaoshan Yu 0001, Meaad Fadhel, Hui Tian 0001, Wenting Wei |
IEEE Trans. Cloud Comput. | 2 |
| 2025 | Iris: Toward Intelligent Reliable Routing for Software-Defined Satellite NetworksabstractSatellite networks have long been regarded as a vital component of space communication systems, which provide integrated satellite-terrestrial broadband access in seamless coverage and cost-effective manner. The inter-satellite routing design for low earth orbit (LEO) satellite constellations is critical for achieving low-latency and high-reliability communication in the space communication systems. However, the inherent dynamic nature of LEO satellites, coupled with the variability in inter-satellite connectivity, imposes significant challenges for routing efficiency and network dependability. Existing routing schemes cannot handle such topological fluctuations due to their insensitivity to real-time network changes, thus suffering from performance degradations in highly dynamic space environments. This paper presents Iris, an intelligent reliable routing scheme for inter-satellite communication, aiming at increasing efficiency and reliability of the packet transmission process. Specifically, we propose a comprehensive deep reinforcement learning (DRL) framework that learns a policy to select routing paths automatically under the emerging software-defined satellite networking (SDSN) architecture. To strengthen fault-tolerance in fluctuating environments, we train an agent in an incremental manner by gradually increasing scenario complexity. Simulation results indicate that our solution significantly outperforms baselines and exhibits advances in adaptability and reliability, especially under dynamic environments with frequent topology changes. Wenting Wei, Liying Fu, Huaxi Gu, Xueyu Lu, Lei Liu 0031, Shahid Mumtaz, Mohsen Guizani |
IEEE Trans. Commun. | 3 |
| 2025 | Grace: Toward Routing in Dynamic Network Environments With Graph EmbeddingabstractRecent efforts have explored adaptive routing via deep reinforcement learning (DRL) techniques without handcrafted parameter engineering. Intrinsically, routing decision-making is essentially a process used to find a subgraph in a graph-structured network. However, previous works seldom took topological relationships into consideration when providing adaptive routing algorithms, causing them to suffer from suboptimal routes in dynamic network environments involving both varying traffic loads and burst traffic. In this paper, we presentGrace, a novel graph embedding-based Deep Reinforcement Learning framework tailored for distributed routing algorithm optimization within the Software-Defined Networking (SDN) paradigm. Specifically,Graceleverages graph embedding to translate graph-structured entities into low-dimensional vectors, thereby enabling multiple DRL agents to learn optimal routing paths under dynamic network environments. Unfortunately, training multiple agents encounters inherent challenges in complicated and dynamic network scenarios. In response, we design an adaptive incremental training method forGracethat makes the model adapt to task complexity in a gradual manner, while speeding up its retraining efforts when environments change. To further accelerate convergence, we integrate intrinsic curiosity intoGraceto tackle large environments with sparse rewards. Extensive experiments conducted on two real-world topologies demonstrate the rationality and effectiveness ofGrace, and the results show throughput improvements of up to 40.1% compared to other state-of-the-art DRL routing algorithms under bursty traffic conditions. Wenting Wei, Huaxi Gu, Liying Fu, Baochun Li |
IEEE Trans. Netw. | 2 |
| 2025 | Halo: An Efficient and Scalable Distributed Control Strategy for AWGR-Based Optical NetworksabstractOptical networks based on arrayed waveguide grating routers (AWGRs) present a promising solution for rapidly evolving data centers (DCs). However, the contention issue remains a significant challenge to their widespread adoption. Existing control strategies typically rely on time-slots that are both synchronized across the network and of fixed duration. This results in the fixed switching granularity. The determination of time-slot duration often involves a trade-off between flexibility and control overhead. In this paper, we propose Halo, a distributed control strategy that avoids contention without relying on time-slots. Halo decentralizes the scheduling tasks across the network. Each node independently issuesGrants, which are passed sequentially among source nodes with communication demands. The node holding theGrantis allowed to transmit a batch of data to the issuing node. Once the demand is no longer present, or the transmitted packet count exceeds a threshold, theGrantis forwarded to the next node. As a result, synchronization operations among nodes is unnecessary, allowing the switching granularity to become flexible. Experimental results demonstrate that Halo significantly outperforms existing control strategies. In trace-based simulations, it achieves up to a 75.13% reduction in flow completion time (FCT). Additionally, we developed a prototype system and executed several real applications to demonstrate the practical feasibility of the Halo. Zhaoxing Zhou, Huaxi Gu, Xiaoshan Yu 0001, Yunhao Wang 0001 |
IEEE Trans. Netw. | 2 |
| 2025 | Energy Efficient and Multi-Resource Optimization for Virtual Machine Placement by Improving MOEA/DabstractThe explosive growth of cloud services has led to the widespread construction of large-scale data centers to meet diverse and multifaceted cloud computing demands. However, this expansion has resulted in substantial energy consumption. Virtual machine placement (VMP) has been extensively studied as a means to provide flexible and scalable cloud services while optimizing energy efficiency. Yet, the increasing complexity and diversity of applications have posted VMP suffering from waste of resources and bottlenecks due to unbalanced utilization of multi-dimensional resources. To address these issues, this article proposes a bi-objective optimization model for VMP that jointly optimizes power consumption and multi-dimensional resource utilization. Solving this large-scale bi-objective model presents a significant challenge in balancing performance and computational complexity. To tackle this, an enhanced decomposition-based multi-objective evolutionary algorithm (MOEA/D) based on$\varepsilon$-domination, termed$\varepsilon$-IMOEA/D-M2M is designed to provide solutions for the proposed optimization. Compared with both heuristics and evolutionary algorithms, performance evaluations demonstrate that our proposed VMP algorithm effectively reduces power consumption and balances multidimensional resource utilization while significantly decreasing running time compared to both heuristic and traditional evolutionary algorithms. Wenting Wei, Huaxi Gu |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | Satformer: Accurate and Robust Traffic Data Estimation for Satellite NetworksabstractThe operations and maintenance of satellite networks heavily depend on traffic measurements. Due to the large-scale and highly dynamic nature of satellite networks, global measurement encounters significant challenges in terms of complexity and overhead. Estimating global network traffic data from partial traffic measurements is a promising solution. However, the majority of current estimation methods concentrate on low-rank linear decomposition, which is unable to accurately estimate. The reason lies in its inability to capture the intricate nonlinear spatio-temporal relationship found in large-scale, highly dynamic traffic data. This paper proposes Satformer, an accurate and robust method for estimating traffic data in satellite networks. In Satformer, we innovatively incorporate an adaptive sparse spatio-temporal attention mechanism. In the mechanism, more attention is paid to specific local regions of the input tensor to improve the model's sensitivity on details and patterns. This method enhances its capability to capture nonlinear spatio-temporal relationships. Experiments on small, medium, and large-scale satellite networks datasets demonstrate that Satformer outperforms mathematical and neural baseline methods notably. It provides substantial improvements in reducing errors and maintaining robustness, especially for larger networks. The approach shows promise for deployment in actual systems. Wenting Wei, Chengbin Liang, Huaxi Gu |
NeurIPS | 5 |
| 2024 | Spatio-temporal communication network traffic prediction method based on graph neural network
Huaxi Gu, Wenting Wei, Zexu Lin, Ning Wang 0001 |
Inf. Sci. | 2 |
| 2024 | Deep Reinforcement Learning Based Dynamic Flowlet Switching for DCNabstractFlowlet switching has been proven to be an effective technology for fine-grained load balancing in data center networks. However, flowlet detection based on static flowlet timeout values, lacks accuracy and effectiveness in complex network environments. In this paper, we propose a new deep reinforcement learning approach, called DRLet, to dynamically detect flowlets. DRLet offers two advantages: first, it provides dynamic flowlet timeout values to detect bursts into fine-grained flowlets; second, flowlet timeout values are automatically configured by the deep reinforcement learning agent, which only requires simple and measurable network states, instead of any prior knowledge, to achieve the pre-defined goal. With our approach, the flowlet timeout value dynamically matches the network load scenario, ensuring the accuracy and effectiveness of flowlet detection while suppressing packet reordering. Our results show that DRLet achieves superior performance compared to existing schemes based on static flowlet timeout values in both baseline and asymmetric topologies. Xinglong Diao, Huaxi Gu, Wenting Wei, Guoyong Jiang, Baochun Li |
IEEE Trans. Cloud Comput. | 2 |
| 2024 | COCSN: A Multi-Tiered Cascaded Optical Circuit Switching Network for Data CenterabstractA cascaded network represents a classic scaling-out model in traditional electrical switching networks. Recent proposals have integrated optical circuit switching at specific tiers of these networks to reduce power consumption and enhance topological flexibility. Utilizing a multi-tiered cascaded optical circuit switching network is expected to extend the advantages of optical circuit switching further. The main challenges fall into two categories. First, an architecture with sufficient connectivity is required to support varying workloads. Second, the network reconfiguration is more complex and necessitates a low-complexity scheduling algorithm. In this work, we propose COCSN, a multi-tiered cascaded optical circuit switching network architecture for data center. COCSN employs wavelength-selective switches that integrate multiple wavelengths to enhance network connectivity. We formulate a mathematical model covering lightpath establishment, network reconfiguration, and reconfiguration goals, and propose theorems to optimize the model. Based on the theorems, we introduce an over-subscription-supported wavelength-by-wavelength scheduling algorithm, facilitating agile establishment of lightpaths in COCSN tailored to communication demand. This algorithm effectively addresses scheduling complexities and mitigates the issue of lengthy WSS configuration times. Simulation studies investigate the impact of flow length, WSS reconfiguration time, and communication domain on COCSN, verifying its significantly lower complexity and superior performance over classical cascaded networks. Huaxi Gu, Xiaoshan Yu 0001, Songyan Wang, Zeshan Chang |
IEEE Trans. Cloud Comput. | 2 |
| 2024 | ICLB: intelligent controllers load balancing for software-defined based optical data center networks
Kassahun Geresu, Huaxi Gu, Meaad Fadhel, Wenting Wei, Xiaoshan Yu 0001 |
J. Supercomput. | 2 |
| 2024 | FlowStar: Fast Convergence Per-Flow State Accurate Congestion Control for InfiniBandabstractAccording to the latest TOP500 list, InfiniBand (IB) is the most widely used network architecture in the top 10 supercomputers. IB relies on Credit-based Flow Control (CBFC) to provide a lossless network and InfiniBand congestion control (IB CC) to relieve congestion, however, this can lead to the problem of victim flow since messages are mixed in the same queue and long-lived congestion spreading due to slow convergence. To deal with these problems, in this paper, we propose FlowStar, a fast convergence per-flow state accurate congestion control for InfiniBand. FlowStar includes two core mechanisms: 1) optimized per-flow CBFC mechanism provides flow state control to detect real congestion; and 2) rate adjustment rules make up for the mismatch between the original IB CC rate regulation and the per-hop CBFC to alleviate congestion spreading. FlowStar implements a per-flow congestion state on switches and can obtain in-flight packet information without additional parameter settings to ensure a lossless network. Evaluations show that FlowStar improves average and tail message complete time under different workloads. Changyun Luo, Huaxi Gu, Lijing Zhu, Huixia Zhang |
IEEE/ACM Trans. Netw. | 2 |
| 2024 | A 3D Hybrid Optical-Electrical NoC Using Novel Mapping Strategy Based DCNN Dataflow AccelerationabstractA large number of multiply-accumulate operations and memory accesses required in deep convolutional neural networks (DCNN) leads to high latency and energy consumption (EC), that hinder their further applications. Dataflow-based acceleration schemes reduce memory accesses by leveraging reusable data in DCNNs. Row Stationary (RS) dataflow is a more advanced dataflow. In the convolutional layer acceleration of RS dataflow, the flexibility of mapping from logical processing element (LPE) sets to physical PE sets is relatively poor. The utilization of processing elements (PEs) is low. In this paper, a novel mapping strategy based on genetic algorithm (GAMS) with the goal of optimizing EC is proposed. GAMS is designed to address the energy inefficiencies faced when mapping RS dataflow. A 3D hybrid optical-electrical Network-on-Chip (3DHOENoC) is proposed to further improve the communication efficiency, energy efficiency and the processing speed of DCNN. Simulation and evaluation results show that GAMS can achieve better mapping flexibility, higher PEs utilization and 15.9% improvement of execution speed on average. In addition, the execution time (ET) performance of processing the DCNN can be further improved by adopting the 3DHOENoC architecture with better communication parallelism. Bowen Zhang 0004, Huaxi Gu, Grace Li Zhang, Yintang Yang, Ziteng Ma, Ulf Schlichtmann |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | Deadline Enables In-Order Flowlet Switching for Load BalancingabstractFine granularity can greatly enhance load balancing opportunities, but packet reordering is still a challenge. In this paper, we propose EDFLet, a flowlet switching mechanism that uses deadlines to achieve in-order flowlet-level load balancing. We assign a deadline value to each packet of a flowlet based on its burst interval, which ensures that an earlier flowlet completes transmission before the next flowlet of the same flow. We also apply Earliest Deadline First scheduling at the switch, which guarantees that packets that packets meet their deadlines and arrive in order at the receiver. Our experimental results show that EDFLet performs better than existing methods in both symmetric and asymmetric topologies. Xinglong Diao, Wenting Wei, Huaxi Gu |
APNet | 3 |
| 2023 | Class-based Quantization for Neural NetworksabstractIn deep neural networks (DNNs), there are a huge number of weights and multiply-and-accumulate (MAC) operations. Accordingly, it is challenging to apply DNNs on resource- constrained platforms, e.g., mobile phones. Quantization is a method to reduce the size and the computational complexity of DNNs. Existing quantization methods either require hardware overhead to achieve a non-uniform quantization or focus on model-wise and layer-wise uniform quantization, which are not as fine-grained as filter-wise quantization. In this paper, we propose a class-based quantization method to determine the minimum number of quantization bits for each filter or neuron in DNNs individually. In the proposed method, the importance score of each filter or neuron with respect to the number of classes in the dataset is first evaluated. The larger the score is, the more important the filter or neuron is and thus the larger the number of quantization bits should be. Afterwards, a search algorithm is adopted to exploit the different importance of filters and neurons to determine the number of quantization bits of each filter or neuron. Experimental results demonstrate that the proposed method can maintain the inference accuracy with low bit-width quantization. Given the same number of quantization bits, the proposed method can also achieve a better inference accuracy than the existing methods. Grace Li Zhang, Huaxi Gu, Bing Li 0005, Ulf Schlichtmann |
DATE | 3 |
| 2023 | SteppingNet: A Stepping Neural Network with Incremental Accuracy EnhancementabstractDeep neural networks (DNNs) have successfully been applied in many fields in the past decades. However, the in-creasing number of multiply-and-accumulate (MAC) operations in DNNs prevents their application in resource-constrained and resource-varying platforms, e.g., mobile phones and autonomous vehicles. In such platforms, neural networks need to provide ac-ceptable results quickly and the accuracy of the results should be able to be enhanced dynamically according to the computational resources available in the computing system. To address these challenges, we propose a design framework called SteppingNet. SteppingNet constructs a series of sub nets whose accuracy is incrementally enhanced as more MAC operations become avail-able. Therefore, this design allows a trade-off between accuracy and latency. In addition, the larger sub nets in SteppingNet are built upon smaller subnets, so that the results of the latter can directly be reused in the former without recomputation. This property allows SteppingNet to decide on-the-fly whether to enhance the inference accuracy by executing further MAC operations. Experimental results demonstrate that SteppingNet provides an effective incremental accuracy improvement and its inference accuracy consistently outperforms the state-of-the-art work under the same limit of computational resources. Grace Li Zhang, Xunzhao Yin, Cheng Zhuo, Huaxi Gu, Bing Li 0005, Ulf Schlichtmann |
DATE | 5 |
| 2023 | Routing and Wavelength Assignment for Multiple Multicasts in Optical Network-on-Chip (ONoC)abstractOptical network-on-chip (ONoC) is an emerging chip-scale optical interconnection technology to realize high-performance and power-efficient intercore communication for many-core processors. Multicast communication is popularly used in parallel applications on chip. However, existing researches for multicast in ONoC mainly focus on the optimization of one multicast. This limits the practical applications of the research outcomes because we often face the dynamic formation of multiple multicast groups in real network systems. In this article, we define the problem of routing and wavelength assignment for multiple multicasts in ONoC with the objective of minimizing the number of wavelengths required. To solve the problem, we first formulate it as an integer programming model for general topologies. Then we design routing policies for special instances that optimally use only one wavelength on mesh topology. For general instances, we design a group-partitioning routing algorithm for multiple multicasts (GPRMM). GPRMM decouples a group of multicasts into a number of subgroups, each of which matching one of the special instances. Theoretical results show that the number of wavelengths required by GPRMM is no more than the Destination Density$\sigma _{d}$, i.e., the maximum number of multicasts with destinations in the same row or column. Moreover, we find the upper bound and the lower bound on the number of wavelengths required for GPRMM. The wavelength requirement is also upper bounded by the network size$n$for an$n\times n$mesh network. Simulation results show that GPRMM can reduce the number of wavelengths by 26.7% compared with previous methods. GPRMM has the advantages of low routing complexity, low wavelength requirement, low power consumption, and good scalability. Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001, Huaxi Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | RibsNet: A Scalable, High-Performance, and Cost-Effective Two-Layer-Based Cloud Data Center Network ArchitectureabstractToday, we witness an unprecedented surge of data traffic due to the explosive growth of smart devices and intelligent applications. Hence, supporting the Data Center Networks (DCNs) infrastructure by accommodating myriads of servers becomes natural and inevitable. This vision can be realized by providing an interconnecting architecture with higher scalability and remarkable capacity while maintaining fault tolerance, minimum latency, and cost and power efficiency. Motivated by this trend, this paper proposes a new network based on dual-centric dubbedRibsNet, which is a symmetric network composed of two layers of switches and dual-port servers. The proposed network enables the servers to establish all kinds of connections (server-switch and server-server), thus enhancing network fault tolerance with high bisection width and maximizing its performance and reliability.RibsNetallows the gradual addition of servers while preserving all its topological properties; it has a stable and low diameter, which does not exceed six regardless of the DCN’s size. Furthermore, an efficient and fault tolerant routing scheme, exclusively designed for this network, is proposed to enhance the routing efficiency and address different types of failures. Theoretical analysis proves thatRibsNetstrikes a good balance among several connections within the network and brings outstanding gains in terms of incremental scalability, bisection width, and cost and power savings against various cutting-edge DCNs architectures. Simulation results demonstrateRibsNetas a promising structure that outperformsBCube,DCell,HSDC, andFiConnin terms of latency and throughput by approximately 8%, 6%, 7%, and 26%, respectively. Moeen Al-Makhlafi, Huaxi Gu, Ahlam Almuaalemi, Eiad Almekhlafi, Musbahu M. Adam |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2023 | GRL-PS: Graph Embedding-Based DRL Approach for Adaptive Path SelectionabstractForwarding path selection for data traffic is one of the most fundamental operations in computer networks, whose performance drastically impacts both transmission efficiency and reliability in network domains. Although deep reinforcement learning (DRL) has attracted considerable attention for path selection instead of hand-tuned heuristics, few works have considered how to exploit graph-structured information in networks to improve routing and forwarding efficiency. In fact, generating routes is essentially a process for finding a subgraph in a graph-structured network. To this end, this paper proposes an effective and novel graph embedding-based DRL framework for adaptive path selection (termed GRL-PS), aiming at reducing end-to-end (E2E) latency and promoting network throughput while maintaining stability in dynamically changing environments. Specifically, graph representation learning (GRL) is deployed as an effective enabler for the DRL agent to learn the relational knowledge of interacting entities for route decisions in networks. However, training such an agent in a dynamically changing environment encounters a knowledge acquisition bottleneck, since the DRL agent is always forced to learn every task from scratch. To improve the adaptation of behaviors and acquire skills beyond what the source policy can teach, we introduce potential-based reward shaping as a means of knowledge transfer to guide the agent in unfamiliar conditions with sparse rewards. Experimental results show that compared with baseline methods, our solution can achieve nearly-optimal performance with both latency and throughput, especially in large-scale dynamic networks. Wenting Wei, Liying Fu, Huaxi Gu, Yan Zhang 0002, Chao Wang 0028, Ning Wang 0001 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2023 | Multi-Dimensional Resource Allocation in Distributed Data Centers Using Deep Reinforcement LearningabstractWith the development of edge-cloud computing technologies, distributed data centers (DCs) have been extensively deployed across the global Internet. Since different users/applications have heterogeneous requirements on specific types of ICT resources in distributed DCs, how to optimize such heterogeneous resources under dynamic and even uncertain environments becomes a challenging issue. Traditional approaches are not able to provide effective solutions for multi-dimensional resource allocation that involves the balanced utilization across different resource types in distributed DC environments. This paper presents a reinforcement learning based approach for multi-dimensional resource allocation (termed as NESRL-MRM) that is able to achieve balanced utilization and availability of resources in dynamic environments. To train NESRL-MRM’s agent with sufficiently quick wall-clock time but without the loss of exploration diversity in the search space, a natural evolution strategy (NES) is employed to approximate the gradient of the reward function. To realistically evaluate the performance of NESRL-MRM, our simulation evaluations are based on real-world workload traces from Amazon EC2 and Google datacenters. Our results show that NESRL-MRM is able to achieve significant improvement over the existing approaches in balancing the utilization of multi-dimensional DC resources, which leads to substantially reduced blocking probability of future incoming workload demands. Wenting Wei, Huaxi Gu, Kun Wang 0001, Jianjia Li, Ning Wang 0001 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2022 | Flex: A flowlet-level load balancing based on load-adaptive timeout in DCN
Xinglong Diao, Huaxi Gu, Xiaoshan Yu 0001, Changyun Luo |
Future Gener. Comput. Syst. | 2 |
| 2022 | ABL-TC: A lightweight design for network traffic classification empowered by deep learning
Wenting Wei, Huaxi Gu, Wenshuai Deng, Xinming Ren |
Neurocomputing | 2 |
| 2022 | An Efficient Dataflow Mapping Method for Convolutional Neural Networks
Zhuangzhuang Liu, Huaxi Gu, Bowen Zhang 0004, Canran Shi |
Neural Process. Lett. | 2 |
| 2022 | A Novel CONV Acceleration Strategy Based on Logical PE Set Segmentation for Row Stationary DataflowabstractDeep convolutional neural networks (DCNNs) have been proposed as enhanced developments of neural networks (NNs) in the field of artificial intelligence (AI) and successfully applied in deep learning (DL) scenarios. With the advancement of technology, the number of network layers has continuously increased, resulting in a huge number of calculations and memory accesses required in the training and inference process of DCNNs and thereby hindering their further deployment and application. Using a specific dataflow formed by reusable DCNN data in the network-on-chip (NoC), reducing the memory access pressure and improving DCNN processing efficiency has become a promising acceleration schemes for the current DCNN. In this paper, a novel convolution layer (CONV) acceleration strategy based on logical PE set segmentation for row stationary (RS) dataflow is proposed to solve the problems of low flexibility and inefficient processing array utilization faced by the conventional folding mapping strategy. The simulation results show that the new mapping strategy based on PE set segmentation can achieve better processing element utilization and CONV acceleration improvement at the expense of little increase in the data movement energy consumption compared with the conventional strategy. Bowen Zhang 0004, Huaxi Gu, Kun Wang 0001, Yintang Yang |
IEEE Trans. Computers | 2 |
| 2022 | HPSTOS: High-Performance and Scalable Traffic Optimization Strategy for Mixed Flows in Data Center NetworksabstractIn data center networks, traffic needs to be distributed among different paths using traffic optimization strategies for mixed flows. Most of the existing strategies consider either distributed or centralized mechanisms to optimize the latency of mice flows or the throughput of elephant flows. However, low network performance and scalability issues are intrinsic limitations of both strategies. In addition, the current elephant flow detection methods are inefficient. In this article, we propose a high-performance and scalable traffic optimization strategy (HPSTOS) based on a hybrid approach that leverages the advantages of both centralized and distributed mechanisms. HPSTOS improves the efficiency of elephant flow detection through sampling and flow-table identification. HPSTOS guarantees preferential transmission of mice flows using priority scheduling and adjusts their transmission rate by coding-based congestion control on the end-host, reducing their latency. Additionally, HPSTOS schedules elephant flows by cost-aware dynamic flow scheduling on a centralized controller to improve their throughput. The controller handles only elephant flows, which constitutes the minority of the flows, allowing effective scalability. Evaluations show that HPSTOS outperforms existing schemes by realizing efficient elephant flow detection and improving network performance and scalability. Yong Liu 0045, Huaxi Gu, Ning Wang 0001 |
IEEE Trans. Cloud Comput. | 2 |
| 2022 | Multi-Objective Optimization for Resource Allocation in Vehicular Cloud Computing NetworksabstractModern transportation is associated with considerable challenges related to safety, mobility, the environment and space limitations. Vehicular networks are widely considered to be a promising approach for improving satisfaction and convenience in transportation. However, with the exploding popularity among vehicle users and the growing diverse demands of different services, ensuring the efficient use of resources and meeting the emerging needs remain challenging. In this paper, we focus on resource allocation in vehicular cloud computing (VCC) and fill the gaps in the previous research by optimizing resource allocation from both the provider’s and users’ perspectives. We model this problem as a multi-objective optimization with constraints that aims to maximize the acceptance rate and minimize the provider’s cloud cost. To solve such an NP-hard problem, we improve the nondominated sorting genetic algorithm II (NSGA-II) by modifying the initial population according to the matching factor, dynamic crossover probability and mutation probability to promote excellent individuals and increase population diversity. The simulation results show that our proposed method achieves enhanced performance compared to the previous methods. Wenting Wei, Ruying Yang, Huaxi Gu, Weike Zhao, Chen Chen 0006, Shaohua Wan 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | RSLB: Robust and Scalable Load Balancing in Software-Defined Data Center NetworksabstractData center networks demand high-performance, robust, and scalable load balancing protocols. Despite progress, existing work still cannot meet these requirements well. Software defined networking (SDN) can bring considerable flexibility to the management of data center networks. In the software-defined data center network, we design, analyze, and evaluate RSLB, a robust and scalable load balancing protocol that overcomes these challenges. RSLB uses fine-grained flowcell as the transmission unit, and uses link delay as the congestion metric. It uses a three-step routing strategy to route flowcells to the path with the least congestion. Through global congestion awareness, RSLB reduces flow completion time (FCT), and is more robust to topological asymmetries compared to existing congestion-agnostic schemes. To collect and store congestion information, RSLB adopts a distributed control structure that monitors the congestion of the entire network through multiple controllers, which makes it much more scalable for implementation in large-scale networks compare to existing congestion-aware schemes. The simulation results show that RSLB can achieve lower FCT for mice flows and higher throughput for elephant flows than existing schemes, no matter in failure-free topology or asymmetric topology. Yong Liu 0045, Huaxi Gu, Zhaoxing Zhou, Ning Wang 0001 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2021 | FLAG: Flow Representation Generator based on Self-supervised Learning for Encrypted Traffic ClassificationabstractDue to its excellent ability in learning features from large scale raw data, deep learning (DL) has attracted much attention for encrypted traffic classification. However, most DL-based traffic classifiers usually rely on enormous labeled samples. Motivated by this, we investigate a self-supervised traffic classifier (FLAG) without sacrifice of identification accuracy, only depending on small labeled traffic samples and highly available unlabeled traffic samples. Specifically, focusing on local short-term characteristics of traffic, we design a preprocessing algorithm, termed as N-phrase Extration, to convert unlabeled raw traffic dataset into sequences of high-frequency phrases as input of Bidirectional Encoder. On account of their significance, potential timing characteristics from input sequences are mined by Bidirectional Encoder and embedded into robust representations with distributed vectors to enhance classifier’s performance significantly. Our comprehensive experiments indicate FLAG can achieve 98.65% in 100% of dataset and 98.07% in 10% of dataset in terms of true positive rate in UNB ISCX VPN-nonVPN dataset, which are better than p-FP, FS-Net and Deep Packet. Wenting Wei, Tianjie Ju, Han Liao, Weike Zhao, Huaxi Gu |
APNet | 5 |
| 2021 | Robustness of Neuromorphic Computing with RRAM-based Crossbars and Optical Neural NetworksabstractRRAM-based crossbars and optical neural networks are attractive platforms to accelerate neuromorphic computing. However, both accelerators suffer from hardware uncertainties such as process variations. These uncertainty issues left unaddressed, the inference accuracy of these computing platforms can degrade significantly. In this paper, a statistical training method where weights under process variations and noise are modeled as statistical random variables is presented. To incorporate these statistical weights into training, the computations in neural networks are modified accordingly. For optical neural networks, we modify the cost function during software training to reduce the effects of process variations and thermal imbalance. In addition, the residual effects of process variations are extracted and calibrated in hardware test, and thermal variations on devices are also compensated in advance. Simulation results demonstrate that the inference accuracy can be improved significantly under hardware uncertainties for both platforms. Grace Li Zhang, Bing Li 0005, Ying Zhu 0008, Yiyu Shi 0001, Xunzhao Yin, Cheng Zhuo, Huaxi Gu, Tsung-Yi Ho, Ulf Schlichtmann |
ASP-DAC | 8 |
| 2021 | Hardware-Software Codesign of Weight Reshaping and Systolic Array Multiplexing for Efficient CNNsabstractThe last decade has witnessed the breakthrough of deep neural networks (DNNs) in various fields, e.g., image/speech recognition. With the increasing depth of DNNs, the number of multiply-accumulate operations (MAC) with weights explodes significantly, preventing their applications in resource-constrained platforms. The existing weight pruning method is considered to be an effective method to compress neural networks for acceleration. However, weights after pruning usually exhibit irregular patterns. Implementing MAC operations with such irregular weight patterns on hardware platforms with regular designs, e.g., GPUs and systolic arrays, might result in an underutilization of hardware resources. To utilize the hardware resource efficiently, in this paper, we propose a hardware-software codesign framework for acceleration on systolic arrays. First, weights after unstructured pruning are reorganized into a dense cluster. Second, various blocks are selected to cover the cluster seamlessly. To support the concurrent computations of such blocks on systolic arrays, a multiplexing technique and the corresponding systolic architecture is developed for various CNNs. The experimental results demonstrate that the performance of CNN inferences can be improved significantly without accuracy loss. Jingyao Zhang 0002, Huaxi Gu, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann |
DATE | 2 |
| 2021 | Tree-RNN: Tree structural recurrent neural network for network traffic classification
Xinming Ren, Huaxi Gu, Wenting Wei |
Expert Syst. Appl. | 2 |
| 2021 | An efficient shortest path algorithm for content-based routing on 2-D mesh accelerator networks
Huaxi Gu, Wenting Wei, Yawen Chen 0001 |
Future Gener. Comput. Syst. | 2 |
| 2021 | Energy-Efficient UAV-Enabled Data Collection via Wireless Charging: A Reinforcement Learning ApproachabstractIn this article, we study the application of unmanned aerial vehicle (UAV) for data collection with wireless charging, which is crucial for providing seamless coverage and improving system performance in the next-generation wireless networks. To this end, we propose a reinforcement learning-based approach to plan the route of UAV to collect sensor data from sensor devices scattered in the physical environment. Specifically, the physical environment is divided into multiple grids, where one spot for UAV hovering as well as the wireless charging of UAV is located at the center of each grid. Each grid has a spot for the UAV to hover, and moreover, there is a wireless charger at the center of each grid, which can provide wireless charging to UAV when it is hovering in the grid. When the UAV lacks energy, it can be charged by the wireless charger at the spot. By taking into account the collected data amount as well as the energy consumption, we formulate the problem of data collection with UAV as a Markov decision problem, and exploit Q-learning to find the optimal policy. In particular, we design the reward function considering the energy efficiency of UAV flight and data collection, based on which Q-table is updated for guiding the route of UAV. Through extensive simulation results, we verify that our proposed reward function can achieve a better performance in terms of the average throughput, delay of data collection, as well as the energy efficiency of UAV, in comparison with the conventional capacity-based reward function. Shu Fu, Yujie Tang 0001, Yuan Wu 0001, Ning Zhang 0007, Huaxi Gu, Chen Chen 0037 |
IEEE Internet Things J. | 5 |
| 2021 | Network Service Chaining and Embedding With Provable BoundsabstractNetwork function virtualization (NFV) is introduced to effectively deliver end-to-end network services for the emerging Internet of Things (IoT), multiaccess edge computing, and 5G communication techniques. In NFV, the network service request can be accommodated in the form of a service function chain (SFC). The SFC will have to reserve abundant resources, such as link bandwidth, service functions, and computation in the physical network to meet the demands of customers. Minimizing the cost from the resource reservation in NFV remains challenging, even though a few works in the literature proposed cost-optimization methodologies with assumptions to guarantee their correctness. In this article, we comprehensively investigate how to minimize the cost when delivering network services as SFCs with provable bounds and fewer assumptions. We formally define the problem of minimum cost service function chaining and embedding (MC-SFCE) and propose an algorithm, namely, cost factor-based SFCE optimization with shortcut (COFO-SC), for MC-SFCE. Novel mathematical analysis is provided to demonstrate the correctness of our approaches and related bounds. Our extensive simulations and analysis also show that the proposed COFO-SC outperforms the schemes directly extended from the existing work. Danyang Zheng 0001, Huaxi Gu, Wenting Wei, Chengzong Peng, Xiaojun Cao |
IEEE Internet Things J. | 2 |
| 2021 | Highly-Efficient Switch Migration for Controller Load Balancing in Elastic Optical Inter-Datacenter NetworksabstractIn elastic optical inter-datacenter networks, multiple software-defined networking (SDN) controllers have been used to improve scalability and reliability. The static controller deployment faces the problem of controller load imbalance under dynamic changes in network traffic. In response to this problem, existing research works proposed to use dynamic switch migration to achieve the controller load balancing. However, the efficiency of the existing switch migration is low, because the load balancing performance of controllers does not improve significantly after the switch migration and the migration activities can also cause additional migration costs. Thus the efficiency of switch migration has not yet been well resolved. In this work, we propose a highly-efficient switch migration (HESM) for controller load balancing. The proposed HESM method defines multiple load metrics to measure the load of controllers, and selects the optimal target controller with the largest remaining resource, which improves the load balancing performance. HESM selects switches based on minimizing migration costs, reducing additional migration costs. In addition, HESM can handle the load of multiple controllers in parallel during a switch migration activity, which improves the efficiency of migration. The simulation results show that HESM significantly improves load balancing performance of controllers and reduces the migration cost compared to existing solutions. Yong Liu 0038, Huaxi Gu, Fulong Yan, Nicola Calabretta |
IEEE J. Sel. Areas Commun. | 2 |
| 2021 | An adaptive failure recovery mechanism based on asymmetric routing for data center networks
Yong Liu 0038, Huaxi Gu, Kun Wang 0001, Xiaoshan Yu 0001, Yunhao Wang 0001 |
J. Supercomput. | 2 |
| 2021 | Universal Method for Constructing Fault-Tolerant Optical Routers Using RRWabstractHigh‐speed data transmission enabled by photonic network‐on‐chip (PNoC) has been regarded as a significant technology to overcome the power and bandwidth constraints of electrical network‐on‐Chip (ENoC). This has given rise to an exciting new research area, which has piqued the public’s attention. Current on‐chip architectures cannot guarantee the reliability of PNoC, due to component failures or breakdowns occurring, mainly, in active components such as optical routers (ORs). When such faults manifest, the optical router will not function properly, and the whole network will ultimately collapse. Moreover, essential phenomena such as insertion loss, crosstalk noise, and optical signal‐to‐noise ratio (OSNR) must be considered to provide fault‐tolerant PNoC architectures with low‐power consumption. The main purpose of this manuscript is to improve the reliability of PNoCs without exposing the network to further blocking or contention by taking the effect of backup paths on signals sent over the default paths into consideration. Thus, we propose a universal method that can be applied to any optical router in order to increase the reliability by using a reliable ring waveguide (RRW) to provide backup paths for each transmitted signal within the same router, without the need to change the route of the signal within the network. Moreover, we proposed a simultaneous transmission probability analysis for optical routers to show the feasibility of this proposed method. This probability analyzes all the possible signals that can be transmitted at the same time within the default and the backup paths of the router. Our research work shows that the simultaneous transmission probability is improved by 10% to 46% compared to other fault‐tolerant optical routers. Furthermore, the worst‐case insertion loss of our scheme can be reduced by 46.34% compared to others. The worst‐case crosstalk noise is also reduced by 24.55%, at least, for the default path and 15.7%, at least, for the backup path. Finally, in the network level, the OSNR is increased by an average of 68.5% for the default path and an average of 15.9% for the backup path, for different sizes of the network. Meaad Fadhel, Huaxi Gu |
Wirel. Commun. Mob. Comput. | 3 |
| 2020 | Multi-Controller Placement Based on Two-Sided Matching in Inter-Datacenter Elastic Optical NetworksabstractSoftware defined networking (SDN) can bring considerable flexibility to the management of inter-DC elastic optical networks. However, in the large-scale network, unreasonable placement of multiple SDN controllers may cause the unbalanced distribution of controller loads. To address this issue, we propose a two-sided matching method (TSMM) and design its corresponding algorithm to implement the optimal multi-controller placement. We solve the controller placement problem as the two-sided matching problem and consider optimizing multiple metrics which affect the controller placement. Different from previous research ideas, TSMM implements placement choice from the bilateral perspective of switch and controller. Additionally, TSMM algorithm can achieve the optimal controller placement by maximizing mutual satisfaction between controller and switch. The simulation results show that TSMM can achieve the optimal multi-controller placement and the balanced distribution of controller loads when compared with the existing schemes. Yong Liu 0038, Huaxi Gu, Fulong Yan, Xiaoshan Yu 0001, Kun Wang 0001 |
ICC | 2 |
| 2020 | Countering Variations and Thermal Effects for Accurate Optical Neural NetworksabstractOptical neural networks (ONNs) have emerged as a promising high-performance computing platform to accelerate deep neural networks. In ONNs, phases of light are modulated through Mach-Zehnder Interferometers (MZIs), and MZIs are connected in a gridlike layout to implement multiply-accumulate operations. However, ONNs are very sensitive to process variations and thermal effects. This sensitivity leads to a significant degradation of inference accuracy of ONNs and thus renders them unusable in practice. In this paper, we propose a framework to calibrate process variations and counter thermal effects by power compensation. Experimental results demonstrate that the proposed framework can recover the inference accuracy under variations and thermal effects, e.g., from as low as 11.05% back to 74.11% for LeNet-5 on Cifar10, so that ONNs can achieve an inference accuracy similar to the accuracy after software training while providing their high bandwidth in neuromorphic computing. Ying Zhu 0008, Grace Li Zhang, Bing Li 0005, Xunzhao Yin, Cheng Zhuo, Huaxi Gu, Tsung-Yi Ho, Ulf Schlichtmann |
ICCAD | 6 |
| 2020 | Lotus: A New Topology for Large-scale Distributed Machine LearningabstractMachine learning is at the heart of many services provided by data centers. To improve the performance of machine learning, several parameter (gradient) synchronization methods have been proposed in the literature. These synchronization algorithms have different communication characteristics and accordingly place different demands on the network architecture. However, traditional data-center networks cannot easily meet these demands. Therefore, we analyze the communication profiles associated with several common synchronization algorithms and propose a machine learning--oriented network architecture to match their characteristics. The proposed design, named Lotus, because it looks like a lotus flower, is a hybrid optical/electrical architecture based on arrayed waveguide grating routers (AWGRs). In Lotus, a complete bipartite graph is used within the group to improve bisection bandwidth and scalability. Each pair of groups is connected by an optical link, and AWGRs between adjacent groups enhance path diversity and network reliability. We also present an efficient routing algorithm to make full use of the path diversity of Lotus, which leads to a further increase in network performance. Simulation results show that the network performance of Lotus is better than Dragonfly and 3D-Torus under realistic traffic patterns for different synchronization algorithms. Huaxi Gu, Xiaoshan Yu 0001, Krishnendu Chakrabarty |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2019 | Joint Energy and Spectrum Efficient Virtual Optical Network embedding in EONsabstractNetwork virtualization facilitates the deployment of diversified services and flexible resource management in elastic optical networks (EONs). However, due to the explosive growth of traffic, considerable energy consumption and spectrum usage have restricted the sustainable development of cloud services. This paper addresses joint energy and spectrum efficient problem for virtual optical network embedding (VONE) over EONs. We propose a heuristic algorithm to improve energy and spectrum efficiency while keeping a high acceptance rate. With consideration of factors influencing energy and spectrum efficiency, a feasible shortest path is preferred; meanwhile, an appropriate modulation format is dynamically selected according to transmission distance and the trade-off between energy and spectrum consumption. To improve the acceptance rate of virtual network requests, a dual mapping is employed to reinforce the embedding process by multi-dimensional resources integrated mapping. The simulation results show that the proposed algorithm can achieve a joint energy and spectrum efficiency with a much lower blocking probability compared with the baseline approach. Wenting Wei, Huaxi Gu, Achille Pattavina, Jiru Wang, Yi Zeng 0005 |
HPSR | 2 |
| 2019 | A Virtual Machine Placement Algorithm Combining NSGA-II and Bin-Packing HeuristicabstractThe servers in the data center networks have multi-dimensional physical resources, and there is a lot of diversity in resource consumption among tasks. When virtual machines carrying different user requests are deployed on the same server at the same time, it is very likely that there is an imbalanced usage of multi-dimensional resources, resulting in the waste of physical resources. In this paper, we focus on virtual machine placement in data centers aiming to balance multi-dimensional resource usage and maximize the service rate. To solve such a bi-objective optimization problem, we present a joint bin-packing heuristic and genetic algorithm to reduce the time complexity while obtaining an approximate optimal solution. Wenting Wei, Kun Wang 0001, Shengjun Guo, Huaxi Gu |
PDCAT | 5 |
| 2019 | Improving Cloud-Based IoT Services Through Virtual Network Embedding in Elastic Optical Inter-DC NetworksabstractWith the boom of Internet of Things (IoT), an increasing amount of data from IoT applications is moved to geo-distributed data centers (DCs) for data analysis. Massive compute-demanding applications call for a more flexible and efficient resource allocation for uncertain and heterogeneous traffic in geo-distributed multi-DC systems. Virtual network embedding, a major part of network virtualization, facilitates to provide different kinds of businesses or services by resource sharing. Moreover, due to their elasticity, elastic optical networks are viewed as a very promising solution to support inter-DC networks. This paper focuses on the effectiveness and spectrum fragmentation problem for virtual optical network embedding in elastic optical inter-DC networks by employing multidimensional resources and a topological attribute. In the node mapping, betweenness of a physical node is considered together with multidimensional resource carrying capacity (MRCC) to identify proper matching. Specifically, to reduce the influence of a spectrum fragment, the available spectrum continuity degree is coupled with the computing capacity of a physical node as the MRCC. In the link mapping, a tightest-matching factor is employed for the selection of paths to accommodate virtual links. Compared with baseline algorithms except for the integer linear programming (ILP) solution, analytical and numerous experiments show that our solution reduces the blocking probability by 30% on average, balances the load by 15% on average and improves spectral efficiency significantly. Moreover, our proposal has a slightly lower spectral efficiency but a better blocking performance and a much better link load balance than that of the ILP formulation. Wenting Wei, Huaxi Gu, Kun Wang 0001, Xiaoshan Yu 0001, Xuanzhang Liu |
IEEE Internet Things J. | 2 |
| 2019 | Mesh-of-Torus: a new topology for server-centric data center networks
Peibo Xie, Huaxi Gu, Kun Wang 0001, Xiaoshan Yu 0001, Shangqi Ma |
J. Supercomput. | 2 |
| 2019 | Wavelength-Reused Hierarchical Optical Network on Chip Architecture for Manycore ProcessorsabstractManycore processor is becoming the mainstream platform for cloud computing applications. However, the design of high-performance and sustainable inter-core communication network is still a challenging problem. Optical Network on Chip (ONoC) is an emerging chip-scale optical communication technology with high bandwidth capacity and energy efficiency. In this paper, we present a Wavelength Reused Hierarchical ONoC architecture, WRH-ONoC. It leverages the nonblocking wavelength-routed λ-router and hierarchical networking to reuse the limited number of wavelengths. In WRH-ONoC, all the cores are grouped into multiple subsystems, and the cores in the same subsystem are directly interconnected using a λ-router for nonblocking communication. For inter-subsystem communication, all subsystems are further connected through multiple λ-routers and gateways in a hierarchical manner. Thus, the available wavelengths can be reused in different λ-routers. Furthermore, WRHm-ONoC, an efficient extension with multicast ability is also proposed. Given the numbers of cores and available wavelengths, we derive the minimum hardware requirement, the expected end-to-end delay, and the maximum data rate. Theoretical analysis and simulation results indicate WRH-ONoC achieves prominent improvement on the communication performance and sustainability, e.g., 46.0 percent of reduction on zero-load delay and 72.7 percent of improvement on throughput for 400 cores with the modest hardware/energy costs. Haibo Zhang 0001, Yawen Chen 0001, Zhiyi Huang 0001, Huaxi Gu |
IEEE Trans. Sustain. Comput. | 5 |
| 2019 | TAONoC: A Regular Passive Optical Network-on-Chip Architecture Based on Comb SwitchesabstractOptical networks on chip (ONoC) has been proposed as a promising alternative paradigm for electronic NoC with the benefit of optical signaling communications such as ultrahigh bandwidth, extremely low energy consumption, and negligible transmission latency. To accommodate the layout of tile-based chip multicore processors, a torus-based passive ONoC architecture, TAONoC, is proposed in this paper. Relying on the unique designs of three function modules, TAONoC can still support contention-free communication without the need for arbitration. TAONoC employs comb switches instead of general microring resonators (MRs). TAONoC has a low demand for the number of MRs because of the ultrahigh utilization of resonant wavelengths owned by a single MR. Simulation results show that TAONoC performs well under three different synthetic traffic patterns. Yintang Yang, Huaxi Gu, Bowen Zhang 0004, Lijing Zhu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Thor: A Scalable Hybrid Switching Architecture for Data CentersabstractOptical interconnects are emerging as a high bandwidth and low energy alternative to traditional electrical networks. However, most exiting designs put optical switching in the core layer due to issues, including limited scalability, potential bottleneck of the control plane, and high cost of the multi-wavelength switch. To this end, we propose Thor, a hybrid network architecture using optical interconnects as the main load bearing portion. It employs the hypercube topology and multi-hop circuit switching to enable server-level optical access, thus solving the scalability and connectivity limitations. Then, in order to build a faster control plane operating short-lived circuits in a large scale, Thor uses a distributed control system over the electrical network, with a new path setup mechanism capable of establishing multiple optical circuits concurrently. Moreover, Thor employs a novel optical switch design to utilize multiple wavelengths more efficiently. Our evaluation results show that Thor consumes at least 59% and 6% less power than existing electrical and optical designs. For performance, Thor is able to deliver 90% bisection bandwidth of a non-blocking network and reduce the end-to-end latency significantly. Finally, by separately delivering mice and elephant flows, Thor achieves salient improvements in average flow completion times compared with existing approaches. Xiaoshan Yu 0001, Hong Xu 0001, Huaxi Gu |
IEEE Trans. Commun. | 3 |
| 2018 | A joint optimization method for NoC topology generation
Kun Wang 0001, Huaxi Gu, Yintang Yang, Yawen Chen 0001, Haibo Zhang 0001 |
J. Supercomput. | 3 |
| 2018 | A highly efficient dynamic router for application-oriented network on chip
Huaxi Gu, Kun Wang 0001, Xiaoshan Yu 0001, Bowen Zhang 0004 |
J. Supercomput. | 2 |
| 2017 | Testudo: A Low Latency and High-Efficient Memory-Centric Network Using Optical InterconnectabstractWith the continuing-scaling of future multicore processors, the performance requirements on memory access has been put forward much higher. Memory- centric network is deemed as a promising communication paradigm for core-to-memory interconnect in future multicore processors. However, the traditional electrical interconnect has the drawbacks of limited capacity, high communication delay, poor scalability and low energy efficiency, which further limits the performance improvement of the system. To support high-performance communication for memory access, we propose the Testudo architecture, an optically connected memory-centric network (MCN), which utilizes the emerging optical interconnect technology and 3D-stacking memory technology to achieve high bandwidth, low power consumption and high scalability. Testudo is designed based on multiple optical crossbar organized in a torus- like topology. Each optical crossbar is in multiple-write- multiple-read construction, which provides high connectivity for IP cores. By employing an all optical, token-based arbitration scheme with low complexity, the memory access communication is contention-free. Simulation results show that Testudo improves the performance significantly compared to the electrical mesh topology. Shixiong Qi, Huaxi Gu, Haibo Zhang 0001, Yawen Chen 0001 |
GLOBECOM | 2 |
| 2017 | 3D network-on-chip design for embedded ubiquitous computing systems
Huaxi Gu, Yawen Chen 0001, Yintang Yang, Kun Wang 0001 |
J. Syst. Archit. | 2 |
| 2016 | Flow Driven Energy-Aware Routing Algorithm in Data Center NetworkabstractRecently, many energy-aware routing algorithms are proposed to decrease the energy consumption of data center network. However, these methods ignore the effect of working time on energy consumption. In this paper, we analyze the energy consumption model and propose an energy-aware routing algorithm by jointly considering power consumption and working time. According to the simulation results, the flow driven energy-aware routing algorithm saves nearly 50 percent of energy compared with flow preemption energy-aware routing algorithm when transmission rate is limited by the available bandwidth, while it saves about 58.3 percent of energy compared with flow aggregation energy-aware routing algorithm when the transmission rate is limited by the forwarding rate of server's NIC. Kun Wang 0001, Xiaoshan Yu 0001, Liangkai Liu, Huaxi Gu, Yantao Guo |
PDCAT | 5 |
| 2016 | RingCube - An incrementally scale-out optical interconnect for cloud computing data center
Xiaoshan Yu 0001, Huaxi Gu, Yintang Yang, Kun Wang 0001 |
Future Gener. Comput. Syst. | 2 |
| 2016 | A Highly Scalable Optical Network-on-Chip With Small Network Diameter and Deadlock FreedomabstractTo increase the performance of chip multiprocessors, optical network-on-chip (ONoC) becomes promising because of its high bandwidth and low energy consumption. In this paper, we propose an architecture called RPNoC (Ring-based Packet-switched NoC), which uses few optical devices. Specifically, Single-waveguide RPNoC employs only one waveguide. Multiwaveguide RPNoC introduces space division multiplexing to make the architecture highly scalable. A novel wavelength assignment method and a deadlock-free deterministic routing algorithm are jointly designed, which make the network diameter quite small. This design also guarantees deadlock freedom, a little resource use, and low complexity at the same time. Evaluation is carried out for the 64-node RPNoC under different synthetic and realistic traffic patterns. The simulation result shows that it yields high throughput and low latency. Comparison with other packet-switched ONoCs shows that RPNoC has the lowest energy consumption. Huaxi Gu, Yintang Yang, Kun Wang 0001, Qinfen Hao |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | WRH-ONoC: A wavelength-reused hierarchical architecture for optical Network on ChipsabstractOptical Network on Chip (ONoC) is a promising technology for the next-generation many-core chip multiprocessors owing to its tremendous advantages in low power consumption, low communication delay, and high bandwidth. In this paper we present WRH-ONoC, a novel wavelength-reused hierarchical architecture that is capable of interconnecting thousands of cores using a limited number of wavelengths while providing extremely high-throughput data communication between connected cores. In WRH-ONoC, the cores are divided into small subsystems that are interconnected using multiple λ-routers and gateways in a hierarchical manner. Each λ-router can provide non-blocking parallel communication among the directly connected cores or gateways, and all λ-routers can reuse the limited number of available wavelengths. Communications between cores in different subsystems are routed via gateways in which optical signals can change their wavelengths via optical-electrical signal conversions. For a given number of cores, we give the minimum number of levels, λ-routers, and gateways required to interconnect these cores, and derive the expected end-to-end data communication delay under the Uniform-Poisson traffic pattern. Both theoretical analysis and simulation results demonstrate that WRH-ONoC can achieve significant improvement on performance and reduction on hardware cost in comparison with the existing solutions. Haibo Zhang 0001, Yawen Chen 0001, Zhiyi Huang 0001, Huaxi Gu |
INFOCOM | 5 |
| 2013 | Analyzing Packet-Level Routing in Data CentersabstractData centers host diverse applications with stringent QoS requirements. The key issue is to eliminate network congestions which severely degrade application performance. One effective solution is to balance the traffic load in the datacenter regular topologies. Many previous strategies focused on optimized flow routing, and these solutions can hardly achieve ideal load balance while guaranteeing QoS of different traffic flows due to the limitations in practical. In this paper, we discuss packet-level routing and analyze its merit for fine-grained load balance in data centers. Though packet-level routing interacts poorly with TCP in traditional network settings, we prove that it can be adapted to datacenter environment. Motived by the work done by Dixit [4] [5], we assert that packet-level routing is the right choice for data centers. Our simulation results demonstrate that packet-level routing better fulfills datacenter requirements. Ruoyan Liu, Huaxi Gu, Yawen Chen 0001, Haibo Zhang 0001 |
DASC | 2 |
| 2013 | A New Approach to Multi-objective Virtual Machine Placement in Virtualized Data CenterabstractIn this paper, a virtual machine placement model to maximize resource utilization, balance multi-dimensional resources use and minimize communication traffic simultaneously within the data center is proposed. The multi-objective problem is simplified by employing average valued inequality and positional constraints. The improved genetic algorithm with local heuristic method and elitism strategy is developed to solve the problem. The simulation results show that performance gains in all aspects can be achieved by the proposed model and algorithm compared to the existing algorithms. Sinong Wang, Huaxi Gu |
NAS | 2 |
| 2012 | A New Two-Layer Topology for Data Center NetworkabstractIn the cloud computing era, the goal of data center network is not only to interconnect a large number of servers, but also to provide low latency and high bandwidth. A new 2-layer architecture, C-tree, is proposed for cloud computing in this paper. It is a flat topology, in which the horizontal traffic's latency is obviously lower than in traditional three-layer architectures. Also, C-tree can provide high bisection bandwidth which is important for bandwidth-intensive applications. We analyze C-tree theoretically and compare it with other architectures in scalability, accommodation capability, network diameter, bisection bandwidth, path diversity and regularity. We have developed a load balancing routing mechanism to balance the traffic on the parallel links. A network simulation platform is set up to verify the performance of C-tree. Lei Chang, Huaxi Gu, Kun Wang 0001, Ruoyan Liu |
PDCAT | 2 |
| 2009 | A low-power fat tree-based optical Network-On-Chip for multiprocessor system-on-chipabstractMultiprocessor system-on-chip (MPSoC) is an attractive platform for high-performance applications. Networks-on-chip (NoCs) can improve the on-chip communication bandwidth of MPSoCs. However, traditional metallic interconnects consume significant amount of power to deliver even higher communication bandwidth required in the near future. Optical NoCs are based on CMOS-compatible optical waveguides and microresonators, and promise significant bandwidth and power advantages. This paper proposes a fat tree-based optical NoC (FONoC) including its topology, floorplan, protocols, and a low-power and low-cost optical router, optical turnaround router (OTAR). Different from other optical NoCs, FONoC does not require building a separate electronic NoC for network control. It carries both payload data and network control data on the same optical network, while using circuit switching for the former and packet switching for the latter. The FONoC protocols are designed to minimize network control data and the related power consumption. An optimized turnaround routing algorithm is designed to utilize the low-power feature of OTAR, which can passively route packets without powering on any microresonator in 40% of all cases. Comparing with other optical routers, OTAR has the lowest optical power loss and uses the lowest number of microresonators. An analytical model is developed to characterize the power consumption of FONoC. We compare the power consumption of FONoC with a matched electronic NoC in 45 nm, and show that FONoC can save 87% power comparing with the electronic NoC on a 64-core MPSoC. We simulate the FONoC for the 64-core MPSoC and show the end-to-end delay and network throughput under different offered loads and packet sizes. Huaxi Gu, Jiang Xu 0001, Wei Zhang 0012 |
DATE | 1 |
| 2007 | rHALB: A New Load-Balanced Routing Algorithm for k-ary n-cube Networks
Huaxi Gu, Jie Zhang 0003, Kun Wang 0001, Changshan Wang |
APPT | 1 |
| 2007 | Enhanced fault tolerant routing algorithms using a concept of "balanced ring"
Huaxi Gu, Jie Zhang 0003, Kun Wang 0001, Zengji Liu, Guochang Kang |
J. Syst. Archit. | 1 |
| 2006 | X-Torus: A Variation of Torus Topology with Lower Diameter and Larger Bisection Width
Huaxi Gu, Qiming Xie, Kun Wang 0001, Jie Zhang 0003, Yunsong Li 0001 |
ICCSA (5) | 1 |
| 2005 | A New Routing Method to Tolerate both Convex and ConcaveabstractTo make the exiting fault routing algorithms tolerate concave fault regions without disabling any healthy nodes, the concept of hole is proposed in this paper. A hole consists of healthy nodes in the concave parts and neighborhood of a given concave fault region. By guiding the packet routing inside and outside the hole, the new routing method empowers the convex fault tolerant routing algorithm to tolerate concave shape regions without disabling any healthy nodes. The proposed modification method is simple and does not add new virtual channels. Moreover, it doesn’t change the rules of the previous algorithms. Finally, the performance of the modified routing algorithm is simulated under various concave fault patterns. Huaxi Gu, Zengji Liu, Guochang Kang |
PDCAT | 1 |