Xiaoshan Yu 0001

dblp:129/4851-1 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0001-8550-5697ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 1 first-author · 6 since 2021Computer networks · 6 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Breaking Robustness Barriers in Cognitive Diagnosis: A One-Shot Neural Architecture Search Perspective
Ziwen Wang 0006, Shangshang Yang, Xiaoshan Yu 0001, Haiping Ma, Xingyi Zhang 0001
KDD (1)3
2026 FlexLoop: A Distributed Scheduling Strategy for AWGR-Based Optical Networks
abstract
Optical networks based on arrayed waveguide grating routers (AWGRs) offer significant advantages, including low latency, high bandwidth, and low power consumption. However, addressing the contention issues associated with AWGRs has become a primary concern in practical deployment. Previous research has focused on centralized control methods to address this challenge, but their scalability is often limited. Additionally, distributed time-slot strategies have been implemented, but fixed time-slot configurations can result in performance loss. To effectively overcome these challenges, we propose a distributed scheduling strategy called FlexLoop. Authorizations are issued by receiving nodes and passed between sending nodes via the software-defined networking (SDN) switches. The authorization holder is permitted to transmit a batch of data to the designated node. Each receiving node has the autonomy to define the authorization passing rules through SDN switch configurations. Similarly, each sending node retains the flexibility to either use the authorization to transmit data or forward it to other nodes. Therefore, FlexLoop operates in a fully distributed manner and shows adaptability to various traffic patterns. Experimental results demonstrate that FlexLoop reduces packet latency by up to 92.01% and improves throughput by up to 77.09% compared to NegotiaToR. Additionally, a prototype system was developed, which validates the practical feasibility of implementing FlexLoop.
Zhaoxing Zhou, Huaxi Gu, Xiaoshan Yu 0001, Yunhao Wang 0001
IEEE Trans. Computers3
2026 RCC: Rate-Based Congestion Control for the Lossless Network
Xiaoshan Yu 0001, Huaxi Gu
IEEE Trans. Netw. Serv. Manag.1
2026 DATCP: A Dynamic and Adaptive Congestion Control Protocol for Reconfigurable Data Center Networks
abstract
With the rapid growth of cloud computing and distributed machine learning, traditional static networks increasingly struggle to meet dynamic communication demands. Optical Circuit Switching brings high bandwidth, low latency, and reconfigurability to data center networks, but the resulting reconfigurable data center networks (RDCNs) introduce dynamic path changes and heterogeneous paths, challenging conventional congestion control protocols. To address this, we propose DATCP, a dynamic and adaptive congestion control protocol designed for RDCNs. DATCP features three key mechanisms: (1) aTTL-based path classification method for accurate and lightweight identification of network types; (2) a receiver-side flow tracking model that detects incast congestion and applies weighted rate control; and (3) an adaptive congestion window algorithm that enables fast convergence and efficient bandwidth utilization. Experiments under synthetic and real-world workloads show that DATCP outperforms existing TCP variants, achieving higher throughput, reducing flow completion time and 99th percentile tail latency by up to 37.8% and 21.1%, respectively, and improving flow fairness.
Hong Zou, Huaxi Gu, Xiaoshan Yu 0001, Boxuan Liu
IEEE Trans. Netw.3
2025 SROdcn: Scalable and Reconfigurable Optical DCN Architecture for High-Performance Computing
abstract
Data Center Network (DCN) flexibility is critical for providing adaptive and dynamic bandwidth while optimizing network resources to manage variable traffic patterns generated by heterogeneous applications. To provide flexible bandwidth, this work proposes a machine learning approach with a new Scalable and Reconfigurable Optical DCN (SROdcn) architecture that maintains dynamic and non-uniform network traffic according to the scale of the high-performance optical interconnected DCN. Our main device is the Fiber Optical Switch (FOS), which offers competitive wavelength resolution. We propose a new top-of-rack (ToR) switch that utilizes Wavelength Selective Switches (WSS) to investigate Software-Defined Networking (SDN) with machine learning-enabled flow prediction for reconfigurable optical Data Center Networks (DCNs). Our architecture provides highly scalable and flexible bandwidth allocation. Results from Mininet experimental simulations demonstrate that under the management of an SDN controller, machine learning traffic flow prediction and graph connectivity allow each optical bandwidth to be automatically reconfigured according to variable traffic patterns. The average server-to-server packet delay performance of the reconfigurable SROdcn improves by 42.33% compared to inflexible interconnects. Furthermore, the network performance of flexible SROdcn servers shows up to a 49.67% latency improvement over the Passive Optical Data Center Architecture (PODCA), a 16.87% latency improvement over the optical OPSquare DCN, and up to a 71.13% latency improvement over the fat-tree network. Additionally, our optimized Unsupervised Machine Learning (ML-UnS) method for SROdcn outperforms Supervised Machine Learning (ML-S) and Deep Learning (DL).
Kassahun Geresu, Huaxi Gu, Xiaoshan Yu 0001, Meaad Fadhel, Hui Tian 0001, Wenting Wei
IEEE Trans. Cloud Comput.3
2025 Halo: An Efficient and Scalable Distributed Control Strategy for AWGR-Based Optical Networks
abstract
Optical networks based on arrayed waveguide grating routers (AWGRs) present a promising solution for rapidly evolving data centers (DCs). However, the contention issue remains a significant challenge to their widespread adoption. Existing control strategies typically rely on time-slots that are both synchronized across the network and of fixed duration. This results in the fixed switching granularity. The determination of time-slot duration often involves a trade-off between flexibility and control overhead. In this paper, we propose Halo, a distributed control strategy that avoids contention without relying on time-slots. Halo decentralizes the scheduling tasks across the network. Each node independently issuesGrants, which are passed sequentially among source nodes with communication demands. The node holding theGrantis allowed to transmit a batch of data to the issuing node. Once the demand is no longer present, or the transmitted packet count exceeds a threshold, theGrantis forwarded to the next node. As a result, synchronization operations among nodes is unnecessary, allowing the switching granularity to become flexible. Experimental results demonstrate that Halo significantly outperforms existing control strategies. In trace-based simulations, it achieves up to a 75.13% reduction in flow completion time (FCT). Additionally, we developed a prototype system and executed several real applications to demonstrate the practical feasibility of the Halo.
Zhaoxing Zhou, Huaxi Gu, Xiaoshan Yu 0001, Yunhao Wang 0001
IEEE Trans. Netw.3
2024 COCSN: A Multi-Tiered Cascaded Optical Circuit Switching Network for Data Center
abstract
A cascaded network represents a classic scaling-out model in traditional electrical switching networks. Recent proposals have integrated optical circuit switching at specific tiers of these networks to reduce power consumption and enhance topological flexibility. Utilizing a multi-tiered cascaded optical circuit switching network is expected to extend the advantages of optical circuit switching further. The main challenges fall into two categories. First, an architecture with sufficient connectivity is required to support varying workloads. Second, the network reconfiguration is more complex and necessitates a low-complexity scheduling algorithm. In this work, we propose COCSN, a multi-tiered cascaded optical circuit switching network architecture for data center. COCSN employs wavelength-selective switches that integrate multiple wavelengths to enhance network connectivity. We formulate a mathematical model covering lightpath establishment, network reconfiguration, and reconfiguration goals, and propose theorems to optimize the model. Based on the theorems, we introduce an over-subscription-supported wavelength-by-wavelength scheduling algorithm, facilitating agile establishment of lightpaths in COCSN tailored to communication demand. This algorithm effectively addresses scheduling complexities and mitigates the issue of lengthy WSS configuration times. Simulation studies investigate the impact of flow length, WSS reconfiguration time, and communication domain on COCSN, verifying its significantly lower complexity and superior performance over classical cascaded networks.
Huaxi Gu, Xiaoshan Yu 0001, Songyan Wang, Zeshan Chang
IEEE Trans. Cloud Comput.3
2024 ICLB: intelligent controllers load balancing for software-defined based optical data center networks
Kassahun Geresu, Huaxi Gu, Meaad Fadhel, Wenting Wei, Xiaoshan Yu 0001
J. Supercomput.5
2022 Flex: A flowlet-level load balancing based on load-adaptive timeout in DCN
Xinglong Diao, Huaxi Gu, Xiaoshan Yu 0001, Changyun Luo
Future Gener. Comput. Syst.3
2021 An adaptive failure recovery mechanism based on asymmetric routing for data center networks
Yong Liu 0038, Huaxi Gu, Kun Wang 0001, Xiaoshan Yu 0001, Yunhao Wang 0001
J. Supercomput.4
2020 Multi-Controller Placement Based on Two-Sided Matching in Inter-Datacenter Elastic Optical Networks
abstract
Software defined networking (SDN) can bring considerable flexibility to the management of inter-DC elastic optical networks. However, in the large-scale network, unreasonable placement of multiple SDN controllers may cause the unbalanced distribution of controller loads. To address this issue, we propose a two-sided matching method (TSMM) and design its corresponding algorithm to implement the optimal multi-controller placement. We solve the controller placement problem as the two-sided matching problem and consider optimizing multiple metrics which affect the controller placement. Different from previous research ideas, TSMM implements placement choice from the bilateral perspective of switch and controller. Additionally, TSMM algorithm can achieve the optimal controller placement by maximizing mutual satisfaction between controller and switch. The simulation results show that TSMM can achieve the optimal multi-controller placement and the balanced distribution of controller loads when compared with the existing schemes.
Yong Liu 0038, Huaxi Gu, Fulong Yan, Xiaoshan Yu 0001, Kun Wang 0001
ICC4
2020 Lotus: A New Topology for Large-scale Distributed Machine Learning
abstract
Machine learning is at the heart of many services provided by data centers. To improve the performance of machine learning, several parameter (gradient) synchronization methods have been proposed in the literature. These synchronization algorithms have different communication characteristics and accordingly place different demands on the network architecture. However, traditional data-center networks cannot easily meet these demands. Therefore, we analyze the communication profiles associated with several common synchronization algorithms and propose a machine learning--oriented network architecture to match their characteristics. The proposed design, named Lotus, because it looks like a lotus flower, is a hybrid optical/electrical architecture based on arrayed waveguide grating routers (AWGRs). In Lotus, a complete bipartite graph is used within the group to improve bisection bandwidth and scalability. Each pair of groups is connected by an optical link, and AWGRs between adjacent groups enhance path diversity and network reliability. We also present an efficient routing algorithm to make full use of the path diversity of Lotus, which leads to a further increase in network performance. Simulation results show that the network performance of Lotus is better than Dragonfly and 3D-Torus under realistic traffic patterns for different synchronization algorithms.
Huaxi Gu, Xiaoshan Yu 0001, Krishnendu Chakrabarty
ACM J. Emerg. Technol. Comput. Syst.3
2019 Improving Cloud-Based IoT Services Through Virtual Network Embedding in Elastic Optical Inter-DC Networks
abstract
With the boom of Internet of Things (IoT), an increasing amount of data from IoT applications is moved to geo-distributed data centers (DCs) for data analysis. Massive compute-demanding applications call for a more flexible and efficient resource allocation for uncertain and heterogeneous traffic in geo-distributed multi-DC systems. Virtual network embedding, a major part of network virtualization, facilitates to provide different kinds of businesses or services by resource sharing. Moreover, due to their elasticity, elastic optical networks are viewed as a very promising solution to support inter-DC networks. This paper focuses on the effectiveness and spectrum fragmentation problem for virtual optical network embedding in elastic optical inter-DC networks by employing multidimensional resources and a topological attribute. In the node mapping, betweenness of a physical node is considered together with multidimensional resource carrying capacity (MRCC) to identify proper matching. Specifically, to reduce the influence of a spectrum fragment, the available spectrum continuity degree is coupled with the computing capacity of a physical node as the MRCC. In the link mapping, a tightest-matching factor is employed for the selection of paths to accommodate virtual links. Compared with baseline algorithms except for the integer linear programming (ILP) solution, analytical and numerous experiments show that our solution reduces the blocking probability by 30% on average, balances the load by 15% on average and improves spectral efficiency significantly. Moreover, our proposal has a slightly lower spectral efficiency but a better blocking performance and a much better link load balance than that of the ILP formulation.
Wenting Wei, Huaxi Gu, Kun Wang 0001, Xiaoshan Yu 0001, Xuanzhang Liu
IEEE Internet Things J.4
2019 Mesh-of-Torus: a new topology for server-centric data center networks
Peibo Xie, Huaxi Gu, Kun Wang 0001, Xiaoshan Yu 0001, Shangqi Ma
J. Supercomput.4
2018 Thor: A Scalable Hybrid Switching Architecture for Data Centers
abstract
Optical interconnects are emerging as a high bandwidth and low energy alternative to traditional electrical networks. However, most exiting designs put optical switching in the core layer due to issues, including limited scalability, potential bottleneck of the control plane, and high cost of the multi-wavelength switch. To this end, we propose Thor, a hybrid network architecture using optical interconnects as the main load bearing portion. It employs the hypercube topology and multi-hop circuit switching to enable server-level optical access, thus solving the scalability and connectivity limitations. Then, in order to build a faster control plane operating short-lived circuits in a large scale, Thor uses a distributed control system over the electrical network, with a new path setup mechanism capable of establishing multiple optical circuits concurrently. Moreover, Thor employs a novel optical switch design to utilize multiple wavelengths more efficiently. Our evaluation results show that Thor consumes at least 59% and 6% less power than existing electrical and optical designs. For performance, Thor is able to deliver 90% bisection bandwidth of a non-blocking network and reduce the end-to-end latency significantly. Finally, by separately delivering mice and elephant flows, Thor achieves salient improvements in average flow completion times compared with existing approaches.
Xiaoshan Yu 0001, Hong Xu 0001, Huaxi Gu
IEEE Trans. Commun.1
2018 A highly efficient dynamic router for application-oriented network on chip
Huaxi Gu, Kun Wang 0001, Xiaoshan Yu 0001, Bowen Zhang 0004
J. Supercomput.4
2016 Flow Driven Energy-Aware Routing Algorithm in Data Center Network
abstract
Recently, many energy-aware routing algorithms are proposed to decrease the energy consumption of data center network. However, these methods ignore the effect of working time on energy consumption. In this paper, we analyze the energy consumption model and propose an energy-aware routing algorithm by jointly considering power consumption and working time. According to the simulation results, the flow driven energy-aware routing algorithm saves nearly 50 percent of energy compared with flow preemption energy-aware routing algorithm when transmission rate is limited by the available bandwidth, while it saves about 58.3 percent of energy compared with flow aggregation energy-aware routing algorithm when the transmission rate is limited by the forwarding rate of server's NIC.
Kun Wang 0001, Xiaoshan Yu 0001, Liangkai Liu, Huaxi Gu, Yantao Guo
PDCAT3
2016 RingCube - An incrementally scale-out optical interconnect for cloud computing data center
Xiaoshan Yu 0001, Huaxi Gu, Yintang Yang, Kun Wang 0001
Future Gener. Comput. Syst.1