EDBT 2026 Demo / reviewers in the wild / expert
Yu Xia 0001
dblp:28/4326-1
· DBLP profile ↗
17ranked-venue papers
6as first author
1since 2021 · last 2021
0000-0001-7488-0491ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 12 · 2 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Interconnection networks and networks-on-chip · 70% Cloud and datacenter computing · 23% Performance modeling and evaluation · 7% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing › job scheduling
batch scheduling |
0.2 | 1 | 2016 | A Practical Large-Capacity Three-Stage Buffered Clos-Network Switch Architecture · IEEE Trans. Parallel Distributed Syst. 2016 |
Interconnection networks and networks-on-chip › switching network
clos network |
0.2 | 1 | 2016 | A Practical Large-Capacity Three-Stage Buffered Clos-Network Switch Architecture · IEEE Trans. Parallel Distributed Syst. 2016 |
Interconnection networks and networks-on-chip
switch architecture |
0.2 | 1 | 2016 | A Practical Large-Capacity Three-Stage Buffered Clos-Network Switch Architecture · IEEE Trans. Parallel Distributed Syst. 2016 |
Interconnection networks and networks-on-chip › network scheduling
switch scheduling |
0.2 | 1 | 2016 | A Practical Large-Capacity Three-Stage Buffered Clos-Network Switch Architecture · IEEE Trans. Parallel Distributed Syst. 2016 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.2combined input-crosspoint queued scheduling · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Multipath-aware TCP for Data Center Traffic Load-balancingabstractTraffic load-balancing is important to data center performance. However, existing data center load-balancing solutions are either limited to simple topologies or cannot provide satisfactory performance. In this paper, we propose a multipath-aware TCP (MA-TCP) which can sense the path migration of TCP flows. With this new mechanism, the reduction in TCP congestion window due to packet reordering during the path migration can be avoided. This, in turn, makes the path migration more timely as soon as the original path is congested. Furthermore, if the new path is congested (again), the flow can securely continue to migrate without worrying about transmitting rate reduction. Through NS-3 simulations, we show that MA-TCP achieves better flow completion time (FCT) than existing data center load-balancing solutions. Yu Xia 0001, Jinsong Wu 0001, Jingwen Xia, Ting Wang 0001, Sun Mao |
IWQoS | 1 |
| 2018 | Achieving Energy Efficiency in Data Centers Using an Artificial Intelligence Abstraction ModelabstractToday's data center networks are usually over-provisioned for peak workloads. This leads to a great waste of energy since in practice traffic rarely ever hits peak capacity resulting in the links being under-utilized most of the time. Furthermore, the traditional non-traffic-aware routing mechanisms worsen the situation. From the perspective of resource allocation and routing, this paper aims to implement a green data center network and save as much energy as possible. With the benefit of blocking island paradigm, we present a general framework trying to maximize the network power conservation and minimize sacrifices of network performance and reliability. The bandwidth allocation mechanism together with power-aware routing algorithm achieve a bandwidth guaranteed green tighter network. Moreover, our fast efficient heuristics for allocating bandwidth enable the system to scale to large sized data centers. The evaluation result shows that achieving up to more than 50 percent power savings are feasible while guaranteeing network performance and reliability. Ting Wang 0001, Yu Xia 0001, Jogesh K. Muppala, Mounir Hamdi |
IEEE Trans. Cloud Comput. | 2 |
| 2016 | Towards cost-effective and low latency data center network architecture
Ting Wang 0001, Zhiyang Su, Yu Xia 0001, Mounir Hamdi |
Comput. Commun. | 3 |
| 2016 | A Practical Large-Capacity Three-Stage Buffered Clos-Network Switch ArchitectureabstractThis paper proposes a three-stage buffered Clos-network switch (TSBCS) architecture along with a novel batch scheduling (BS) mechanism. We found that TSBCS/BS can be mapped to a “fat” combined input-crosspoint queued (CICQ) switch. Consequently, the well-studied CICQ scheduling algorithms can be directly applied in TSBCS. Moreover, BS drastically reduces the time complexity of TSBCS scheduling when compared with ordinary CICQ switches of the same number of switch ports, which enables us to build a larger-capacity switch with reasonable scheduling complexity. We further show that TSBCS/BS can achieve 100 percent throughput under any admissible traffic if a stable CICQ scheduling algorithm is used. Direct cell forwarding schemes are proposed to overcome the performance drawback of BS under light traffic loads. With extensive simulations, we show that the performance of TSBCS/BS is comparable to that of output-queued switches and the latter are usually considered as theoretical optimal. Yu Xia 0001, Mounir Hamdi, H. Jonathan Chao |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | CLOT: A cost-effective low-latency overlaid torus-based network architecture for data centersabstractIn this paper, we present the design, analysis, and implementation of a novel data center network architecture named CLOT, which delivers significant reduction in the network diameter, network latency, and infrastructure cost. CLOT is built based on a switchless torus topology by adding only a number of most beneficial low-end switches in a proper way. Forming the servers in close proximity of each other in torus topology well implements the network locality. The extra layer of switches largely shortens the average routing path length of torus network, which increases the communication efficiency. We show that CLOT can achieve lower latency, smaller routing path length, higher bisection bandwidth and throughput, and better fault tolerance compared to both conventional hierarchical data center networks as well as the recently proposed CamCube network. Coupled with the coordinate based translated IP addresses, the carefully designed POW routing algorithm helps CLOT achieve its maximum theoretical performance. The sufficient mathematical analysis and theoretical derivation prove both guaranteed and ideal performance of CLOT. Ting Wang 0001, Zhiyang Su, Yu Xia 0001, Mounir Hamdi |
ICC | 3 |
| 2015 | CeMon: A cost-effective flow monitoring system in software defined networks
Zhiyang Su, Ting Wang 0001, Yu Xia 0001, Mounir Hamdi |
Comput. Networks | 3 |
| 2015 | Designing efficient high performance server-centric data center network architecture
Ting Wang 0001, Zhiyang Su, Yu Xia 0001, Jogesh K. Muppala, Mounir Hamdi |
Comput. Networks | 3 |
| 2014 | FlowCover: Low-cost flow monitoring scheme in software defined networksabstractNetwork monitoring and measurement are crucial in network management to facilitate quality of service routing and performance evaluation. Software Defined Networking (SDN) makes network management easier by separating the control plane and data plane. Network monitoring in SDN is lightweight as operators only need to install a monitoring module into the controller. Active monitoring techniques usually introduce too many overheads into the network. The state-of-the-art approaches utilize sampling method, aggregation flow statistics and passive measurement techniques to reduce overheads. However, little work in literature has focus on reducing the communication cost of network monitoring. Moreover, most of the existing approaches select the polling switch nodes by sub-optimal local heuristics. Inspired by the visibility and central control of SDN, we propose FlowCover, a low-cost high-accuracy monitoring scheme to support various network management tasks. We leverage the global view of the network topology and active flows to minimize the communication cost by formulating the problem as a weighted set cover, which is proved to be NP-hard. Heuristics are presented to obtain the polling scheme efficiently and handle flow changes practically. We build a simulator to evaluate the performance of FlowCover. Extensive experiment results show that FlowCover reduces roughly 50% communication cost without loss of accuracy in most cases. Zhiyang Su, Ting Wang 0001, Yu Xia 0001, Mounir Hamdi |
GLOBECOM | 3 |
| 2014 | NovaCube: A low latency Torus-based network architecture for data centersabstractThis paper presents the design, analysis, and implementation of a novel data center network architecture, named NovaCube. Based on regular Torus topology, NovaCube is constructed by adding a number of most beneficial jump-over links, which offers many distinct advantages and practical benefits. Moreover, in order to enable NovaCube to achieve its maximum theoretical performance, a probabilistic oblivious routing algorithm PORA is carefully designed. PORA is a both deadlock and livelock free routing algorithm, which achieves near-optimal performance in terms of average routing path length with better load balancing thus leading to higher throughput. Theoretical derivation and mathematical analysis further prove the good performance of NovaCube and PORA. Ting Wang 0001, Zhiyang Su, Yu Xia 0001, Mounir Hamdi |
GLOBECOM | 3 |
| 2014 | A general framework for performance guaranteed green data center networkingabstractFrom the perspective of resource allocation and routing, this paper aims to save as much energy as possible in data center networks. We present a general framework, based on the blocking island paradigm, to try to maximize the network power conservation and minimize sacrifices of network performance and reliability. The bandwidth allocation mechanism together with power-aware routing algorithm achieve a bandwidth guaranteed tighter network. Besides, our fast efficient heuristics for allocating bandwidth enable the system to scale to large sized data centers. The evaluation result shows that up to more than 50% power savings are feasible while guaranteeing network performance and reliability. Ting Wang 0001, Yu Xia 0001, Jogesh K. Muppala, Mounir Hamdi, Sebti Foufou |
GLOBECOM | 2 |
| 2014 | Fine-grained power control for combined input-crosspoint queued switchesabstractReducing the power consumption of packet switches is becoming increasingly significant to future networks. However, previous research all focused on reducing power in crossbar-based switches, which is either complex or not effective, especially in some extreme cases. This paper proposes to leverage the dynamic voltage and frequency scaling (DVFS) technique in the buffered crossbar-based switches, which is more flexible and simple. The basic idea is to decrease the working frequencies of the crosspoint buffers while still preserving the maximum throughput and the satisfactory delay. Traffic estimators are used at the input and output ports to estimate the traffic arrival rates, based on which the power controller can adjust the working frequencies of the crosspoint buffers at a fine-grained level. Simulation results show that the scheme is effective. Yu Xia 0001, Ting Wang 0001, Zhiyang Su, Mounir Hamdi |
GLOBECOM | 1 |
| 2014 | CheetahFlow: Towards low latency software-defined networkabstractSoftware defined networking (SDN), which enables programmability, has the advantage of global visibility and high flexibility. However, when forwarding new flows in SDN, the interaction between the switch and the controller imposes extra latency such as round-trip time and routing path search time. Even though such latency is acceptable for elephant flows since it only takes limited ratio of total transmission time of elephant flows, it is an overkill to pay certain overheads for mice flows due to the their short transmission time. Moreover, the controller is frequently invoked by the mice flows since the number of mice flows accounts for a large portion of the total number of flows. Hence, the frequent controller invocation is mainly responsible for the controller performance degradation, and thus increasing the flow setup latency significantly. To solve this problem, we propose CheetahFlow, a novel scheme to predict frequent communication pairs via support vector machine and proactively setup wildcard rules to reduce flow setup latency. Particularly, in order to avoid congestion along a fixed path, elephant flows are detected and rerouted to the non-congestion path efficiently by applying blocking island paradigm. Extensive experiments show that CheetahFlow prominently reduces latency without any loss of flexibility of SDN. Zhiyang Su, Ting Wang 0001, Yu Xia 0001, Mounir Hamdi |
ICC | 3 |
| 2014 | SprintNet: A high performance server-centric network architecture for data centersabstractThis paper presents the design, implementation and evaluation of SprintNet, a novel network architecture for data centers. SprintNet achieves high performance in network capacity, fault tolerance, and network latency. SprintNet is also a scalable, yet low-diameter network architecture where the maximum shortest distance between any pair of servers can be limited by no more than four and is independent of the number of layers. The specially designed routing schemes for SprintNet strengthen its merits. Both theoretical analysis and simulation experiments are conducted to evaluate its overall performance with respect to average path length, aggregate bottleneck throughput, and fault tolerance. Ting Wang 0001, Zhiyang Su, Yu Xia 0001, Yang Liu 0081, Jogesh K. Muppala, Mounir Hamdi |
ICC | 3 |
| 2014 | Improving the efficiency of server-centric data center network architecturesabstractData center network architecture is regarded as one of the most important determinants of network performance. As the most typical representatives of architecture design, the server-centric scheme stands out due to its good performance in various aspects. However, there still exist some critical shortcomings in these server-centric architectures. In order to provide an efficient solution to these shortcomings and improve the efficiency of server-centric architectures, in this paper, we propose a hardware based approach, named “Forwarding Unit”. Furthermore, we put forward a traffic aware routing scheme for FlatNet to further evaluate the feasibility and efficiency of our approach. Both theoretical analysis and simulation experiments are conducted to measure its overall performance with respect to cost-effectiveness, fault-tolerance, system latency, packet loss ratio, aggregate bottleneck throughput, and average path length. Ting Wang 0001, Yu Xia 0001, Dong Lin, Mounir Hamdi |
ICC | 2 |
| 2013 | On practical stable packet scheduling for bufferless three-stage Clos-network switchesabstractIn this paper, we extend our previous work of StablePlus, a stable scheduling algorithm for single-stage packet switches, to bufferless three-stage Clos-network switches. StablePlus is based on an existing stable distributed scheduling algorithm, called DISQUO. We further improve the switching performance by incorporating a heuristic scheduling algorithm after the DISQUO scheduling. In a three-stage Clos-network switch, DISQUO is first used to solve the output contention which generates a stable matching between the input and output ports, then Karol's algorithm is used to find the feasible internal paths for the matched input and output pairs. However, the latter requires multiple mini-cycles to complete the path-finding task. Worse is that the number of mini-cycles increases as the switch size does, limiting the Clos-network to a small implementable size. By replacing the Hamiltonian Walk in DISQUO with time-division multiplexing (TDM) scheme, we show that the number of required mini-cycles for Karol's algorithm can be reduced to only two, independent of the switch size. Moreover, with the help of a parallel hardware approach, we can implement packet scheduling in O(1) time complexity. To support high data rates, e.g., 100 Gbps, we can also make the scheduling work on a frame basis. We prove that StablePlus can achieve 100% throughput under any admissible traffic, and by simulations we show that it also has good delay performance. Yu Xia 0001, H. Jonathan Chao |
HPSR | 1 |
| 2012 | Module-level matching algorithms for MSM clos-network switchesabstractIn this paper, we propose a simple module-level matching scheme for memory-space-memory Clos-network switches to avoid complex path-allocation algorithms in bufferless Clos networks, as well as cell out-of-order and saturation-tree problems in buffered Clos networks. We show that the module-level matching scheme can achieve 100% throughput.We propose static and dynamic dispatching cell schemes in addition to the module-level matching to improve the delay performance. The static cell dispatching scheme requires no additional scheduling; while the dynamic cell dispatching scheme is more adaptive to the traffic than the static one, thus can achieve better delay performance under non-uniform traffic loads. However, the wiring complexity of the scheduler for dynamic cell dispatching is high. Thus, the grouped dynamic cell dispatching scheme is proposed as a trade-off between the complexity and performance. In practice, embedded memory size is restricted, thus the queue length limitation in each switch module is also considered in this paper. We propose an efficient scheme to prevent queues to overflow in this situation which makes our work more practical. Yu Xia 0001, H. Jonathan Chao |
HPSR | 1 |
| 2011 | StablePlus: A practical 100% throughput scheduling for input-queued switchesabstractThis paper proposes a practical stable packet scheduling algorithm for input-queued switches, called StablePlus, which combines a stable matching with a heuristic matching. It not only achieves 100% throughput under any admissible traffic but also has good delay performance. StablePlus can be implemented with today's technology for high line rates, e.g., 100Gbps, and a relatively large input-queued switch, e.g., a few hundred ports. Yu Xia 0001, H. Jonathan Chao |
HPSR | 1 |