Shuqiang Zhang

dblp:63/8970 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
3 papers
Datacenter networks · 33% Internet architecture and protocols · 29% Optical networks · 21%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 33% High-performance computing · 33% Interconnection networks and networks-on-chip · 33%
Databases, data mining, and information retrieval
2 papers
Recommender systems · 88% Machine learning and data management · 12%
Artificial intelligence
1 paper
Probabilistic and Bayesian machine learning · 100%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.912025
Addressing Correlated Latent Exogenous Variables in Debiased Recommender Systems · KDD (2) 2025
Recommender systems
debiased recommendation
0.912025
Addressing Correlated Latent Exogenous Variables in Debiased Recommender Systems · KDD (2) 2025
Recommender systems › debiased recommendation
selection bias
0.912025
Addressing Correlated Latent Exogenous Variables in Debiased Recommender Systems · KDD (2) 2025
Datacenter networks › RDMA
RDMA over Converged Ethernet
0.812024
RDMA over Ethernet for Distributed Training at Meta Scale · SIGCOMM 2024
High-performance computing
collective communication
0.812024
RDMA over Ethernet for Distributed Training at Meta Scale · SIGCOMM 2024
Distributed systems › distributed machine learning
distributed training
0.812024
RDMA over Ethernet for Distributed Training at Meta Scale · SIGCOMM 2024
Interconnection networks and networks-on-chip › remote direct memory access
RDMA networks
0.812024
RDMA over Ethernet for Distributed Training at Meta Scale · SIGCOMM 2024
Internet architecture and protocols
wide area network
0.712023
EBB: Reliable and Evolvable Express Backbone Network in Meta · SIGCOMM 2023
Routing and switching › traffic engineering
MPLS traffic engineering
0.212023
EBB: Reliable and Evolvable Express Backbone Network in Meta · SIGCOMM 2023
Routing and switching
traffic engineering
0.212023
EBB: Reliable and Evolvable Express Backbone Network in Meta · SIGCOMM 2023
Optical networks › traffic grooming
dynamic traffic grooming
0.212013
Dynamic Traffic Grooming in Elastic Optical Networks · IEEE J. Sel. Areas Commun. 2013
Optical networks › elastic optical networks
routing and spectrum assignment
0.212013
Dynamic Traffic Grooming in Elastic Optical Networks · IEEE J. Sel. Areas Commun. 2013
Optical networks
traffic grooming
0.212013
Dynamic Traffic Grooming in Elastic Optical Networks · IEEE J. Sel. Areas Commun. 2013

Methods — techniques the papers use, named apart from their topics

structural causal model · 1.7monte carlo algorithm · 1.7likelihood maximization · 1.7distributed control agents · 0.7centralized traffic engineering · 0.7integer linear programming · 0.2heuristic algorithm · 0.2auxiliary graph · 0.2
YearPublicationVenuePosition
2026 Connecting 100K+ GPUs: Building the Communication Stack for Large-Scale LLM Training
abstract
The arrival of 100K+ GPU clusters marks a new frontier in AI infrastructure. Standard communication stack meets new challenges as physical topologies span multiple datacenter buildings, introducing high bandwidth-delay product links where latency increases by up to 30× compared to intra-rack traffic. Furthermore, the transition toward Mixture-of-Experts architectures generating bursty all-to-all patterns that create transient congestion hotspots. These constraints, combined with an operational environment where hardware failures shift from anomalies to frequent occurrences, renders traditionally lightweight operations like initialization and resource management challenging.
Hongyi Zeng, Min Si, Pavan Balaji, Yongzhou Chen, Ching-Hsiang Chu, Adithya Gangidi, Prashanth Kannan, Bingzhe Liu, Saif Hasan, Deep Shah, Ashmitha Jeevaraj Shetty, Gregory R. Steinbrecher, Srikanth Sundaresan, Yulun Wang, Yexin Wu, Mingran Yang, Kenny Yu, Minlan Yu, Cen Zhao, Shengbao Zheng, Wesley Bland, Denis Boyda, Suman Gumudavelli, Subodh Iyengar, Cristian Lumezanu, Rui Miao 0001, Venkat Ramesh, Jingliang Ren, Maxim Samoylov, Jan Seidel, Qiye Tan, Xinfeng Xie, Yimeng Zhao, Shuqiang Zhang, Art Zhu
SIGCOMM40
2025 Addressing Correlated Latent Exogenous Variables in Debiased Recommender Systems
abstract
Recommendation systems (RS) aim to provide personalized content, but they face a challenge in unbiased learning due to selection bias, where users only interact with items they prefer. This bias leads to a distorted representation of user preferences, which hinders the accuracy and fairness of recommendations. To address the issue, various methods such as error imputation based, inverse propensity scoring, and doubly robust techniques have been developed. Despite the progress, from the structural causal model perspective, previous debiasing methods in RS assume the independence of the exogenous variables. In this paper, we release this assumption and propose a learning algorithm based on likelihood maximization to learn a prediction model. We first discuss the correlation and difference between unmeasured confounding and our scenario, then we propose a unified method that effectively handles latent exogenous variables. Specifically, our method models the data generation process with latent exogenous variables under mild normality assumptions. We then develop a Monte Carlo algorithm to numerically estimate the likelihood function. Extensive experiments on synthetic datasets and three real-world datasets demonstrate the effectiveness of our proposed method. The code is at https://github.com/WallaceSUI/kdd25-background-variable.
Shuqiang Zhang, Yuchao Zhang 0001, Jinkun Chen, Haochen Sui
KDD (2)1
2024 RDMA over Ethernet for Distributed Training at Meta Scale
abstract
The rapid growth in both computational density and scale in AI models in recent years motivates the construction of an efficient and reliable dedicated network infrastructure. This paper presents the design, implementation, and operation of Meta's Remote Direct Memory Access over Converged Ethernet (RoCE) networks for distributed AI training.
Adithya Gangidi, Rui Miao 0001, Shengbao Zheng, Sai Jayesh Bondu, Guilherme Loch Waltrick Goes, Hany Morsy, Rohit Puri, Mohammad Riftadi, Ashmitha Jeevaraj Shetty, Shuqiang Zhang, Mikel Jimenez Fernandez, Shashidhar Gandham, Hongyi Zeng
SIGCOMM11
2023 EBB: Reliable and Evolvable Express Backbone Network in Meta
abstract
We present the design, implementation, evaluation, deployment and production experiences of EBB (Express BackBone), a private WAN (Wide Area Network) connecting Meta's global data centers (DCs). Initiated in 2015, EBB now carries 100% of DC-DC traffic, witnessing remarkable growth over the years. A key design aspect of EBB is its multi-plane architecture, facilitating seamless deployment of a new control plane while ensuring operational simplicity. This architecture allows for efficient failure mitigation, standard maintenance, and capacity expansion by draining one or two planes without impacting service level objectives (SLOs). Another critical design decision is the hybrid model, combining distributed control agents and a central controller. EBB's centralized traffic engineering utilizes an MPLS-TE based solution to allocate paths periodically for different traffic classes based on service requirements, while its distributed control agents enable fast local failure recovery by pre-installing pre-computed backup paths in the data plane. We delve into our eight-year production experience, highlighting the successful deployment of multiple generations of EBB.
Marek Denis, Yuanjun Yao, Ashley Hatch, Chiunlin Lim, Shuqiang Zhang, Kyle Sugrue, Henry Kwok, Mikel Jimenez Fernandez, Petr Lapukhov, Sandeep Hebbani, Gaya Nagarajan, Omar Baldonado, Lixin Gao 0001, Ying Zhang 0022
SIGCOMM6
2013 Connecting the clouds with low-latency, low-cost virtual private lines enabled by sliceable optical networks
abstract
Cloud computing is evolving as a major priority for many enterprises. Better wide-area networks with low latency and cost are needed to interconnect geographically-distributed data centers and offices using virtual private lines (VPLs). We propose novel network architectures based on sliceable optical networks to implement future VPLs. Optimal designs of the new network architectures and the traditional packet-over-optical network architecture are proposed and compared. It is found that the new architectures can achieve the `lowest-possible' latency with potentially lower cost than traditional architecture.
Shuqiang Zhang, Rui Wang 0025, Uttam Mandal, M. Farhan Habib, Biswanath Mukherjee
GLOBECOM1
2013 Dynamic Traffic Grooming in Elastic Optical Networks
abstract
Spectrum elastic optical networks support flexible central frequency and spectrum assignment for lightpaths. When provisioning a new connection in an elastic optical network that allows traffic grooming, the control plane has to solve two problems: the electrical-layer routing and optical-layer routing and spectrum assignment (RSA). The electrical-layer routing determines how to route the new connection through a combination of new and existing lightpaths, while the optical-layer RSA decides how to establish new lightpaths under the spectrum-continuity constraint. The flexibility (e.g., bandwidth variability of lightpaths) provided by elastic optical networks makes it suitable for accommodating dynamic traffic. It is important and challenging to exploit the full potential of the flexibility when dealing with the above two problems. In this study, we propose a multi-layer auxiliary graph to jointly solve the electrical-layer routing and optical-layer RSA. Various traffic-grooming policies (objectives) can be achieved by properly adjusting the edge weights in the auxiliary graph. Also, we propose a spectrum reservation scheme that can efficiently utilize the bandwidth variability of lightpaths by reserving bandwidth for non-fully utilized lightpaths and grooming future connections onto them. We show that there is a tradeoff among different traffic-grooming policies, and the spectrum reservation scheme can be easily incorporated into various traffic-grooming policies and lead to a significant reduction in operational expenditure (OPEX) and better spectrum efficiency.
Shuqiang Zhang, Chip Martel, Biswanath Mukherjee
IEEE J. Sel. Areas Commun.1
2012 Energy-efficient dynamic provisioning for spectrum elastic optical networks
abstract
Spectrum elastic optical networks support flexible central frequency and spectrum assignment for lightpaths. In this paper, we investigate energy-efficient dynamic provisioning for such networks. When provisioning a connection, the routing problems in both electrical layer (routing over multiple lightpaths) and optical layer (routing over multiple fibers) have to be addressed. Also, the control plane has to determine the optical-layer spectrum assignment considering the spectrum-continuity constraint as well as the bandwidth variability of transponders. We adopt a novel auxiliary graph based on which a new dynamic provisioning policy called Time-Aware Provisioning with Bandwidth Reservation (TAP-BR) is proposed. TAP-BR incorporates two important factors to facilitate energy-efficient provisioning: time awareness and bandwidth reservation. We compare TAP-BR with previously-proposed dynamic provisioning policies and show that TAP-BR can save significant amount of energy and make efficient use of spectrum resources.
Shuqiang Zhang, Biswanath Mukherjee
ICC1
2010 Energy Efficient Time-Aware Traffic Grooming in Wavelength Routing Networks
abstract
In this paper, we investigate both static and dynamic traffic grooming problems in a wavelength routing network, so as to minimize the total energy consumption of the core network, with the additional consideration of the holding times of the lightpaths and connection requests. In static case, all connection requests with their setup and tear-down times are known in advance, we formulate an Integer Linear Programming (ILP) to minimize the energy consumption. In dynamic case, we adopt a layered graph model called Grooming Graph and propose a new traffic grooming heuristics called Time-Aware Traffic Grooming (TATG) which takes the holding time of a new arrival connection request and the remaining holding time of existing lightpaths into consideration. We compare the energy efficiency of different traffic grooming policies under various traffic loads, and the results provide implications to choose the most energy-efficient traffic grooming policies under various scenarios.
Shuqiang Zhang, Chun-Kit Chan
GLOBECOM1