Wu Xiang

dblp:12/7846 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0002-6733-0947ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CXL-CCL: Inter-Node Collective GPU-Communication Using a CXL Shared Memory Pool
abstract
Large language models (LLMs) training or inference across multiple nodes introduces significant pressure on GPU memory and interconnect bandwidth. The Compute Express Link (CXL) shared memory pool offers a scalable solution by enabling memory sharing across nodes, reducing over-provisioning and improving resource utilization. We propose CXL-CCL, a collective communication library, leveraging the CXL shared memory pool to support cross-node GPU operations without relying on traditional RDMA-based networking. Our design addresses the challenges in synchronization, data interleaving, and communication parallelization faced by using the CXL shared memory pool for collective communications. Evaluating on multiple nodes with a TITAN-II CXL switch and six Micron CZ120 memory cards, we show that CXL-CCL achieves highly efficient collective operations across hosts, demonstrating CXL’s potential for scalable, memory-centric GPU communication. Our evaluation demonstrates that CXL-CCL achieves average performance improvements of 1.34 × for AllGather, 1.84 × for Broadcast, 1.94 × for Gather, and 1.07 × for Scatter, compared to the original RDMA-based implementation over 200 Gbps InfiniBand. In addition, an LLM training case study shows 1.11 × speedup compared with the InfiniBand while reducing interconnect hardware cost by 2.75 ×.
Dong Xu 0024, Han Meng, Dengcheng Zhu, Liguang Xie, Wu Xiang, Henry Hu, Hui Zhang 0033, Dong Li 0001
ICS8
2026 Coarse-Fine: A Novel VPA Strategy for Boosting Resource Utilization of Transient Offline Tasks
abstract
To improve resource utilization in Cloud clusters and reduce operational costs, this paper proposes a two-level collaborative Vertical Pod Autoscaler (VPA) strategy for offline tasks in hybrid deployment environments. Offline tasks feature significant runtime variation and low periodicity—characteristics that make existing VPA strategies ineffective. The proposed approach combines coarse-grained adjustments (dynamically recommending resource settings based on cluster-wide utilization trends) with fine-grained adjustments (performing instance-level resource recalibration via sliding windows). It addresses the averaging perspective limitation of coarse-grained methods without requiring predictive models or prior knowledge, ensuring interpretability and operational simplicity. Production deployment verifies that the strategy elevates offline resource utilization to near-target levels, achieving an 8.62% improvement in offline resource utilization and a 41.5% increase in offline task deployment volume under Out-of-Memory (OOM) and computational constraints.
Yingying Wen, Yingzhi Chu, Zhecheng Lin, Wu Xiang
WWW6
2025 From Bottleneck to Breakthrough: Optimizing Scheduling for Hyperscale Containerized Clusters
abstract
Container orchestration platforms have become the backbone of modern private cloud infrastructure, offering flexibility, scalability, and reliability for managing containerized workloads. Among them, Kubernetes has emerged as the de facto standard, widely adopted across the industry, with many organizations operating tens or even hundreds of clusters globally. However, as infrastructure scales to ultra-large clusters (exceeding 10,000 nodes) and faces extreme provisioning scenarios—such as abrupt workload surges, irrespective of available resource headroom—scheduling latency becomes a critical bottleneck. This issue is not adequately addressed by the default Kubernetes scheduler or existing alternative solutions, which struggle to maintain low latency under such conditions.
Yuquan Ren, Xinyi Song, Zhilei Liu, Caixue Lin, Wu Xiang
SoCC8
2024 Resource Allocation with Service Affinity in Large-Scale Cloud Environments
abstract
Containerization has garnered substantial favor among cloud service providers. Nevertheless, the notable network overhead incurred between containers has prompted concerns within the community. In cloud resource scheduling, collocating service containers that frequently communicate to the same machine - termed “service affinity” - is instrumental in enhancing application performance. In response to this concern, we present a solution that harnesses service affinity and collocates containers to enhance the overall system performance and stability. To maximize the benefits of collocating containers, it is necessary to calculate a new schedule that optimally and efficiently maximizes service affinity, especially within the expansive domain of industry-scale cloud environments. In pursuit of this, we leverage the skewness property of affinity and machine learning to fuse solver-based algorithms, thereby assuring both quality and efficiency for problems at scale. Our methodology encompasses the partitioning of a given task into discrete subproblems, with a keen focus on resolving the most critical ones. Via a graph neural network classifier, we assign each subproblem to be solved independently using methods based on off-the-shelf solvers in our algorithm pool - namely, MIP-based, or column generation. This strategic approach enables the efficient computation of a schedule for a cloud cluster that fully optimizes the overall service affinity. We further propose a heuristic algorithm to compute executable container migration plans for practical use, facilitating the transition to the new placement where service affinity is well optimized. Our solution has been deployed in our large-scale production environment, covering over a million cores within ByteDance. Through the successful real-world production deployment, our approach exhibits an average improvement in end-to-end latency by 23.75% and a reduction in request error rates by 24.09% compared to the original system.
Zuzhi Chen, Fuxin Jiang, Binbin Chen 0005, Yu Li 0003, Yunkai Zhang 0002, Jianjun Chen 0001, Wu Xiang, Guozhu Cheng, Wei Zhang 0172, Tieying Zhang
ICDE10
2023 Gödel: Unified Large-Scale Resource Management and Scheduling at ByteDance
abstract
Over the last few years, at ByteDance, our compute infrastructure scale has been expanding significantly due to expedited business growth. In this journey, to meet hyper-scale growth, some business groups resorted to managing their own compute infrastructure stack running different scheduling systems such as Kubernetes, YARN which created two major pain points: the increasing resource fragmentation across different business groups and the inadequate resource elasticity between workloads of different business priorities. Isolation across different business groups (and their compute infrastructure management) leads to inefficient compute resource utilization and prevents us from serving the business growth needs in the long run.
Wu Xiang, Yuquan Ren, Chaohui Xin, Chao Xiang, Xinyi Song, Kaiyang Shao, Yuqi Fu, Wilson Wang, Caixue Lin, Yuming Liang
SoCC1
2014 Local Hybrid Coding for Image Classification
abstract
Sparse coding has received considerable research attentions due to its competitive performance for SPM-based image classification algorithms. In sparse coding, each low-level image descriptor (e.g., SIFT) is quantized into a sparse vector using an over-complete dictionary. Two typical schemes for achieving the code sparsity are imposing ℓ1-sparsity penalty on the coding coefficients, or selecting a set of fc-nearest-neighbor bases from the dictionary for locality-aware encoding. In this paper, we discover that different coding schemes usually produce substantially inconsistent coefficients, each preferring either ℓ1-sparsity or bases-locality. We therefore conjecture that different schemes should be explored simultaneously to further enhance the quantization quality. To this end, we propose a novel ensemble framework, Local Hybrid Coding (LHC), to formalize a unified optimization problem for different coding schemes. Specifically, we quantize each image descriptor using two disjoint sets of dictionaries, fcNN bases and non-fcNN bases, from which we efficiently compute a hybrid representation comprising of local coding and sparse coding, respectively. Extensive experiments on three benchmarks verify that LHC can remarkably outperform several state-of-the-art methods for image classification tasks, and bare comparable complexity to the most efficient coding methods.
Wu Xiang, Jianmin Wang 0001, Mingsheng Long
ICPR1