Xi Peng 0006

dblp:149/7762-6 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0002-8912-3351ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora
abstract
We present AutoSchemaKG, a framework for fully autonomous knowledge graph construction that eliminates the need for predefined schemas. Our system leverages large language models to simultaneously extract knowledge triples and induce comprehensive schemas directly from text, modeling both entities and events while employing conceptualization to organize instances into semantic categories. Processing over 50 million documents, we construct ATLAS (Automated Triple Linking And Schema induction), a family of knowledge graphs with 900+ million nodes and 5.9 billion edges. This approach outperforms state-of-the-art baselines on multi-hop QA tasks and enhances LLM factuality. Notably, our schema induction achieves 92\% semantic alignment with human-crafted schemas with zero manual intervention, demonstrating that billion-scale knowledge graphs with dynamically induced schemas can effectively complement parametric knowledge in large language models.
Jiaxin Bai, Wei Fan 0001, Qing Zong, Hong Ting Tsang, Hongyu Luo, Yauwai Yim, Tianshi Zheng, Xi Peng 0006, Xin Yao 0008, Huiwen Yang, Leijie Wu, J. I Yi, Gong Zhang 0001, Renhai Chen, Yangqiu Song
ACL (1)13
2026 PP-OpenNet: Privacy-Preserved Open Set Classification for Network Traffic
Jingze Zhang, Leijie Wu, Xi Peng 0006, Ruilun Liu, Hong Xu 0001
INFOCOM3
2025 HourglassSketch: An Efficient and Scalable Framework for Graph Stream Summarization
abstract
Graph stream is a special kind of data stream, where every item coming in sequence represents an edge in a dynamic graph. Graph stream has wide application in many fields, including cyber security, social networks and financial fraud detection. In this paper, we propose HourglassSketch, a two-stage data structure, for high-accuracy graph stream summarization. In Stage 1, HourglassSketch uses a CocoSketch to accurately record a partial collection of large-weight edges. In Stage 2, HourglassSketch integrates a TowerSketch with a TCMSketch to approximately record the statistics of most small-weight edges. In addition, we propose a key technique named Error Funnel to further reduce its error margin. Theoretical analysis and experimental results demonstrate that HourglassSketch supports various kinds of query operation and adapts well to graph stream storage. HourglassSketch achieves up to 100x smaller error and 2.7x higher speed than prior work. We also explore the versatility of HourglassSketch as a hardware-friendly framework by implementing it on FPGA and P4 platforms. We have released our codes on GitHub.
Jiarui Guo, Boxuan Chen, Kaicheng Yang 0001, Tong Yang 0003, Zirui Liu 0002, Qiuheng Yin, Yuhan Wu 0001, Bin Cui 0001, Xi Peng 0006, Renhai Chen, Gong Zhang 0001
ICDE12
2024 Speal: Achieving a More Accurate Model with Less Training Data in Performance Evaluation of Storage System through Sampling Optimization
Liang Bao, Hua Wang 0008, Ke Zhou 0001, Ji Zhang 0010, Xi Peng 0006, Renhai Chen, Gong Zhang 0001
DASFAA (2)6
2024 zQoS: Unleashing full performance capabilities of NVMe SSDs while enforcing SLOs in distributed storage systems
abstract
Nowadays, data centers consolidate latency-critical (LC) tenants and best-effort (BE) tenants on the same cloud platform to increase resource utilization and reduce costs. In such a scenario, the underlying distributed storage systems are responsible for guaranteeing SLOs for LC tenants while maximizing bandwidth for BE tenants. As high-performance NVMe SSDs are widely deployed, how to make full use of their performance capabilities and guarantee SLOs has become an urgent problem. However, current methods restrict the performance capabilities of NVMe SSDs based on a conservative offline model, and also ignore runtime changes in tenant loads and device states, which definitely affect the performance capabilities.
Liuying Ma, Zhenqing Liu, Jin Xiong, Renhai Chen, Xi Peng 0006, Gong Zhang 0001, Dejun Jiang 0001
ICPP6
2023 LBFF: Load-Balancing First Fit Algorithm for Tenant Placement Problem
abstract
To meet the prevalence of cloud services, it is of great significance for operators to design a cost-effective tenant placement strategy on physical machines while the services are still guaranteed. Moreover, workload balancing between physical machines is critical for high performance. We characterize the problem of tenant placement as a mixed-integer programming problem that efficiently minimizes the number of active physical machines and balances the workloads among them. To address the NP-hardness of the optimization problem, we first investigate the global lower bound and other inherent properties of the proposed model, and then design two efficient heuristic algorithms. Through numerous simulations at different scales and settings, we demonstrate the superiority of our proposed algorithms over state-of-the-art works in terms of the objective function values and computational time.
Zhenyu Ming, Xi Peng 0006, Liping Zhang 0008
ICC3
2022 DeepQueueNet: towards scalable and generalized network performance estimation with packet-level visibility
abstract
Network simulators are an essential tool for network operators, and can assist important tasks such as capacity planning, topology design, and parameter tuning. Popular simulators are all based on discrete event simulation, and their performance does not scale with the size of modern networks. Recently, deep-learning-based techniques are introduced to solve the scalability problem, but, as we show with experiments, they have poor visibility in their simulation results, and cannot generalize to diverse scenarios. In this work, we combine scalable and generalized continuous simulation techniques with discrete event simulation to achieve high scalability, while providing packet-level visibility. We start from a solid queueing-theoretic modeling of modern networks, and carefully identify the mathematically-intractable or computationally-expensive parts, only which are then modeled using deep neural networks (DNN). Dubbed DeepQueueNet, our approach combines prior knowledge of networks, and supports arbitrary topology and device traffic management mechanisms (given sufficient training data). Our extensive experiments show that DeepQueueNet achieves near-linear speedup in the number of GPUs, and its estimation accuracy for average and 99th percentile round-trip time outperforms existing end-to-end DNN-based performance estimators in all scenarios.
Xi Peng 0006, Li Chen 0008, Libin Liu 0001, Jingze Zhang, Hong Xu 0001, Baochun Li, Gong Zhang 0001
SIGCOMM2
2021 A MAP-based Performance Analysis on 5G-powered Cloud VR Streaming
abstract
Despite that cloud virtual reality (VR) is the most promising service in 5G networks, a reasonable traffic model for it is still unknown. Based on statistics of real cloud VR traces, we justify that the Markovian arrival process (MAP) can well characterize the inter-arrival times (IATs) among packets. Moreover, we discuss possible methods to efficiently estimate parameters for flow aggregation. Since MAPs are analytically tractable, we quantify the performance of delivering cloud VR traffic from a queueing theory perspective. Aiming at supporting quantile metrics, we provide distributions of queue-length and latency for an arbitrary packet. The accuracy of estimated latency is validated by comparing with measured latency on an industrial 5G platform. The MAP-based traffic model will potentially serve as an input for performance evaluation and network planning for assuring high requirements of user experience.
Xi Peng 0006, Fan Zhang 0016, Li Chen 0008, Gong Zhang 0001
ICC1
2018 Bit-Level Power-Law Queueing Theory with Applications in LTE Networks
abstract
Though the classical packet-level queueing theory, which treats each packet as an entry, has achieved a great success in network analysis, it can be inaccurate when directly applied to long-term evolution (LTE) networks. This is because arriving packets could be broken down at the LTE base station server across adjacent transmission time intervals (TTIs), which are the smallest scheduling time units in LTE networks. In this paper, we first propose an innovative bit-level queueing theory to address the challenges in performance analysis of LTE networks. To consider the randomness in packet arrivals and packet lengths, we propose two representative compound network traffic models-Poisson-Exponential (PE) and Zeta-Pareto (ZP) models-to approximate light-tailed and heavy-tailed network traffic, respectively. PE models are suitable for conventional voice and low-speed services, while ZP models, which compound power-law distributions, describe complicated high-speed network applications. We derive tail asymptotics for the distributions of the number of bit arrivals in one TTI and the corresponding waiting time. Based on the results in bit-level queueing theory, we present engineering applications that take into account user experience, including estimating the user experience rate (UER), the busy UER and hourly traffic volume. The theoretical results are then validated through extensive simulations. Our novel traffic estimation approach has been adopted by Wireless Product Line at Huawei for network capacity planning and also projected to International Telecommunication Union (ITU) to compose 5G standards.
Xi Peng 0006, Bo Bai 0001, Gong Zhang 0001, Haofeng Qi, Don Towsley
GLOBECOM1
2017 Layered Group Sparse Beamforming for Cache-Enabled Green Wireless Networks
abstract
The exponential growth of mobile data traffic is driving the deployment of dense wireless networks, which will not only impose heavy backhaul burdens, but also generate considerable power consumption. Introducing caches to the wireless network edge is a potential and cost-effective solution to address these challenges. In this paper, we will investigate the problem of minimizing the network power consumption of cache-enabled wireless networks, consisting of the base station (BS) and backhaul power consumption. The objective is to develop efficient algorithms that unify adaptive BS selection, backhaul content assignment, and multicast beamforming, while taking account of user QoS requirements and backhaul capacity limitations. To address the NP-hardness of the network power minimization problem, we first propose a generalized layered group sparse beamforming (LGSBF) modeling framework, which helps to reveal the layered sparsity structure in the beamformers. By adopting the reweighted ℓ1/ℓ2-norm technique, we further develop a convex approximation procedure for the LGSBF problem, followed by a three-stage iterative LGSBF framework to induce the desired sparsity structure in the beamformers. Simulation results validate the effectiveness of the proposed algorithm in reducing the network power consumption, and demonstrate that caching plays a more significant role in networks with higher user densities and less power-efficient backhaul links.
Xi Peng 0006, Yuanming Shi, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Commun.1
2016 Cache size allocation in backhaul limited wireless networks
abstract
Caching popular content at base stations is a powerful supplement to existing limited backhaul links for accommodating the exponentially increasing mobile data traffic. Given the limited cache budget, we investigate the cache size allocation problem in cellular networks to maximize the user success probability (USP), taking wireless channel statistics, backhaul capacities and file popularity distributions into consideration. The USP is defined as the probability that one user can successfully download its requested file either from the local cache or via the backhaul link. We first consider a single-cell scenario and derive a closed-form expression for the USP, which helps reveal the impacts of various parameters, such as the file popularity distribution. More specifically, for a highly concentrated file popularity distribution, the required cache size is independent of the total number of files, while for a less concentrated file popularity distribution, the required cache size is in linear relation to the total number of files. Furthermore, we study the multi-cell scenario, and provide a bisection search algorithm to find the optimal cache size allocation. The optimal cache size allocation is verified by simulations, and it is shown to play a more significant role when the file popularity distribution is less concentrated.
Xi Peng 0006, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief
ICC1
2015 Backhaul-Aware Caching Placement for Wireless Networks
abstract
As the capacity demand of mobile applications keeps increasing, the backhaul network is becoming a bottleneck to support high quality of experience (QoE) in next-generation wireless networks. Content caching at base stations (BSs) is a promising approach to alleviate the backhaul burden and reduce user-perceived latency. In this paper, we consider a wireless caching network where all the BSs are connected to a central controller via backhaul links. In such a network, users can obtain the required data from candidate BSs if the data are pre-cached. Otherwise, the user data need to be first retrieved from the central controller to local BSs, which introduces extra delay over the backhaul. In order to reduce the download delay, the caching placement strategy needs to be optimized. We formulate such a design problem as the minimization of the average download delay over user requests, subject to the caching capacity constraint of each BS. Different from existing works, our model takes BS cooperation in the radio access into consideration and is fully aware of the propagation delay on the backhaul links. The design problem is a mixed integer programming problem and is highly complicated, and thus we relax the problem and propose a low-complexity algorithm. Simulation results will show that the proposed algorithm can effectively determine the near-optimal caching placement and provide significant performance gains over conventional caching placement strategies.
Xi Peng 0006, Juei-Chin Shen, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM1
2014 Joint data assignment and beamforming for backhaul limited caching networks
abstract
Caching at wireless access points is a promising approach to alleviate the backhaul burden in wireless networks. In this paper, we consider a cooperative wireless caching network where all the base stations (BSs) are connected to a central controller via backhaul links. In such a network, users can get the required data locally if they are cached at the BSs. Otherwise, the user data need to be assigned from the central controller to BSs via backhaul. In order to reduce the network cost, i.e., the back-haul cost and the transmit power cost, the data assignment for different BSs and the coordinated beamforming to serve different users need to be jointly designed. We formulate such a design problem as the minimization of the network cost, subject to the quality of service (QoS) constraint of each user and the transmit power constraint of each BS. This problem involves mixed-integer programming and is highly complicated. In order to provide an efficient solution, the connection between the data assignment and the sparsity-introducing norm is established. Low-complexity algorithms are then proposed to solve the joint optimization problem, which essentially decouple the data assignment and the transmit power minimization beamforming. Simulation results show that the proposed algorithms can effectively minimize the network cost and provide near optimal performance.
Xi Peng 0006, Juei-Chin Shen, Jun Zhang 0004, Khaled Ben Letaief
PIMRC1