Yuchao Zhang 0004

dblp:42/7674-4 · DBLP profile ↗
← Back
57ranked-venue papers
15as first author
45since 2021 · last 2026
0000-0002-0135-8915ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 46 · 11 first-author · 36 since 2021Systems, architecture and hardware · 5 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LCMP: Distributed Long-Haul Cost-Aware Multi-Path Routing for Inter-Datacenter RDMA Networks
abstract
RDMA-empowered cloud services are gradually deployed across datacenters (DCs) with multiple paths, which exhibit new properties of path asymmetry, delayed congestion signals, and simultaneous flow routing collisions, and further fail existing routing methods.
Dong-Yang Yu 0001, Yuchao Zhang 0004, Jun Wang 0178, Wenfei Wu, Haipeng Yao, Wendong Wang 0003, Ke Xu 0002
EuroSys2
2026 Fairness-Oriented Strategies for Video Streaming Competition in Shared Network Environments
Yuchao Zhang 0004, Xiaoxi Xue, Zeming Gao, Ye Tian 0008, Haipeng Yao, Wendong Wang 0003
ICC1
2026 CoPHo: Classifier-guided Conditional Topology Generation with Persistent Homology
abstract
The structure of topology underpins much of the research on performance and robustness, yet available topology data are typically scarce, necessitating the generation of synthetic graphs with desired properties for testing or release. Prior diffusion-based approaches either embed conditions into the diffusion model, requiring retraining for each attribute and hindering real-time applicability, or use classifier-based guidance post-training, which does not account for topology scale and practical constraints. In this paper, we show from a discrete perspective that gradients from a pre-trained graph?level classifier can be incorporated into the discrete reverse diffusion posterior to steer generation toward specified structural properties. Based on this insight, we propose Classifier-guided Conditional Topology Generation with Persistent Homology (CoPHo), which builds a persistent homology filtration over intermediate graphs and interprets features as guidance signals that steer generation toward the desired properties at each denoising step. Experiments on four generic/network datasets demonstrate that CoPHo outperforms existing methods at matching target metrics, and we further validate its transferability on the QM9 molecular dataset. The code is available at https://github.com/Lrbomchz/CoPHo.
Gongli Xi, Ye Tian 0008, Mengyu Yang, Yuchao Zhang 0004, Xiangyang Gong, Xirong Que, Wendong Wang 0003
KDD (1)5
2026 Packet-Level DDoS Data Augmentation Using Dual-Stream Temporal-Field Diffusion
Gongli Xi, Ye Tian 0008, Yannan Hu, Yuchao Zhang 0004, Yapeng Niu, Xiangyang Gong
SECON4
2026 RepLLM: Toward Automatically Reproducing Network Research Results
abstract
Result reproduction of computer networking research is challenging as the scarcity of open-source implementations and the complexity of heterogeneous system architectures. Even though Large Language Models have demonstrated potential in code generation, existing code generation frameworks often fail to address the long-context constraints and intricate logical dependencies, which are vital in reproducing network systems from academic papers. Thus, we introduce RepLLM, an end-to-end multi-agent framework designed to automate code reproduction from paper content. RepLLM features a collaborative architecture comprising four specialized agents—Content Parsing, Architecture Design, Code Generation, and Audit & Repair, which are coordinated through Shared Memory mechanism to ensure global context consistency. With the enhancement of Structured Chain-of-Thought LLM reasoning and a sandbox-isolated static-dynamic debugging methodology, our framework effectively resolves semantic discrepancies and runtime errors, thereby improving reliable reproductions. Extensive evaluations on representative papers in top conferences demonstrate that RepLLM outperforms state-of-the-art system-level LLM frameworks in generating compile-ready and logically correct systems. Our results show that, with the aid of RepLLM, we can reproduce 95% of the original benchmarks within approximately two hours while reducing token consumption by up to 10% compared with state-of-the-art baselines.
Yining Jiang, Yunxin Xu, Wenyun Xu, Yufan Zhu, Tangtang He, Letian Zhu, Qingyu Song 0002, Lizhao You, Lu Tang 0004, Wanjian Feng, Yuchao Zhang 0004, Linghe Kong, Qiao Xiang, Jiwu Shu
SIGCOMM14
2026 Towards Efficient Verification of Distributed In-Network Computing Programs
Mingyuan Song, Huan Shen, Jinghui Jiang, Qingyu Song 0002, Yuchao Zhang 0004, Wanjian Feng, Fei Yuan 0001, Yitao Xing, Wenjia Wei, Qiao Xiang, Jiwu Shu
SIGCOMM7
2026 EPIC: Abstraction and Polymorphism of In-Network Collectives on Ethernet
Yitao Yuan, Jianglong Nie, Tianyu Bai, Ruizhe Zhou, Siyuan Cao, Xujie Fan, Yuchen Xu 0003, Junkai Chen, Chenqi Zhao, Nengyuan Zhang, Shaoke Fang, Jiangyuan Chen, Yuanfeng Chen, Zhan Wang 0003, Yuchao Zhang 0004, Yang Liu 0038, Xiangrui Yang 0002, Xiaohe Hu, Limin Xiao 0001, Weifeng Zhang 0003, Yazhu Lan, Jianbo Dong, Binzhang Fu, Wenfei Wu
SIGCOMM17
2026 Mico: efficient query scheduling for multi-cloud deployed LLM inference service
Peizhuang Cong, Tong Yang 0003, Yuchao Zhang 0004, Wendong Wang 0003, Ke Xu 0002
Sci. China Inf. Sci.3
2026 I2BGP: A Privacy-Preserving Intra-AS State-Assisted Inter-AS Routing Scheme
abstract
BGP is the most widely employed inter-AS routing protocol, connecting millions of ASes worldwide. While it is possible to select the egress for outgoing flows based on administrators’ configurations, such schemes are localized due to the privacy of the intra-AS network state. TheAS_Pathfield of BGP records all crossed ASes, which can be used to prevent routing loops and select paths,i.e., selecting the minimum AS-hop path among available paths. Although this scheme is simple, effective, and offers a certain degree of global perspective, selecting paths at AS granularity ignores the transmission performance within each intra-AS, which may result in selecting non-optimal routing paths. To enable the use of private intra-AS data for inter-AS routing, we proposed a privacy-preserving intra-AS state-assisted inter-AS routing scheme, which can select optimal inter-AS paths without disclosing specific intra-AS state data. Specifically, we added an additional BGP header field to carry path performance features and designed a three-step data masking mechanism to protect intra-AS state data, enabling the selection of inter-AS paths with intra-AS state awareness. I2BGP has been deployed in the Greater Bay Area Future Network and a large-scale network simulator based on real network topologies. The results show that I2BGP outperforms BGP in terms of specified forwarding hops, delay, and bandwidth metrics.
Peizhuang Cong, Yuchao Zhang 0004, Jun Wang 0178, Wendong Wang 0003, Tong Yang 0003, Dan Li 0001, Ke Xu 0002
IEEE Trans. Netw.2
2026 TitanLog: Hierarchical and Elastic Logging for High-Speed Network Data Stream
abstract
Logging network traffic plays a crucial role as it serves as the foundation for various network applications. As network scale continues to expand, contemporary network traffic becomes increasingly high-speed, high-volume, and dynamic. This growth poses challenges to traditional server-based solutions. In this paper, we proposeTitanLog, ahierarchicalandelasticlogging system designed specifically for large-scale network traffic. TitanLog utilizes thehierarchical loggingmethodology, which aims to identify the importance of each packet in real-time and log packet data of different importance at different levels. To enhance efficiency, we propose a co-design of the emerging programmable switch and the server, incorporating sketches and RDMA to boost performance. To achieve elasticity, we design mechanisms for run-time adjustments and monitoring for resource insufficiency. TitanLog possesses the capability to switch between these modes at run-time. We fully implement TitanLog on a testbed and conduct extensive evaluations. The experimental results demonstrate that TitanLog supports logging of 100Gbps traffic with a zero packet loss rate and reduces the log volume by up to 96.28%.
Yuanpeng Li 0002, Xian Niu, Yikai Zhao 0001, Tong Yang 0003, Yannan Hu, Yuchao Zhang 0004, Xiangwei Deng, Qiuheng Yin, Ruwen Zhang, Yisen Hong, Kaicheng Yang 0001, Ruijie Miao, Kun Meng, Dahui Wang, Yong Cui 0001
IEEE Trans. Netw.6
2025 Understanding DNS Failures at Scale: Fine-grained Classification and Root Cause Analysis
Shaoxuan Yun, Yuchao Zhang 0004, Yingqiang Wang, Changhua Pei
APNet2
2025 Temporal Quality as a Metric: The MORS Routing Protocol for Model Training
Chenyue Zheng, Yuchao Zhang 0004, Wenfei Wu, Zhuo Jiang, Jianglong Nie, Wendong Wang 0003
APNet2
2025 Not All Gradients Are Equal: Dual Transport and Queue Dropping Strategy Guided by Spatiotemporal Gradient Importance
abstract
Distributed training has become a cornerstone of large-scale deep learning, yet gradient synchronization remains a critical bottleneck—particularly in high packet loss environments such as wide-area networks (WANs), where frequent retransmissions lead to long-tail latency and degraded training performance. In this paper, we propose a novel approach that leverages spatiotemporal gradient importance modeling to optimize both the transmission protocol and switch queue dropping strategy. Our method dynamically evaluates the importance of each gradient based on its position in the model layers and the current training stage, enabling adaptive packet loss control. In terms of implementation, we integrate a hybrid UDP/TCP transmission protocol: UDP ensures high-throughput gradient delivery, while TCP carries control signals to provide reliable loss feedback. Additionally, our system uses DSCP marking to map gradients to different switch queues, with each queue configured with a distinct WRED (Weighted Random Early Detection) loss threshold that is periodically updated based on the dynamically assessed gradient importance. Extensive experiments demonstrate that our model effectively captures the intrinsic spatiotemporal variations in gradient behavior and significantly reduces long-tail latency while maintaining model accuracy. This provides a robust framework for optimizing gradient synchronization and network resource allocation in distributed deep learning systems.
Zhichuan Zuo, Ye Tian 0008, Zeming Gao, Yuchao Zhang 0004, Xiangyang Gong, Wendong Wang 0003
GLOBECOM4
2025 Towards Feature-Consistent Parameter Collaboration for Personalized Federated Learning
abstract
Personalized federated learning (PFL) aims to improve the performance of the local model on each client with the non-IID data among different clients. This paper introduces FedFPC, a PFL method that allows effective and robust parameter-wise collaboration to achieve outperforming performance. Stem from the idea that similar clients should share more consistent feature representation and benefit more from each other, two strategies are designed to ensure feature consistency during training. First, we present an Attention-Guided Critical Parameter Selection strategy for critical parameter selection, which utilizes the attention prior from current expressive all-purpose features to identify the parameter with the most contribution to feature representation for the local data. Then, a Feature-Consistent Parameter Collaboration strategy is proposed to provide robust parameter collaboration for the local model with help from feature-consistent clients, which obtain similar feature representations to the target client. Experimental results demonstrate that FedFPC stands out by its superior performance in various PFL tasks compared to state-of-the-art methods, meanwhile with better robustness in diverse complex scenarios.
Jiahe Li 0007, Yuchao Zhang 0004, Wendong Wang 0003
ICASSP3
2025 Resolving Congestion Packet Losses for Small Flows in Datacenter Networks with Link-local Retransmission
abstract
With the increasing demand for high-performance computing (HPC) and artificial intelligence (AI), the transmission rates of data center networks have rapidly increased. However, this rise in link rate has not been accompanied by a proportional increase in switch buffer sizes, leading to an increase in the number of small flows and exacerbating buffer overflow problems. Traditional TCP and PFC mechanisms are ineffective in alleviating the flow completion time (FCT) issues of small flows. To address this, this paper proposes an improved solution based on the Link Local Retransmission (LLR) technique, called LLR-CoLoR. The approach works by backing up packets during congestion and precisely retransmitting lost packets, thereby avoiding the negative impacts of the PFC mechanism. This significantly reduces the FCT of small flows and improves the overall performance of data center networks. Large-scale topology experiments conducted on NS3 show that LLR-CoLoR reduces the FCT slowdown of small flows by more than 4X compared to other algorithms.
Jun Wang 0178, Yuchao Zhang 0004, Wendong Wang 0003
ICCCN2
2025 Extendible RDMA-Based Remote Memory KV Store with Dynamic Perfect Hashing Index
abstract
Perfect hashing is a special hashing function that maps each item to a unique location without collision, which enables the creation of a KV store with small and constant lookup time. Recent dynamic perfect hashing attains high load factor by increasing associativity, which impacts bandwidth and throughput. This paper proposes a novel dynamic perfect hashing index without sacrificing associativity, and uses it to devise an RDMA-based remote memory KV store called CuckooDuo. CuckooDuo simultaneously achieves high load factor, fast speed, minimal bandwidth, and efficient expansion without item movement. We theoretically analyze the properties of CuckooDuo, and implement it in an RDMA-network based testbed. The results show CuckooDuo achieves 1.9~17.6x smaller insertion latency and 9.0~18.5x smaller insertion bandwidth than prior works.
Zirui Liu 0002, Xian Niu, Wei Zhou 0077, Yisen Hong, Zhouran Shi, Tong Yang 0003, Yuchao Zhang 0004, Yuhan Wu 0001, Yikai Zhao 0001, Zhuochen Fan, Bin Cui 0001
ICDE7
2025 A CoT Reasoning-Based Computation and Network Resource Deployment Intent Translation Framework
abstract
Advancements in distributed machine learning have led to a demand for engineers with knowledge in AI, hardware, and networking. Research aims to automate deployment based on user intent to improve development efficiency. Large Language Models (LLMs) assist in translating user intent for resource allocation but face challenges in translation accuracy and contextual integrity. The Chain-of-Thought (CoT) improves LLM reasoning by breaking down problems into sequential steps. However, the reasoning steps and dependencies between the model requirements and the computation and network configurations are complex, requiring the selection of an appropriate thought path to construct the CoT framework. To address these issues, we propose a computation and network resource deployment intent translation framework based on CoT reasoning and create a benchmark for user intent in distributed learning. This framework translates user intent into hardware and network configurations, using LLMs to optimize translation accuracy through logical reasoning. Experiments on four LLM models show significant accuracy improvements with CoT compared to traditional prompts. The feasibility of our framework has been validated through its implementation and testing on a real-world testbed.
Jialu Du, Lintong Du, Zeming Gao, Yuchao Zhang 0004, Ye Tian 0008, Xiangyang Gong
ICNP4
2025 Sub-RTT Congestion Control for Inter-Datacenter Networks
abstract
With the explosive growth in the scale and complexity of large language models (LLMs), there is an urgent need to extend training and inference workloads from within a single data center to across multiple data centers. However, this also introduces new challenges for network transport protocols. To address these issues, we propose SRCC (Sub-RTT Congestion Control), a method designed for inter-datacenter networks. Specifically, SRCC introduces a flowset-based mechanism along with shared node tables, enabling Datacenter Interconnect (DCI) switches to be aware of the path status of each flow. By leveraging information shared among different flows, SRCC can accurately adjust the sending rate at a sub-RTT timescale, thereby significantly improving network performance. Building on this approach, we design detailed mechanisms to address the following challenges: (1) applying INT technology in wide-area networks; (2) acquiring INT information with low overhead; and (3) achieving precise congestion window adjustments under sub-RTT perception.We conducted large-scale simulations using NS3, and the experimental results show that our scheme reduces the average FCT slowdown by 44.17% and 53.86% compared to HPCC and DCTCP, respectively.
Jun Wang 0178, Yuchao Zhang 0004, Gaoxiong Zeng, Chenyue Zheng, Wendong Wang 0003, Haipeng Yao
ICNP2
2025 MORS: Traffic-Aware Routing based on Temporal Attributes for Model Training Clusters
abstract
To train large AI models, clusters are constructed with abundant connectivity and bandwidth; but the commodity protocol ECMP and recent proposals fail to fully utilize the network bandwidth for AI traffic pattern. As model training jobs and AI clusters exhibit a predictable and periodic traffic pattern, so in this paper, we propose a MOdel training Routing System — MORS — for traffic routing in AI clusters. MORS defines temporal attributes to characterize the periodic traffic pattern of flows and network links, and temporal quality to quantify whether a path could deliver a flow quickly in the near future. MORS runs In-band Network Telemetry (INT) to collect temporal attributes of the network, and periodic analysis to extend the collected attributes in the time domain. Based on the time series of link utilization and latency, MORS computes the temporal quality of candidate paths. It enforces high-quality path selection while maintaining compatibility with commodity ECMP by manipulating the source UDP port to ensure the flow complies with the target path in the ECMP protocol. MORS is light-weight and readily deployable in the RDMA commodity cluster. Our prototype and experiments demonstrate that MORS achieves performance comparable to adaptive routing and delivers up to 14% and 50% better FCT than PLB and ECMP, respectively.
Yuchao Zhang 0004, Chenyue Zheng, Wenfei Wu, Zhuo Jiang, Huichen Dai, Jianglong Nie, Wendong Wang 0003
ICNP1
2025 DeCross: Toward Accountable and Efficient Cross-Chain Collaboration in IIoT
abstract
Blockchain-based industrial Internet-of-Things (IIoT) systems have seen rapid adoption and development in recent years. The increasing diversity of IIoT application scenarios is driving the growth of multichain ecosystem, making cross-chain communication a key issue in multichain collaboration. However, existing centralized cross-chain architectures risk derailing the blockchain's trust-free decentralization and suffer from single point failure. The design of decentralized cross-chain collaboration mainly faces two key challenges. First, implicit cross-chain accountability, which is caused by collusion among malicious distributed participants and, thus, compromises cross-chain security. Second, low cross-chain efficiency, which is induced by the highly dynamic environment in practical cross-chain networks, i.e., changing memberships, and adaptive attacks to corrupt honest nodes. In this article, we propose a cross-chain consensus protocolDeCrossto solve the abovementioned problems.DeCrossachieves decentralized cross-chain collaboration with explicit accountability and high efficiency. Specifically, we audit participants with a succinct auditable data object constructed from the protocol to hold nodes accountable for misbehaving. Furthermore, we propose a parallel processing workflow that leverages both CPUs and GPUs to guarantee efficient and stable cross-chain communication. Finally, we implement a prototype based on the hyperledger fabric with both local and geo-distributed clusters. Our extensive experiments show thatDeCrossachieves 44% better throughput over the existing cross-chain approaches.
Yuchao Zhang 0004, Ke Xu 0002
IEEE Trans. Ind. Informatics1
2024 HybridCom: Improve Federated Learning Efficiency on Unstable Data
abstract
Federated learning (FL) has made significant advancements in recent years. However, its efficiency on unstable distributed data remains a critical challenge. This stems from oversights in existing FL frameworks regarding the instability of global and client private data distributions or their assumption of stable distributions over time. To address this challenge, we present HybridCom, an efficient FL framework on unstable data. The core concept of HybridCom is to adjust the client participation probability based on their contributions to adapting the global model to data distribution changes. This is achieved through a hybrid contribution indicator that includes a performance-based client-side indicator and a gradient-based server-side indicator. Based on the results of the contribution indicator, HybridCom integrates a probabilistic participation controller to dynamically adjust the participation probability of each client during the FL process. By utilizing Hybrid-Com, clients undergoing data distribution changes that are not perceived by the global model have a higher probability of participating. This makes HybridCom more efficient in adapting to unstable data distributions. The experimental results demonstrate that HybridCom surpasses the baseline models, achieving an approximate 1.3% improvement across diverse simulation settings and communication resource constraints.
Yuchao Zhang 0004, Xiangyang Gong, Wendong Wang 0003
ICC2
2024 Gradient Rotation Unit for Non-I.I.D. Federated Learning
abstract
Federated Learning (FL) enables collaborative training of a global model without exposing raw data by aggregating local updates from clients. However, the convergence efficiency on non-i.i.d. data remains challenging, leading to performance loss and resource bottleneck. Meanwhile, the nature of non-i.i.d. challenge are not yet fully understood. In this paper, we first reveal that non-i.i.d. data leads to server-side multi-objective aggregation conflict challenge which hampers the convergence efficiency. We then propose Federated Gradient Rotation Unit (FGRU), a simple yet general approach to mitigate this challenge by deliberately aligning optimization trajectories across clients. FGRU is a server-side plugin that rotates gradients to other before aggregation. On a series of challenging non-i.i.d. FL tasks, FGRU leads to significant gains in convergence efficiency and performance. Experimental results demonstrate that FGRU improves inference accuracy by approximately 3-5% and accelerates convergence by 2-3X on simulated non-i.i.d. data using MNIST, CIFAR-10, and Fashion-MNIST datasets. The results also show that FGRU is model agnostic and can be combined with existing non-i.i.d. FL frameworks such as SCAFFOLD and FedProx to further improve performance.
Yuchao Zhang 0004, Xiangyang Gong, Wendong Wang 0003
IJCNN2
2024 Nexus: Efficient and Conflict-Equivalent DAG-Based Permissioned Blockchain Sharding
abstract
Sharding is one of the most promising solutions to tackle the well-known scalability issue in blockchain by dividing the network into multiple parallel committees. However, inefficient and insecure handling of cross-shard transactions can lead to a decline in the system service quality. Existing permissioned blockchains are mainly based on the classic two-phase commit (2PC) and two-phase locking (2PL) from traditional distributed databases to implement cross-shard transactions with deterministic safety, but this also leads to issues such as decreased system throughput and frequent transaction aborts. In this paper, we introduce Nexus, an efficient and conflict-equivalent sharded permissioned blockchains to overcome the aforementioned challenges. The core idea of Nexus is to construct a directed acyclic graph (DAG) among all shards to establish a partial order of blocks, while leaving transaction dissemination and final ordering to be completed in parallel by each shard. Nexus alleviate the overhead of global ordering through a efficient consensus process, and ensure that all transactions with overlapping read and write sets can be committed in the same relative order across different shards. A prototype of Nexus is implemented and evaluated, and experimental results show that Nexus achieves approximately 70% and 50% throughput improvement over AHL and SharPer under default settings.
Yuchao Zhang 0004
IWQoS1
2024 ActiveDNS: Is There Room for DNS Optimization Beyond CDNs?
abstract
Domain Name System (DNS) converts domain names into IP addresses. The specific IP it returns to a client has significant implications for optimizing user experiences on web services. The establishment of Content Distribution Networks (CDNs) facilitates the spread of content across diverse cache servers, thereby enabling quick responses to user requests. But not all distributed Internet services have sufficient budget or resources to deploy CDN servers. For these services, is there a cost-efficient way to enhance users’ access experience to them? In this paper, we build a data-driven DNS resolver system named ActiveDNS. We conduct a historical analysis of billions of DNS request-response logs from a public DNS resolver. Inspired by these findings, we probed dozens of public DNS servers to obtain available IPs behind carefully selected target domain names and measure their performance metrics. We built a domain ip performance database of 2.9 million records that can be easily incorporated into the open-source DNS framework BIND via DLZ technology. ActiveDNS has been deployed in real-world scenarios and the evaluations reveal promising results. Compared to currently used DNS resolver, the 75th service round-trip time (RTT) for the IP returned by ActiveDNS decreased from 67 ms to 43 ms.
Changhua Pei, Yingqiang Wang, Guo Chen 0001, Yuchao Zhang 0004, Gaogang Xie
LCN6
2024 Galaxy: A Scalable BFT and Privacy-Preserving Pub/Sub IoT Data Sharing Framework Based on Blockchain
abstract
The emergence of the Internet of Things (IoT) technology in recent years has led to a considerable amount of data to be shared across different organizations. The publish and subscribe (Pub/Sub) paradigm, with its asynchronous, one-to-many, and decoupling characteristics, is considered to be a promising communication model in IoT. However, designing a Pub/Sub framework for IoT data sharing confronts two challenges: 1) Byzantine faults and 2) privacy concerns. Byzantine nodes that are subjectively malicious or hacked by attackers may discard or forge data in the broker network composed of untrusted IoT organizations. Unauthorized brokers or clients may try to obtain the content of publications or subscriptions, thus violating the IoT data privacy. Existing works have limitations in terms of relatively low scalability and high overhead in tackling these two challenges. In this article, we propose Galaxy, a blockchain-based Pub/Sub IoT data sharing framework. To achieve Byzantine fault-tolerant (BFT) Pub/Sub, Galaxy adopts sharding to improve scalability and achieve efficient BFT Pub/Sub workflow within each shard with a novel leader rotation scheme. In attaining privacy-preserving Pub/Sub, a secret key sharing and encrypted Pub/Sub scheme is designed in Galaxy to achieve low overhead without breaking the decoupling of the system. We implemented a prototype of Galaxy and deployed it on Alibaba Cloud for experimental evaluation. The experiment results show the feasibility and efficiency of Galaxy.
Yuchao Zhang 0004, Ning Zhang 0007, Zibin Zheng, Ke Xu 0002
IEEE Internet Things J.1
2024 FLAIR: A Fast and Low-Redundancy Failure Recovery Framework for Inter Data Center Network
abstract
Due to the fast developments of 5G and IoT technologies, Inter-Datacenter (Inter-DC) networks are facing unprecedented pressure to duplicate large volumes of geographically distributed user data in a real-time manner. Meanwhile, with the expansion of Inter-DC networks scale, link/node failures also become increasingly frequent, negatively affecting the data transmission efficiency. Therefore, link failure recovery methods become of utmost importance. Many works investigated fast failure recovery, yet none of them consider the deployment overhead of such recovery schemes. While in this paper, we found that the side-effect of deploying recovery strategies and the future availability of the recovered transmissions are also crucial for fast recovery. So we propose a fast and low-redundancy failure recovery framework, FLAIR, which consists of a fast recovery strategy FRAVaR and a redundancy removal algorithm ROSE. FRAVaR takes full consideration of deployment overhead by minimizing shuffle traffic. On its base, ROSE regularly eliminates the cumulative rerouting redundancy by removing unnecessary routing updates. The experiment results on 4 realistic network topologies show that FLAIR successfully reduces up to 48.2% deployment overhead compared with the state-of-the-art solutions, and thus reduces up to 70.2% recovery speed and improves up to 36% network utilization.
Yuchao Zhang 0004, Haoqiang Huang, Ahmed M. Abdelmoniem, Gaoxiong Zeng, Chenyue Zheng, Xirong Que, Wendong Wang 0003, Ke Xu 0002
IEEE Trans. Cloud Comput.1
2023 ABC: Adaptive Bitrate Algorithm Commander for Multi-Client Video Streaming
abstract
With the improvement of live streaming technology, ensuring high QoE and fairness of different ABR algorithm clients sharing the same LAN is becoming a pressing issue. However, aggressive and conservative algorithm will make different bitrate adjustment decisions when they share network resources, which leads to unfairness. In this poster, we proposed a regulation mechanism ABC, adjusting the sensitive parameters such as latency, delay and buffer, to coordinate overall system QoE by 68% and improve the fairness problem.
Xiaoxi Xue, Yuchao Zhang 0004
APNet2
2023 FaCa: Fast Aware and Competition-Avoided Balancing for Data Center Network
Haiyang Jiang 0005, Yuchao Zhang 0004, Haoqiang Huang, Xirong Que, Zhuo Jiang, Wendong Wang 0003
ICA3PP (6)2
2023 Jointly Optimal Routing and Caching with Bounded Link Capacities
abstract
We study a cache network in which intermediate nodes equipped with caches can serve requests. We model the problem of jointly optimizing caching and routing decisions with link capacity constraints over an arbitrary network topology. This problem can be formulated as a continuous diminishing-returns (DR) submodular maximization problem under multiple continuous DR-supermodular constraints, and is NP-hard. We propose a poly-time alternating primal-dual heuristic algorithm, in which primal steps produce solutions within 1 - approximation factor from the optimal. Through extensive experiments, we demonstrate that our proposed algorithm significantly out-performs competitors.
Yuchao Zhang 0004, Stratis Ioannidis, Jon Crowcroft
ICC2
2023 Grandet: Cost-aware Traffic Scheduling without Prior Knowledge in SD-WAN
abstract
The rapid growth of traffic demands on wide-area networks (WANs) has resulted in escalated transmission costs for cross-national enterprises. Many researchers have proposed traffic scheduling methods that can effectively reduce transmission costs and improve network performance. However, the majority of research in this field assumes that traffic demands and network link quality are known in advance, disregarding the impact of information agnostic. While some works try to obtain this knowledge through prediction, they lack awareness of prediction errors, which makes it difficult for their scheduling strategies to achieve theoretical results. In this paper, we propose a novel scheduler Grandet that aims to reduce transmission costs without any prior knowledge. First, instead of requiring prior knowledge or accurate prediction, Grandet determines the intervals of flow sizes and link quality parameters through confidence-based Bootstrap method combined with neural network model, thus quantifying the uncertainty of these information. Then, we design a cost-aware online traffic scheduling framework using the uncertainty intervals from interval determination to optimize the cost minimization problem. Through rigorous theoretical analysis, we prove the approximate optimality of Grandet in minimizing transmission costs. Trace-driven and large-scale simulations show that Grandet successfully reduces transmission costs by over 23%, reduces deadline miss rate by over 31%, and reduces Service Level Agreement (SLA) dissatisfaction rate by over 37%.
Yuchao Zhang 0004, Huahai Zhang, Peizhuang Cong, Wendong Wang 0003, Ke Xu 0002
IWQoS1
2023 FRAVaR: A Fast Failure Recovery Framework for Inter-DC Network
abstract
Along with the development of 5G and IoT technologies in recent years, Inter Data Center (Inter-DC) network is facing an explosive growth of geographically distributed user data, which needs to be duplicated among DCs in a real-time manner. Transmission-based applications require high availability that is going beyond 99.99%. However, with the expansion of Inter-DC network scale, link failures are also growing, which seriously affects data transmission efficiency, so fast link failure recovery is then urgently needed. Many previous works have been done to achieve fast failure recovery, but most of them ignore two key points, 1) the cost of deploying recovery strategies, and 2) the side-effect of re-transmission to network availability. These two factors make the existing failure recovery process too slow to be practical in real-time online industrial environments. To achieve realistic fast recovery from Inter-DC network failures, we propose a failure recovery framework FRAVaR, which achieves high network availability with very little deployment overhead. Particularly, FRAVaR reduces the deployment overhead by a novel incremental routing strategy to isolate link failures. In other words, it only needs to shuffle a tiny amount of traffic within a small failure isolation domain. On this base, FRAVaR further adopts a risk assessment theory named Value-at-Risk (VaR) to control flow re-transmission. We implement a prototype of FRAVaR and conduct a series of experiments on 4 real InterDC network topologies (ATT North America, IBM, GlobalCenter, AGIS). Experiment results show that FRAVaR outperforms state-of-the-art solutions on the recovery speed by 70.2%.1
Haoqiang Huang, Yuchao Zhang 0004, Qiao Xiang, Wendong Wang 0003, Xirong Que, Ke Xu 0002
WCNC2
2023 A delayed eviction caching replacement strategy with unified standard for edge servers
Pengmiao Li, Yuchao Zhang 0004, Huahai Zhang, Wendong Wang 0003, Ke Xu 0002
Comput. Networks2
2023 DIT and Beyond: Interdomain Routing With Intradomain Awareness for IIoT
abstract
Along with the ever-increasing amount of data generated from industrial devices, the cross domain [also known as autonomous systems (ASs)] data transmission problem has attracted more and more attention in the Industrial Internet of Things (IIoT). As mature and widely used interdomain routing protocols, border gateway protocol-based solutions often take the number of domains (i.e., AS hops) of each path as a criterion to make routing decisions, which is simple and effective. However, such protocols can only meet the reachability requirements while ignoring the performance requirements. That is, the path with the minimum AS hops will be selected to carry flows, even if the actual performance of this path does not meet the transmission requirements due to the unawareness of intradomain information on that path. But it is not impractical to directly access intradomain information for making better routing decisions given data privacy concerns. In this article, we propose M-DIT, which can make interdomain routing decisions with the assistance of desensitized intradomain information for multiple-requirement transmissions. To do so, we design a homomorphic encrypted-based private number comparison scheme to export intradomain information securely and, thus, assist in routing decisions. The results of some experiments based on five real topologies (ATMnet,Claranet,Compuserve,NSFnet, andPeer1) with thousands of interdomain flows demonstrate that M-DIT reduced flow completion time by about 60% or selected high bandwidth paths flexibly for interdomain routing for IIoT scenarios.
Peizhuang Cong, Yuchao Zhang 0004, Wendong Wang 0003, Xiangyang Gong, Tong Yang 0003, Dan Li 0001, Ke Xu 0002
IEEE Internet Things J.2
2022 A Scalable Nested Blockchain Framework with Dynamic Node Selection Approach for IoT
abstract
A high level of scalability is needed to support the large-scale Internet-of-Things (IoT) networks. To address the issue of distributed trust in different IoT devices, blockchain technology can be effectively used to safely manage IoT data due to its ability to provide transactions traceability and security. However, massive real-time IoT application data has brought huge challenges to the scalability of the integration framework of blockchain and IoT. This paper proposes a nested-chain architecture, which consists of one main chain and multiple sub chains to address the aforementioned challenges. The main chain stores identity credential used for distributed identity (DID) management, while the sub chain stores the IoT data. A notary module that involves access nodes from both chains is designed for cross-chain transactions. In addition, considering the transaction information, node characteristics, and network status, we further introduce a node selection algorithm based on Graph Convolutional Network (GCN), which can effectively reduce the cost of cross-chain communications. We implement and evaluate a prototype of our framework on the Hyperledger Fabric platform to demonstrate its feasibility and superiority. The analyzed results have shown that our proposed framework outperforms traditional schemes, by reducing system latency up to 23.2% and increasing system throughput up to 12.5%.
Yuchao Zhang 0004
IPCCC2
2022 Break the Blackbox! Desensitize Intra-domain Information for Inter-domain Routing
abstract
Along with the ever-increasing amount of data generated from edge networks, cross domain (also known as Autonomous Systems, AS) transmission problem has attracted more and more attention. As mature and widely used inter-domain routing protocols, BGP-based solutions often use the number of domains (i.e. AS hops) of each path to make inter-domain routing decisions, which is simple and effective, but usually can not get the optimal routing results due to the lack of real state/information within ASes. These protocols choose the path with less AS hops as the forwarding path, even if the total latency or cost of the domains on this path is higher. While to solve this problem, directly access to intra-domain information as the assistance to make routing decisions is impractical due to data privacy.In this paper, we propose DIT, which makes near-optimal inter-domain routing decisions with desensitized intra-domain information. To do so, we design a homomorphic encrypted-based private number comparison scheme to export intra-domain information securely and thus assist in routing decisions. We conduct a series of experiments according to five real network topologies with nearly 900 simulated flows, and the results show that DIT reduces the number of forwarding hops by about 45% in average and reduces flow completion time by about 60%.
Peizhuang Cong, Yuchao Zhang 0004, Wendong Wang 0003, Xiangyang Gong, Tong Yang 0003, Dan Li 0001, Ke Xu 0002
IWQoS2
2022 Chameleon: A Self-adaptive cache strategy under the ever-changing access frequency in edge network
Pengmiao Li, Yuchao Zhang 0004, Wendong Wang 0003, Weiliang Meng, Ke Xu 0002
Comput. Commun.2
2022 A&B: AI and Block-Based TCAM Entries Replacement Scheme for Routers
abstract
With the ever-increasing deployment of 5G and IoT, the number of end-hosts/terminals is increasing rapidly, so that routers have to cache more and more forwarding entries to guarantee communication reachability of these terminals, which makes Ternary Content Addressable Memory (TCAM)-based routers keep expanding resource requirements. However, the design and implementation of large-capacity TCAM-based routers are faced with such challenges: difficult circuit design, high production cost and energy consumption, thereby posing an urgent requirement on a lightweight TCAM that can still maintain those massive communication connections. In this paper, we aim to design a lightweight router with small storage requirement while still retaining the original communication connection performance, which is not straightforward due to the following two challenges: First, under the condition of massive sequential flow data, it’s difficult to accurately and timely select the entries to cache for a small capacity TCAM. Second, given the strict prefix matching principle, how to efficiently insert the selected entries into TCAM is also challenging. To address these problems, we propose A&B: an AI-based Routing entry prediction strategy (AIR) and a Block-based entry Insertion Tactic (BIT). AIR can precisely select entries by conducting accurate entry predictions, which converts dynamic flow-based prediction into stable and parallelizable entry-based prediction by decoupling spatio-temporal characteristics. BIT optimizes entry insertion by isolating TCAM into several blocks, thus eliminating the time-consuming entry movements. The experiment results based on real backbone traffic show that our lightweight A&B achieves comparable performance compared to the traditional schemes by using only 1/8 TCAM storage.
Peizhuang Cong, Yuchao Zhang 0004, Bin Liu 0001, Wendong Wang 0003, Zehui Xiong, Ke Xu 0002
IEEE J. Sel. Areas Commun.2
2021 TPA based content popularity prediction for caching and routing in edge-cloud cooperative network
abstract
The rapid development and application of 5G/B5G generate tremendous amount of traffic which in turn cause great burden for the corresponding transmission network. One typical way to address such challenge is to sink the content (e.g., 4K and 8K videos) from the remote cloud to the edge servers. In this case, how to efficiently visiting and getting these contents becomes a new problem, in which the cooperation between cloud and edge should be taken into consideration. In this regard, this work builds an edge and cloud cooperative routing and caching system which consists of three main modules of content popularity prediction, cooperative caching and cooperative routing. Specifically, the content prediction is designed by jointly leveraging the technologies of Long Short-Term Memory (LSTM) and Temporal Pattern Attention (TPA) to dig the traffic features and predict the future content popularity. Based on the prediction results and the technology of reinforce learning, the cooperative caching module designs both a reactive content replacement and an active content caching strategies. After that, the cooperative routing is carried out to help customers visiting and obtaining these content efficiently with the objective of minimizing the overhead. The experimental results indicate that the proposed methods outperform the state-of-the-art benchmarks in terms of the caching hit rate, the average throughput, the successful content delivery rate and the average routing overhead.
Bo Yi 0002, Fuliang Li, Yuchao Zhang 0004, Xingwei Wang 0001
GLOBECOM3
2021 A Deep Reinforcement Learning-based Routing Scheme with Two Modes for Dynamic Networks
abstract
With the development of communication and transmission technologies, more and more applications, like Internet of vehicles and tele-medicine, become more sensitive to network latency and accuracy, which requires routing schemes to be more efficient. In order to meet such urgent need, learning-based routing strategies emerges, with the advantages of high flexibility and accuracy. These strategies can be divided into two categories, centralized and distributed, enjoying the advantages of high precision and high efficiency, respectively. However, routing become more complex in dynamic network, where the link connections and access states are time-varying, so these learning-based routing mechanisms are required to be able to adapt to network changes in real time. In this paper, we designed and implemented both two of centralized and distributed reinforcement learning-based routing schemes (RLR-T). By conducting a series of experiments, we deeply analyzed the results and gave the conclusion that the centralized is better to cope with dynamic networks due to its faster reconvergence, while the distributed is better to handle with large-scale networks by its high scalability.
Peizhuang Cong, Yuchao Zhang 0004, Wendong Wang 0003, Ke Xu 0002, Ruidong Li 0001, Fuliang Li
ICC2
2021 FreeVM: A Server Release Algorithm in DataCenter Network
abstract
With the development of 5G access technology and the corresponding explosive growth of user requests, service providers have to activate more and more physical machines (PM) in cloud datacenters. This simple expansion of PMs results in not only low utilization of servers but also high maintenance cost. Existing researches that try to release unnecessary servers by migrating VMs always face a severe challenge-the large searching space for the optimal VM placement solution. In this paper, we prose a two-stage variable neighborhood searching (STVNS) algorithm, named FreeVM, which can significantly reduce the number of occupied servers. FreeVM works harmoniously with all kinds of VM placement schemes under various scenarios. We conducted an extensive series of experiments using real traces, and the results show that FreeVM can release at least 15% physical machines compared with existing solutions.
Shiyan Zhang, Yuchao Zhang 0004, Xiangyang Gong
ICC2
2021 AIR: An AI-based TCAM Entry Replacement Scheme for Routers
abstract
Ternary Content Addressable Memory (TCAM) is an important hardware used to store route entries in routers, which is used to assist routers to make fast decision on forwarding packets. In order to cope with the explosion of route entries due to massive IP terminals brought by 5G and the Internet of Things (IoT), today’s commercial TCAM has to keep the corresponding growth in capacity. But large TCAM capacity is causing many problems such as circuit design difficulties, production costs, and high energy consumption, so it is urgent to design a lightweight TCAM with small capacity while still maintains the original query performance.Designing such a TCAM faces two fundamental challenges. Firstly, it is essential to accurately predict the incoming flows in order to cache correct entries in limited TCAM capacity, but prediction on aggregated time-sequential data is challenging in the massive IoT scenarios. Secondly, the prediction algorithm needs to be real-time as the lookup process is in line-rate. In order to address the above two challenges, in this paper, we proposed a lightweight AI-based solution, called AIR, where we successfully decoupled the route entries and designed a parallel-LSTM prediction method. The experiment results under real backbone traffic showed that we successfully achieved comparable query performance by using just 1/8 TCAM size.
Yuchao Zhang 0004, Peizhuang Cong, Bin Liu 0001, Wendong Wang 0003, Ke Xu 0002
IWQoS1
2021 CRATES: A Cache Replacement Algorithm for Low Access Frequency Period in Edge Server
abstract
In recent years, with the maturity of 5G and Internet of Things technologies, the traffic in mobile network is growing explosively. To reduce the burden of cloud data centers and CDN network, edge servers that are closer to users are widely deployed, caching hot contents and providing higher Quality of Service (QoS) by shortening access latency. Storage resources on edge servers are much limited compared with CDN servers, so the research on cache replacement strategy of edge servers is critical to edge computing and storage area. Many efforts have been made to improve caching performance on edge servers. Existing caching strategies only focus on the high access frequency period to solve the caching problem, they ignore low access frequency period with two characteristics, including that hot contents are difficult to predict and hot topics usually change unstably, which makes it inefficient to improve the hit rate on edge servers.In this paper, we deeply analyzed the real traces from Chuang-Cache and found some specific user groups are playing more important roles than general users during low access frequency period, and the contents accessed by these specific user groups have a much higher possibility to become hot contents. Therefore, we firstly classify such users to core users, and treat others as common users. Then we adopt the principal component analysis algorithm to analyze the relationship between hot contents and core users. On this basis, we finally propose a hot contents pre-cache protection mechanism, which is a significant part of our cache replacement algorithm CRATES. To improve CRATES’s efficiency, we extract key part of historical data by designing a sliding window method. Through a series of experiments using real application data, we demonstrate that CRATES reaches about 98% in caching hit rate and outperforms the state-of-the-art algorithm LRB by 1.4X.
Pengmiao Li, Yuchao Zhang 0004, Huahai Zhang, Wendong Wang 0003, Ke Xu 0002
MSN2
2021 A deep reinforcement learning-based multi-optimality routing scheme for dynamic IoT networks
Peizhuang Cong, Yuchao Zhang 0004, Zheli Liu, Thar Baker, Hissam Tawfik, Wendong Wang 0003, Ke Xu 0002, Ruidong Li 0001, Fuliang Li
Comput. Networks2
2021 DND: Driver Node Detection for Control Message Diffusion in Smart Transportations
abstract
Along with the development of IoT and mobile edge computing in recent years, smart transportation holds great potential to improve road safety and efficiency. The network that carries smart transportation service is highly dynamic. Controllability has long been recognized as one of the fundamental properties of such temporal networks, which can provide valuable insights for the construction of new infrastructures, and thus is in urgent need to be explored. In this article, under the smart transportation scenario, we first disclose the controllability problem in Internet of Vehicles (IoV), and then design DND (Driver Node Detection) algorithm based on Kalman's controllability rank condition to analyze the controllability and control message diffusion in such a dynamic temporal network. Moreover, we use the control message diffusion efficiency as a metric to assist in selecting suitable driver nodes. At last, we conduct a series of experiments to analyze the controllability of the IoV network, and the results show the effects of vehicle density, speed, coverage radius on network controllability, and the efficiency of the control message diffusion algorithm and its feedback effect on driver nodes selection. These insights are critical for varieties of applications in the future smart transportation.
Peizhuang Cong, Yuchao Zhang 0004, Wendong Wang 0003, Ning Zhang 0007
IEEE Trans. Netw. Serv. Manag.2
2021 BDS+: An Inter-Datacenter Data Replication System With Dynamic Bandwidth Separation
abstract
Many important cloud services require replicating massive data from one datacenter (DC) to multiple DCs. While the performance of pair-wise inter-DC data transfers has been much improved, prior solutions are insufficient to optimize bulk-data multicast, as they fail to explore the rich inter-DC overlay paths that exist in geo-distributed DCs, as well as the remaining bandwidth reserved for online traffic under fixed bandwidth separation scheme. To take advantage of these opportunities, we present BDS+, a near-optimal network system for large-scale inter-DC data replication. BDS+ is an application-level multicast overlay network with a fully centralized architecture, allowing a central controller to maintain an up-to-date global view of data delivery status of intermediate servers, in order to fully utilize the available overlay paths. Furthermore, in each overlay path, it leverages dynamic bandwidth separation to make use of the remaining available bandwidth reserved for online traffic. By constantly estimating online traffic demand and rescheduling bulk-data transfers accordingly, BDS+ can further speed up the massive data multicast. Through a pilot deployment in one of the largest online service providers and large-scale real-trace simulations, we show that BDS+ can achieve 3- 5× speedup over the provider's existing system and several well-known overlay routing baselines of static bandwidth separation. Moreover, dynamic bandwidth separation can further reduce the completion time of bulk data transfers by 1.2 to 1.3 times.
Yuchao Zhang 0004, Xiaohui Nie, Junchen Jiang, Wendong Wang 0003, Ke Xu 0002, Youjian Zhao, Martin J. Reed, Kai Chen 0005, Guang Yao
IEEE/ACM Trans. Netw.1
2020 Auction-based High Timeliness Data Pricing under Mobile and Wireless Networks
abstract
Data is the cornerstone of intelligent algorithms such as deep learning, and the explosive development of mobile and wireless networks has prompted more devices to share data in time via the Internet. Meanwhile, data is highly time sensitive. It has been found that the value of data is becoming more and more critical to any application areas, significantly highlighting the importance of data pricing mechanisms in data transactions. Although traditional auction mechanisms for ordinary commodities are gradually becoming matures, they fail in the high timeliness data pricing market due to the following key challenges: Firstly, the value and price of the high timeliness data is ever changing with time, making existing mechanisms with fixed prices expired. Secondly, the price changing of such data is uncertain and dynamic, requiring the auction mechanisms to work stably under different price variations of the high timeliness data. To address these challenges, we for the first time innovatively propose an efficient auction mechanism for High Timeliness Data Pricing, namely HTDP. The newly proposed HTDP can maximize the profit of auctioneer in the high timeliness data transactions. And the key factor for HTDP's success is the consideration of the price changing in the high timeliness data, which fills the blank of traditional auction mechanisms in this area. We further evaluate the newly proposed HTDP on the overall auction profit, and compare the results with the benchmark. Experimental results demonstrate that HTDP not only achieves high profit under proper settings, but also is stable and efficient.
Yi Zhao 0011, Ke Xu 0002, Yuchao Zhang 0004
ICC4
2020 Software-Defined Networking-Assisted Content Delivery at Edge of Mobile Social Networks
abstract
With the explosive growth of mobile devices at the edge of mobile social networks (MSNs), the amount of the content that needs to be transmitted is exploded. Traditional content delivery mechanisms leverage only local information to make routing decisions, which results in both high latency and low delivery rate. Software-defined networking (SDN) is a novel network paradigm, the design philosophy of which could be applied to MSN for improving the content delivery performance. In this article, the centralized control thought of SDN is introduced into MSN to efficiently process social information. The classical routing algorithm of BubbleRap is improved from the perspective of network density, which is the basis of designing the sparse and dense routing mechanisms for MSNs. In addition, flexibly switching between these two routing mechanisms is implemented by a discriminating scheme, achieving efficient yet adaptive routing. The experimental results show that the delivery ratio of sparse routing is up to 83%, and the dense routing could reach up to 93%.
Fuliang Li, Yaoguang Lu, Xingwei Wang 0001, Yuanguo Bi, Tian Pan 0001, Yuchao Zhang 0004, Weichao Li 0001, Yi Wang 0004
IEEE Internet Things J.6
2019 Webpage Fingerprinting using Only Packet Length Information
abstract
Encrypted web traffic can reveal sensitive information of a user, such as their browsing histories. Existing studies on encrypted traffic analysis attacks usually focus on traffic fingerprinting of different websites rather than that of webpages from a same website. Fine-grained webpage fingerprinting allows exploiting more private information of users, e.g., their interests within a news website or an online shopping website. Since webpages from a same website usually have very similar features (e.g., statistical information) that make them indistinguishable, existing solutions may end up with low accuracy. In this paper, we propose a novel webpage fingerprinting method based on a simple and comprehensible idea. We make an observation that the length information of packets in bidirectional interaction between clients and servers can be a distinctive feature in webpage fingerprinting. Then, we extract the cumulative length of a sequence of packets to represent the fingerprint of a specific webpage. More precisely, only the first 100 packets in the loading process of a webpage is considered, thus enabling early-stage fingerprinting. The experimental results with real-world datasets demonstrate that our method is superior to other state-of-the-art approaches in terms of classification accuracy and time complexity. To the best of our knowledge, this is the first work on fine-grained webpage fingerprinting.
Meng Shen 0001, Liehuang Zhu, Yuchao Zhang 0004
ICC5
2019 Communication-Aware Container Placement and Reassignment in Large-Scale Internet Data Centers
abstract
Containerization has been used in many applications for isolation purposes due to its lightweight, scalable, and highly portable properties. However, to apply containerization in large-scale Internet data centers faces a big challenge. Services in data centers are always instantiated as a group of containers, which often generate heavy communication workloads and therefore resulting in inefficient communications and downgraded service performance. Although assigning the containers of the same service to the same server can reduce the communication overhead, this may cause heavily imbalanced resource utilization since containers of the same service are usually intensive to the same resource. To reduce communication cost as well as balance the resource utilization in large-scale data centers, we further explore the container distribution issues in a real industrial environment and find that such conflict lies in two phases-container placement and container reassignment. The objective of this paper is to address the container distribution problem in these two phases. For the container placement problem, we propose an efficient communication aware worst fit decreasing algorithm to place a set of new containers into data centers. For the container reassignment problem, we propose a two-stage algorithm called Sweep&Search to optimize a given initial distribution of containers by migrating containers among servers. We implement the proposed algorithms in Baidu's data centers and conduct extensive evaluations. Compared with the state-of-the-art strategies, the evaluation results show that our algorithms perform better up to 70% and increase the overall service throughput up to 90% simultaneously.
Yuchao Zhang 0004, Yusen Li, Ke Xu 0002, Dan Wang 0002, Wendong Wang 0003, Xuan Cao, Qingqing Liang
IEEE J. Sel. Areas Commun.2
2018 BDS: a centralized near-optimal overlay network for inter-datacenter data replication
abstract
Many important cloud services require replicating massive data from one datacenter (DC) to multiple DCs. While the performance of pair-wise inter-DC data transfers has been much improved, prior solutions are insufficient to optimize bulk-data multicast, as they fail to explore the capability of servers to store-and-forward data, as well as the rich inter-DC overlay paths that exist in geo-distributed DCs. To take advantage of these opportunities, we present BDS, an application-level multicast overlay network for large-scale inter-DC data replication. At the core of BDS is a fully centralized architecture, allowing a central controller to maintain an up-to-date global view of data delivery status of intermediate servers, in order to fully utilize the available overlay paths. To quickly react to network dynamics and workload churns, BDS speeds up the control algorithm by decoupling it into selection of overlay paths and scheduling of data transfers, each can be optimized efficiently. This enables BDS to update overlay routing decisions in near realtime (e.g., every other second) at the scale of multicasting hundreds of TB data over tens of thousands of overlay paths. A pilot deployment in one of the largest online service providers shows that BDS can achieve 3-5 x speedup over the provider's existing system and several well-known overlay routing baselines.
Yuchao Zhang 0004, Junchen Jiang, Ke Xu 0002, Xiaohui Nie, Martin J. Reed, Guang Yao, Kai Chen 0005
EuroSys1
2017 A Communication-Aware Container Re-Distribution Approach for High Performance VNFs
abstract
Containers have been used in many applications for isolation purposes due to the lightweight, scalable and highly portable properties. However, to apply containers in virtual network functions (VNFs) faces a big challenge because high-performance VNFs often generate frequent communication workloads among containers while the container communications are generally not efficient. Compared with hardware modification solutions, properly distributing containers among hosts is an efficient and low-cost way to reduce communication overhead. However, we observe that this approach yields a trade-off between the communication overhead and the overall throughput of the cluster. In this paper, we focus on the communication-aware container redistribution problem to optimize the communication overhead and the overall throughput jointly for VNF clusters. We propose a solution called FreeContainer which utilizes a novel two-stage algorithm to re-distribute containers among hosts. We implement FreeContainer in Baidu clusters with 6000 servers and 35 services deployed. Extensive experiments on real networks are conducted to evaluate the performance of the proposed approach. The results show that FreeContainer can increase the overall throughput up to 90% with significant reduction on communication overhead.
Yuchao Zhang 0004, Yusen Li, Ke Xu 0002, Dan Wang 0002, Xuan Cao, Qingqing Liang
ICDCS1
2016 Towards Minimal Tardiness of Data-Intensive Applications in Heterogeneous Networks
abstract
The increasing data requirement of Internet applications has driven a dramatic surge in developing new programming paradigms and complex scheduling algorithms to handle data-intensive workloads. Due to the expanding volume and the variety of such flows, their raw data are often processed on intermediate processing nodes before being sent to servers. The intermediate processing constraints are however not yet considered in existing task and flow computing models. In this paper, we aim to minimize the total tardiness of all flows in the presence of intermediate processing constraints. We build a model to consider Tardiness-aware Flow Scheduling with Processing constraints (TFS-P), which is unfortunately NP-Hard. Hence, we propose a heuristic Routing and Scheduling duplex MATching (RSMAT) framework based on the classic Gale-Shapley Matching Theory. We find that the problem can be well-addressed by classic Deferred Acceptance (DA) algorithm, in which the match is stable but inefficient for the model. We therefore propose the Tardiness-aware Deferred Acceptance algorithm with Dynamical Quota (TDA-DQ). This algorithm is enhanced by overcoming the inefficient stability and smartly considering the dynamical quota in the system. The evaluation compares TDA-DQ to the lower bound obtained by a modified subgradient optimization algorithm. The result indicates that TDA-DQ can achieve near-optimal performance for data-intensive applications.
Tong Li 0014, Ke Xu 0002, Meng Sheng, Kun Yang 0001, Yuchao Zhang 0004
ICCCN6
2016 PieBridge: A Cross-DR scale Large Data Transmission Scheduling System
abstract
Cross-DR WAN (Datacenter Region Wide Area Network) with various services are deployed to provide timely data information and analytics for users in a wide range of geographical locations. For its reliability and performance, data duplication synchronization is essential among different IDCs (Internet datacenters). However, this problem poses a challenge. First, data duplication requires huge amount of bandwidth whereas the bandwidth of cross-DR links and the upload/download rates of server interfaces are limited. Second, data transmissions are time sensitive, but the current network cannot complete such tasks in a timely manner. In this work, we present PieBridge, a cross-RD data duplicate transmission platform that accommodates hundreds of TBs of data generated from user applications online data analytics. We deployed PieBridge on the IDCs of Baidu and obtained promising performance results in comparison with the prevalent approaches.
Yuchao Zhang 0004, Ke Xu 0002, Guang Yao, Xiaohui Nie
SIGCOMM1
2016 Continuous double auction for cloud market: Pricing and bidding analysis
abstract
Cloud computing has recently attracted a substantial amount of attention from both industry and academia. Its growing demand gives normal users an opportunity to sell their local resources to the cloud market, which introduces new challenges for the existing coarse-grained pricing models. In this paper, we examine the potential of applying continuous double auction framework to handle these heterogeneous cloud resources. First, we establish an e-auction platform, on which cloud service providers and users can trade computing and storage resources online. Then we formulate a continuous double auction model for cloud market and further develop a novel belief-based hybrid bidding strategy (BH-strategy) for cloud players to ensure their profit maximization. At last, we conduct three simulation scenarios to compare the performance between BH-strategy and other dominating bidding strategies, and plenty of simulation results show that our BH-strategy outperforms others in all the scenarios on user surpluses by 20% or above. Besides, the BH-strategy can obtain a 16% higher efficiency in 1/3 the amount of time of other strategies.
Yuchao Zhang 0004, Ke Xu 0002, Xuelin Shi, Jiangchuan Liu
WCNC1
2015 Towards shorter task completion time in datacenter networks
abstract
Datacenters are now used as the underlying infrastructure of many modern commercial operations, powering both large Internet services and a growing number of data-intensive scientific applications. The tasks in these applications always consist of rich and complex flows which require different resources at different time slots. The existing data center scheduling frameworks are however base on either task or flow level metrics. This simplifies the design and deployment, but hardly unleashes the potentials of obtaining low task completion time for delay sensitive applications. In this paper, we show that the performance (e.g., tail and average task completion time) of existing flow-aware and task-aware network scheduling is far from being optimal. To address such a problem, we carefully examine the possibility to consider both task and flow level metrics together and present the design of TAFA (Task-Aware and Flow-Aware) in data center networks. This approach seamlessly combines the existing flow and task metrics together while successfully avoids their problems as flow-isolation and flow indiscrimination. The evaluation result shows that TAFA can obtain a near-optimal performance and reduce over 35% task completion time for the existing data center systems.
Yuchao Zhang 0004, Ke Xu 0002, Meng Shen 0001
IPCCC1
2015 Performance and incentive of teamwork-based channel allocation in spectrum access networks
abstract
Recent years have witnessed the great popularity of dynamic spectrum access networks. Such an approach is adopted between three players: government, Internet Service Providers (ISPs) and end-users. ISPs need to purchase spectrum from the government before subletting it to end-users, but currently most researches focus on the subletting process and ignore the purchasing process. In this paper, we try to investigate the game between government and ISPs in spectrum access networks. In this framework, the former aims to optimize user experience yet the later want to maximize their own profits. Such a conflict of interests introduces significant challenges to ensure end-user's performance and thus leads to a severe bottleneck to the spectrum access networks. Inspired by cooperative trends among users, we proposed a novel Channel Allocation model based on Teamwork (CAT). This approach considers both ISP's respective bands and end-user's experience and enables a smart profit sharing algorithm to address the problem. The evaluation results indicate that CAT improves the overall social welfare by about 30% than the Vickrey Clarke Groves (VCG) mechanism and obtains higher stability.
Yuchao Zhang 0004, Ke Xu 0002, Jiangchuan Liu, Yifeng Zhong
IWQoS1
2014 Online combinatorial double auction for mobile cloud computing markets
abstract
The emergence of cloud computing as an efficient means of providing computing as a form of utility can already be felt with the burgeoning of cloud service companies. Notable examples including Amazon EC2, Rackspace, Google App and Microsoft Azure have already attracted an increasing number of users over the Internet. However, due to the dynamic behaviors of some users, the traditional cloud pricing models cannot well support such popular applications as Mobile Cloud Computing (MCC). To mitigate this problem, we take our first steps towards the design of an efficient double-sided combinatorial auction model in the context of mobile cloud computing. In particular, we carefully develop the framework of online combinatorial double auctions and apply a Winner Determination Problem (WDP) model for the proposed auction mechanism. The experiment results indicate that the allocation efficiency of our proposed online auction mechanism is comparable to the social optimal solution.
Ke Xu 0002, Yuchao Zhang 0004, Xuelin Shi, Meng Shen 0001
IPCCC2