Yong Cui 0001

dblp:91/2346-1 · DBLP profile ↗
← Back
182ranked-venue papers
32as first author
64since 2021 · last 2026
0000-0002-5171-739XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 134 · 23 first-author · 47 since 2021Systems, architecture and hardware · 24 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 since 2021Security and privacy · 6 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Software engineering, systems software and programming languages · 3Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
abstract
Efficient inference of large language models (LLMs) is hindered by an ever-growing key-value (KV) cache, making KV cache compression a critical research direction. Traditional methods selectively evict less important KV cache entries, which leads to information loss and hallucinations. Recently, merging-based strategies have been explored to retain more information by merging KV pairs that would be discarded; however, these existing approaches inevitably introduce inconsistencies in attention distributions before and after merging, causing degraded generation quality. To overcome this challenge, we propose KeepKV , a novel adaptive KV cache merging method designed to preserve performance under strict memory constraints, achieving single-step lossless compression and providing error bounds for multi-step compression. KeepKV introduces the Electoral Votes mechanism that records merging history and adaptively adjusts attention scores. Moreover, it further leverages a novel Zero Inference-Perturbation Merging method, compensating for attention loss resulting from cache merging. Extensive experiments on various benchmarks and LLM architectures demonstrate that KeepKV substantially reduces memory usage while successfully retaining essential context information, achieving over 2 times inference throughput improvement and maintaining superior generation quality even with only 10% KV cache budgets.
Yuxuan Tian 0001, Yebo Peng, Aomufei Yuan, Bairen Yi, Yong Cui 0001, Tong Yang 0003
AAAI8
2026 Toward Efficient LLM Agents for Emulator-Based Network Experiment Automation
abstract
Agent-driven scientific experimentation is emerging across domains such as chemistry, biology, and materials, yet each tool class imposes its own execution discipline. Network experimentation requires more than one-shot topology or configuration synthesis: an experimenter must plan a task, operate a live and evolving network, interpret feedback, refine intermediate state, and validate the resulting behavior. This poster presents a Network Experimentation Harness for emulator-backed network experiments, helping LLM agents operate across these stateful workflows. The Harness pairs a semantic action interface with reusable experimentation skills to handle sequencing, timing, and verification that a careful experimenter would perform by hand. A preliminary study on GNS3-based network protocol experiments shows that this approach reduces wall-clock time by 47% and 37%, and token use by 81% and 76%, on average versus raw GNS3 access and a Python wrapper (GNS3Fy), respectively.
Chenguang Du, Chang Liu 0021, Lei Zhang 0157, Yong Cui 0001
APNet4
2026 AlignSketch: A Framework for Aligning Theoretical and Practical Estimation Errors
Hanyue Zheng, Jingwei Shi, Xinye Xu, Wei Zhou 0077, Tong Yang 0003, Zhenyu Guan 0002, Yong Cui 0001
ICDE8
2026 DMKG: DDoS Defense Enhanced by Multimodal Knowledge Graph Data
Jialin Niu, Mingzhe Xing, Da An, Yong Cui 0001
IWQoS5
2026 Bias in the Shadows: Explore Shortcuts in Encrypted Network Traffic Classification
abstract
Pre-trained models operating directly on raw bytes have achieved promising performance in encrypted network traffic classification (NTC), but often suffer from shortcut learning-relying on spurious correlations that fail to generalize to real-world data. Existing solutions heavily rely on model-specific interpretation techniques, which lack adaptability and generality across different model architectures and deployment scenarios. In this paper, we propose BiasSeeker, the first semi-automated framework that is both model-agnostic and data-driven for detecting dataset-specific shortcut features in encrypted traffic. By performing statistical correlation analysis directly on raw binary traffic, BiasSeeker identifies spurious or environment-entangled features that may compromise generalization, independent of any classifier. To address the diverse nature of shortcut features, we introduce a systematic categorization and apply category-specific validation strategies that reduce bias while preserving meaningful information. We evaluate BiasSeeker on 19 public datasets across three NTC tasks. By emphasizing context-aware feature selection and dataset-specific diagnosis, BiasSeeker offers a novel perspective for understanding and addressing shortcut learning in encrypted network traffic classification, raising awareness that feature selection should be an intentional and scenario-sensitive step prior to model training.
Chuyi Wang, Xiaohui Xie, Tongze Wang, Yong Cui 0001
IWQoS4
2026 NetRadar: Enabling Robust Carpet Bombing DDoS Detection
Junchen Pan, Lei Zhang 0157, Xiaoyong Si, Xinggong Zhang, Yong Cui 0001
NDSS6
2026 Adaptive DDoS attack detection via packet payload feature selection
abstract
Abstract Distributed Denial of Service attacks (DDoS) are a common and influential network malicious behavior. The timely and accurate detection of Distributed Denial-of-Service (DDoS) attacks constitutes a critically significant research imperative in cyber security. Most current research focuses on classification based on statistical characteristics of network traffic, but less considers the significance of packet payload feature for DDoS attack identification. This paper proposes an adaptive DDoS detection framework integrating machine learning with payload feature engineering. The methodology comprises three phases: 1) constructing a heterogeneous task classification system based on packet metadata analysis, 2) establishing a hierarchical keyword lexicon through payload decomposition and feature pattern mining, followed by feature vector transformation via numerical encoding, and 3) implementing supervised learning algorithms for discriminative model training and feature validity verification. This multilevel feature engineering approach demonstrates enhanced adaptability in DDoS attack pattern recognition compared to conventional detection paradigms. Test results on the public datasets CIC-DDoS-2019, ISCX-SlowDoS-2016 and DoS/DDoS-MQTT-IoT show that the average detection rate of the method in this paper reaches 98.9% for attack behaviors, and the false alarm rate is only 0.1%.
Fengjun Zhang, Yong Cui 0001, Guangcan Cui, Lisheng Huang, Yunhai Lan
Cybersecur.2
2026 State-Aware Perturbation Optimization for Robust Deep Reinforcement Learning
abstract
Recently, deep reinforcement learning (DRL) has emerged as a promising approach for robotic control. However, the deployment of DRL in real-world robots is hindered by its sensitivity to environmental perturbations. While existing whitebox adversarial attacks rely on local gradient information and apply uniform perturbations across all states to evaluate DRL robustness, they fail to account for temporal dynamics and statespecific vulnerabilities. To combat the above challenge, we first conduct a theoretical analysis of white-box attacks in DRL by establishing the adversarial victim-dynamics Markov decision process (AVD-MDP), to derive the necessary and sufficient conditions for a successful attack. Based on this, we propose a selective state-aware reinforcement adversarial attack method, named STAR, to optimize perturbation stealthiness and state visitation dispersion. STAR first employs a soft mask-based state-targeting mechanism to minimize redundant perturbations, enhancing stealthiness and attack effectiveness. Then, it incorporates an information-theoretic optimization objective to maximize mutual information between perturbations, environmental states, and victim actions, ensuring a dispersed state-visitation distribution that steers the victim agent into vulnerable states for maximum return reduction. Extensive experiments demonstrate that STAR outperforms state-of-the-art benchmarks
Zongyuan Zhang, Tianyang Duan, Zheng Lin 0001, Dong Huang 0005, Zihan Fang 0003, Zekai Sun, Ling Xiong, Hongbin Liang, Heming Cui, Yong Cui 0001
IEEE Trans. Mob. Comput.10
2026 One Sketch is Enough: Accurate Per-Flow Tail Latency Estimation With SketchPolymer
Jiarui Guo, Yuqi Dong, Yuhan Wu 0001, Yisen Hong, Xiaolin Wang 0001, Yong Cui 0001, Bin Cui 0001, Tong Yang 0003
IEEE Trans. Netw.7
2026 TitanLog: Hierarchical and Elastic Logging for High-Speed Network Data Stream
abstract
Logging network traffic plays a crucial role as it serves as the foundation for various network applications. As network scale continues to expand, contemporary network traffic becomes increasingly high-speed, high-volume, and dynamic. This growth poses challenges to traditional server-based solutions. In this paper, we proposeTitanLog, ahierarchicalandelasticlogging system designed specifically for large-scale network traffic. TitanLog utilizes thehierarchical loggingmethodology, which aims to identify the importance of each packet in real-time and log packet data of different importance at different levels. To enhance efficiency, we propose a co-design of the emerging programmable switch and the server, incorporating sketches and RDMA to boost performance. To achieve elasticity, we design mechanisms for run-time adjustments and monitoring for resource insufficiency. TitanLog possesses the capability to switch between these modes at run-time. We fully implement TitanLog on a testbed and conduct extensive evaluations. The experimental results demonstrate that TitanLog supports logging of 100Gbps traffic with a zero packet loss rate and reduces the log volume by up to 96.28%.
Yuanpeng Li 0002, Xian Niu, Yikai Zhao 0001, Tong Yang 0003, Yannan Hu, Yuchao Zhang 0004, Xiangwei Deng, Qiuheng Yin, Ruwen Zhang, Yisen Hong, Kaicheng Yang 0001, Ruijie Miao, Kun Meng, Dahui Wang, Yong Cui 0001
IEEE Trans. Netw.16
2026 NegotiaToR: Toward a Simple Yet Effective On-Demand Reconfigurable Datacenter Network
abstract
Recent advances in fast optical switching show promise in meeting the high goodput and low latency requirements of datacenter networks. We present NegotiaToR, a simple network architecture for optical reconfigurable DCNs that utilizes on-demand scheduling to handle dynamic traffic. In NegotiaToR, racks exchange scheduling messages through an in-band control plane and distributedly calculate non-conflicting paths from binary traffic demand information. Optimized for incasts, it also provides opportunities to bypass scheduling delays. NegotiaToR is compatible with prevalent flat topologies, and is tailored towards a minimalist design for on-demand reconfigurable DCNs, enhancing practicality. Through large-scale simulations, we show that NegotiaToR achieves both small mice flow completion time and high goodput on two representative flat topologies, especially under heavy loads. Particularly, the flow completion time of mice flows is one to two orders of magnitude better than the state-of-the-art traffic-oblivious reconfigurable DCN design.
Cong Liang 0005, Xiangli Song, Mowei Wang, Yashe Liu, Zhenhua Liu 0008, Shizhen Zhao, Yong Cui 0001
IEEE Trans. Netw.9
2025 FlowSentry: Accelerating NetFlow-based DDoS Detection
abstract
Distributed Denial of Service (DDoS) attacks threaten the stability of online services by overwhelming them with excessive traffic. NetFlow-based DDoS detection systems are widely adopted by Internet Service Providers (ISPs) in upstream multi-point detection scenarios to provide robust detection for volumetric DDoS attacks. However, these systems face inherent delays, as NetFlow detection is non-instantaneous—routers aggregate and summarize flow records over a period before reporting, which impacts timely detection. Existing research primarily focuses on optimizing the NetFlow reporting mechanism at the router side. Unfortunately, the need for either software or hardware upgrades for routers would incur a high deployment cost, which is impractical for ISPs in the short term. In this paper, we propose FlowSentry, a novel NetFlow detection framework to accelerate DDoS attack identification at the server side. The system operates on a dual-layer filtering paradigm to handle the high-frequency NetFlow records, incorporating two core technologies: ADWindow and STAnalyzer. ADWindow is a sketch-based sliding window mechanism designed to retain possibly anomalous flow information, filtering out benign flows to reduce the computational overhead. STAnalyzer leverages the cross-router traffic correlation to efficiently infer abnormal growth patterns of potential malicious traffic based on partially reported flow records, thus significantly reducing the detection delay. Our extensive experiments in simulated backbone network environments demonstrate that FlowSentry achieves better detection accuracy while reducing the detection delay by up to 65.63% compared to existing methods.
Xiaohui Xie, Xin Wang 0001, Lei Zhang 0157, Kun Xie 0001, Yong Cui 0001
CCS7
2025 Rethinking Adversarial Attacks in Reinforcement Learning from Policy Distribution Perspective
abstract
Deep Reinforcement Learning (DRL) suffers from uncertainties and inaccuracies in the observation signal in real-world applications. Adversarial attack is an effective method for evaluating the robustness of DRL agents. However, existing attack methods targeting individual sampled actions have limited impacts on the overall policy distribution, particularly in continuous action spaces. To address these limitations, we propose the Distribution-Aware Projected Gradient Descent attack (DAPGD). DAPGD uses distribution similarity as the gradient perturbation input to attack the policy network, which leverages the entire policy distribution rather than relying on individual samples. We utilize the Bhattacharyya distance in DAPGD to measure policy similarity, enabling sensitive detection of subtle but critical differences between probability distributions. Our experiment results demonstrate that DAPGD achieves SOTA results compared to the baselines in three robot navigation tasks, achieving an average 22.03% higher reward drop compared to the best baseline.
Tianyang Duan, Zongyuan Zhang, Zheng Lin 0001, Yue Gao 0001, Ling Xiong, Yong Cui 0001, Hongbin Liang, Xianhao Chen, Heming Cui, Dong Huang 0005
ICASSP6
2025 Modeling Flow-level Traffic Demand for Network Performance Evaluation and Optimization
abstract
Modeling traffic demand at the flow level is essential for accurate network performance evaluation and optimization. However, despite its prevalence, the common practice is oversimplified and relies on unsubstantiated assumptions of traffic homogeneity and independent arrivals. In this paper, we analyze real-world traffic data collected from production environments to challenge these assumptions. Our findings reveal notable fidelity issues in the common practice, compromising the reliability of network performance evaluation and optimization. To address these limitations, we introduce Encore, a flow-level traffic demand modeling framework that captures key traffic characteristics and generates high-fidelity synthetic traces. Encore adopts a divide-and-conquer strategy, employing tailored machine learning models for distributional and sequential modeling, along with problem-specific enhancements. Systematic evaluations demonstrate that Encore outperforms existing traffic modeling methods in terms of accuracy and coverage in distribution modeling, and fidelity in sequential modeling. In addition to accurately restoring key characteristics of real traffic, Encore improves simulation performance consistency by a factor of 4 to 17 over the common practice. Moreover, Encore achieves a ~0.88 correlation in parameter ranking compared to the ground truth, showcasing its practical utility for network optimization.
Sijiang Huang, Xiaohui Xie, Mowei Wang, Lingfeng Peng, Yong Zhang 0062, Yingjie Qin, Yong Cui 0001
ICNP9
2025 INTA: Intent-Based Translation for Network Configuration with LLM Agents
abstract
Translating configurations between different network devices is a common yet challenging task in modern network operations. This challenge arises in typical scenarios such as replacing obsolete hardware and adapting configurations to emerging paradigms like Software Defined Networking (SDN) and Network Function Virtualization (NFV). Engineers need to thoroughly understand both source and target configuration models, which requires considerable effort due to the complexity and evolving nature of these specifications. To promote automation in network configuration translation, we propose INTA, an intent-based translation framework that leverages Large Language Model (LLM) agents. The key idea of INTA is to use configuration intent as an intermediate representation for translation. It first employs LLMs to decompose configuration files and extract fine-grained intents for each configuration fragment. These intents are then used to retrieve relevant manuals of the target device. Guided by a syntax checker, INTA incrementally generates target configurations. The translated configurations are further verified and refined for semantic consistency. We implement INTA and evaluate it on real-world configuration datasets from the industry. Our approach outperforms state-of-the-art methods in translation accuracy and exhibits strong generalizability. INTA achieves an accuracy of 98.15% in terms of both syntactic and view correctness, and a command recall rate of 84.72% for the target configuration. The semantic consistency report of the translated configuration further demonstrates its practical value in real-world network operations.
Yunze Wei, Xiaohui Xie, Tianshuo Hu, Yiwei Zuo, Kaiwen Chi, Yong Cui 0001
ICNP7
2025 6GQoS: A Flow-Level QoS Assurance Framework for Next-Generation 6G Networks
abstract
The 6G network aspires to deliver everyone-centric services with stringent and dynamic QoS demands. Compared to 4G/5G, 6G requires finer-grained, per-flow resource allocation that accounts for real-time traffic variations and highly dynamic wireless channel conditions to improve user satisfaction. However, enabling per-flow QoS introduces substantial complexity in scheduling, and existing heuristic-based approaches—lacking long-term resource planning and global channel awareness—struggle to ensure fairness and service satisfaction, especially under high load and large-scale scenarios.In this paper, we present 6GQoS, a flow-level QoS assurance framework that integrates real-time QoS target setting, long-term resource management, and global channel awareness. 6GQoS continuously monitors user channel conditions, traffic demands, and service-level objectives to dynamically adjust per-flow service targets. It models long-term resource allocation via Lyapunov control theory and incorporates a novel BestUsage algorithm to guide scheduling decisions based on queue states, bandwidth gains, and resource costs. Extensive evaluations show that 6GQoS outperforms all baselines, achieving 40%–80% higher service satisfaction, reducing radio resource usage by 20%–80%, and consistently ensuring 100% fairness.
Gang Yi, Tongze Wang, Xiaohui Xie, Ling Deng, Shixian Deng, Yong Cui 0001
ICNP7
2025 Robust Deep Reinforcement Learning in Robotics via Adaptive Gradient-Masked Adversarial Attacks
abstract
Deep reinforcement learning (DRL) has emerged as a promising approach for robotic control, but its real-world deployment remains challenging due to its vulnerability to environmental perturbations. Existing white-box adversarial attack methods, adapted from supervised learning, fail to effectively target DRL agents as they overlook temporal dynamics and indiscriminately perturb all state dimensions, limiting their impact on long-term rewards. To address these challenges, we propose the Adaptive Gradient-Masked Reinforcement (AGMR) Attack, a white-box attack method that combines DRL with a gradient-based soft masking mechanism to dynamically identify critical state dimensions and optimize adversarial policies. AGMR selectively allocates perturbations to the most impactful state features and incorporates a dynamic adjustment mechanism to balance exploration and exploitation during training. Extensive experiments demonstrate that AGMR outperforms state-of-the-art adversarial attack methods in degrading the performance of the victim agent and enhances the victim agent’s robustness through adversarial defense mechanisms.
Zongyuan Zhang, Tianyang Duan, Zheng Lin 0001, Dong Huang 0005, Zihan Fang 0003, Zekai Sun, Ling Xiong, Hongbin Liang, Heming Cui, Yong Cui 0001, Yue Gao 0001
IROS10
2025 xWitch: Towards Fast and Accurate Performance Evaluation for Hierarchical QoS
abstract
Hierarchical Quality of Service (HQoS) has been designed to meet the diverse needs of different users and applications, and it is widely applied in commercial routers. However, the large number of users and applications results in numerous HQoS queue parameters that need to be configured. Fast and accurate performance evaluation of these configurations is crucial. Traditional discrete event network simulators experience significant slowdowns as the traffic increases. Existing machine learning-based methods for performance evaluation also face challenges, such as low accuracy and poor generalization, due to the long input sequences caused by large network traffic.In this paper, we propose xWitch, a packet-level performance evaluation scheme that supports HQoS. xWitch uses a sequence-to-sequence model to achieve short sequence performance prediction. A long sequence parallel prediction scheme based on dependency prediction is proposed to support fast and accurate prediction of long sequences. Experimental results show that xWitch outperforms all baselines. It achieves an average error of under 3% in latency prediction for short sequences. For long sequence prediction, it can reduce the error by more than 20% and consistently maintains latency prediction errors below 10% across different traffic distributions.
Gang Yi, Mowei Wang, Chuxuan Zeng, Feipeng Li, Yong Cui 0001
IWQoS7
2025 BCT: Modeling Block Completion Time for Transport Protocols
abstract
Many applications based on block transmission have strict latency requirements. Existing solutions to reduce latency primarily focus on enhancing packet loss resistance through redundancy and scheduling transmission orders. However, the performance of these algorithms depends on accurate block transmission time estimation. Current evaluation methods, which do not account for the impacts of complex network protocols and packet loss, result in significant estimation errors.In this paper, we propose Block Completion Time (BCT) as a transmission delay index for block-based applications. We develop a BCT distribution model that incorporates packet-level block transmission under complex protocols and apply it to predict BCT for typical TCP and QUIC protocols. As a use case, we demonstrate how our model can assist in optimizing redundancy configurations. The model is evaluated under varying network conditions, application types, and transport protocols, showing improved accuracy in predicting both the mean and distribution of BCT compared to baseline methods.
Gang Yi, Lei Zhang 0157, Yong Cui 0001
IWQoS4
2025 Fast Inference for Augmented Large Language Models
abstract
Augmented Large Language Models (LLMs) enhance standalone LLMs by integrating external data sources through API calls. In interactive applications, efficient scheduling is crucial for maintaining low request completion times, directly impacting user engagement. However, these augmentations introduce new scheduling challenges: the size of augmented requests (in tokens) no longer correlates proportionally with execution time, making traditional size-based scheduling algorithms like Shortest Job First less effective. Additionally, requests may require different handling during API calls, which must be incorporated into scheduling. This paper presents MARS, a novel inference framework that optimizes augmented LLM latency by explicitly incorporating system- and application-level considerations into scheduling. MARS introduces a predictive, memory-aware scheduling approach that integrates API handling and request prioritization to minimize completion time. We implement MARS on top of vLLM and evaluate its performance against baseline LLM inference systems, demonstrating improvements in end-to-end latency by 27%-85% and reductions in TTFT by 4%-96% compared to the existing augmented-LLM system, with even greater gains over vLLM. Our implementation is available online.
Rana Shahout, Cong Liang 0005, Shiji Xin, Qianru Lao, Yong Cui 0001, Minlan Yu, Michael Mitzenmacher
NeurIPS5
2025 Accelerating Distributed Graph Learning by Using Collaborative In-Network Multicast and Aggregation
Jiawei Huang 0001, Yijun Li 0002, Jingling Liu, Junxue Zhang 0001, Hui Li 0120, Shengwen Zhou, Xiaojuan Lu, Qichen Su, Jianxin Wang 0001, Chee-Wei Tan 0001, Yong Cui 0001, Kai Chen 0005
USENIX ATC14
2025 vClos: Network contention aware scheduling for distributed machine learning tasks in multi-tenant GPU clusters
Xinchi Han, Shizhen Zhao, Yongxi Lv, Peirui Cao, Qinwei Yang, Yunzhuo Liu, Shengkai Lin, Bo Jiang 0003, Ximeng Liu, Yong Cui 0001, Chenghu Zhou, Xinbing Wang
Comput. Networks11
2025 Advances in Attack-Defense Game Models for IIoT: A Review
abstract
The industrial internet of things (IIoT) significantly increases industrial productivity but also brings more network security threats.The evolving diversity of cyber attacks has exacerbated security risks in IIoT systems, rendering conventional passive defense mechanisms inadequate against sophisticated intrusions, such as advanced persistent threat(APT), distributed denial of services(DDoS) etc.. This necessitates the adoption of proactive defense strategies, where game theory emerges as a powerful mathematical framework for modeling dynamic attack-defense interactions. Unlike existing surveys that focus solely on game theory fundamentals or IIoT security mechanisms, this paper establishes a novel three-dimensional methodological framework for systematically reviewing game-theoretic applications in IIoT security: (1) theoretical foundations (game components and taxonomy), (2) model implementations (eight principal game types), and (3) comparative analyses (advantages/limitations per model). Crucially, we introduce an innovative classification perspective based on player relationships and payoff computation methods - a significant departure from prior categorizations. We first establish the theoretical framework by examining essential game components and taxonomy classifications. Subsequently, we investigate current research progress through the lens of these eight principal game models, conducting a comprehensive comparative analysis of their category,advantages,drawbacks. The study further identifies three fundamental limitations in existing game-theoretic approaches: imperfect information processing, dynamic adaptation constraints, and multi-agent coordination challenges. Our critical analysis proposed future research directions emphasizing hybrid defense mechanisms, machine learning-enhanced game models, and real-time response architectures. This structured review not only fills the gap in IIoT-specific game-theoretic security surveys but also provides researchers with a unified conceptual framework for developing adaptive security solutions.
Fengjun Zhang, Yong Cui 0001, Guangcan Cui
IEEE Internet Things J.2
2025 Edge-Assisted Adaptive Configuration for Serverless-Based Video Analytics
abstract
The growth of video volumes and increased DNN capabilities have led to a growing desire for video analytics, which demands intensive computation resources. Traditional resource provisioning strategies, such as configuring a cluster per peak utilization, lead to low resource efficiency. Serverless computing is a promising way to avoid wasteful resource provisioning since video analytics regularly encounters bursty input workloads and fine-grained video content dynamics. For serverless-based video analytics, the application configuration (frame rate, detection model, and computation resources) will impact several metrics, such as computation cost and analytics accuracy. In this paper, we investigate the joint configuration adjustment problem for video knobs and computation resources provided by the serverless platform. We propose an algorithm that can efficiently adapt configurations for video streams to address two key challenges in serverless-based video analytics systems, including the complex relationships between the configurations and the key performance metrics, and the dynamically best configuration. Our adaptive configuration adjustment algorithm is developed based on Markov approximation to minimize the computation cost. To guarantee the accuracy, we then design the keyframe selection algorithm based on the secretary algorithm to identify significant changes in video content. We have developed a prototype over AWS Lambda and conducted extensive experiments with real-world video streams. The results show that our algorithm can greatly reduce the computation cost under the constraint of target accuracy.
Ziyi Wang 0002, Songyu Zhang, Wei Cheng 0008, Wendong Wang 0003, Yong Cui 0001
IEEE Trans. Netw.6
2025 Automatic Dual Threshold Tuning for Switch Buffer Sharing in Datacenter Networking
abstract
For the widely deployed on-chip shared buffer, efficient buffer management is the key to absorbing bursts and avoiding packet loss during transient congestion. However, as the buffer-per-port-per-Gbps in production data centers decreases, it becomes more challenging to provide efficient buffer management to meet the requirements of heterogeneous traffic. We observe that typical shared buffer management policies have two steps: first, they identify short flows arriving at ports and then allocate more buffer room for these ports. Unfortunately, the lack of isolation between long and short flows leads to increased queue buildup and even packet loss of short flows. To address this limitation, we propose D2T, which uses different queue length thresholds for long and short flows. Specifically, we first design a compact data structure to distinguish between long and short flows. Then when two kinds of flows coexist at the same port, the threshold of long flows will decrease to absorb the bursty short flows. What’s more, we introduce D2T${}^{*}$which combines D2T with advanced DRL techniques to move toward mastering buffer management for further improving performance across various scenarios. We implement D2T at a P4-programmable switch and large-scale simulations. The results demonstrate that D2T reduces both average and tail flow completion times (FCT) of short flows by up to 29% and 62% compared with the state-of-the-art policies, respectively.
Jingling Liu, Hui Li 0120, Jiawei Huang 0001, Ping Zhong 0002, Boyan Huang, Pingping Dong, Wensheng Tang, Wanchun Jiang, Jianxin Wang 0001, Yong Cui 0001
IEEE Trans. Netw.11
2024 StarTCP: Handover-aware Transport Protocol for Starlink
abstract
Legacy transport protocols such as TCP and QUIC suffer from high packet loss and low link utilization in Starlink. From the measurement data, we figure out the ground-satellite link (GSL) handover is mainly to blame. The periodic handovers result in link interruptions and bursty losses with a fixed interval of 15s, which impair TCP’s performance. Based on this finding, we present a handover-aware transport protocol, StarTCP, which proactively stalls transmission during handovers to avoid bursty losses and erroneous congestion signals. Preliminary results indicate that StarTCP can efficiently reduce packet loss and enhance throughput in Starlink.
Li Jiang 0021, Yihang Zhang 0007, Yannan Hu, Yong Cui 0001, Xinggong Zhang
APNet4
2024 ShieldGPT: An LLM-based Framework for DDoS Mitigation
abstract
The constantly evolving Distributed Denial of Service (DDoS) attacks pose a significant threat to the cyber realm, which underscores the importance of DDoS mitigation as a pivotal area of research. While existing AI-driven approaches, including deep neural networks, show promise in detecting DDoS attacks, their inability to elucidate prediction rationales and provide actionable mitigation measures limits their practical utility. The advent of large language models (LLMs) offers a novel avenue to overcome these limitations. In this work, we introduce ShieldGPT, a comprehensive DDoS mitigation framework that harnesses the power of LLMs. ShieldGPT comprises four components: attack detection, traffic representation, domain-knowledge injection and role representation. To bridge the gap between the natural language processing capabilities of LLMs and the intricacies of network traffic, we develop a representation scheme that captures both global and local traffic features. Furthermore, we explore prompt engineering specific to the network domain and design two prompt templates that leverage LLMs to produce traffic-specific, comprehensible explanations and mitigation instructions. Our preliminary experiments and case studies validate the effectiveness and applicability of ShieldGPT, demonstrating its potential to enhance DDoS mitigation efforts with nuanced insights and tailored strategies.
Tongze Wang, Xiaohui Xie, Lei Zhang 0157, Chuyi Wang, Yong Cui 0001
APNet6
2024 D2T: Dynamic Dual Threshold Policy of Shared-Memory in Data Center Switches
abstract
Nowadays the data center switches employ the on-chip shared buffer to absorb bursts and avoid packet loss during transient congestion. However, as the buffer-per-port-per-Gbps in production data centers decreases, it becomes more challenging to provide efficient buffer management to meet the requirements of heterogeneous traffic. We observe that typical shared buffer management policies have two steps: first, they identify short flows arriving at ports and then allocate more buffer room for these ports. Unfortunately, the lack of isolation between long and short flows leads to increased queue buildup and even packet loss of short flows. To address this limitation, we propose D2T, which uses different queue length thresholds for long and short flows. Specifically, we first design a compact data structure to distinguish between long and short flows. Then when two kinds of flows coexist at the same port, the threshold of long flows will decrease to absorb the bursty short flows. We implement D2T at a P4- programmable switch and large-scale simulations. The results demonstrate that D2T reduces both average and tail flow completion times (FCT) of short flows by up to 29% and 62% compared with the state-of-the-art policies, respectively.
Jiawei Huang 0001, Hui Li 0120, Jingling Liu, Wenlu Zhang, Yijun Li 0002, Sitan Li, Shengwen Zhou, Ping Zhong 0002, Jianxin Wang 0001, Wanchun Jiang, Yong Cui 0001
ICDCS14
2024 E-DDoS: An Evaluation System for DDoS Attack Detection
abstract
Research in the area of Distributed Denial of Service (DDoS) attack detection is of paramount importance in the field of network security. Many existing studies employ static evaluation methods that fail to account for the reduction in accuracy due to the impact of inference latency on the timeliness of classification results. Furthermore, these studies frequently rely on simulated datasets for experimentation, which often lack the complexity and challenge of real-world attacks. These limitations significantly hinder the applicability of such research in practical scenarios. To overcome these challenges, we propose an evaluation methodology for real-time DDoS attack detection incorporating inference latency considerations. Additionally, we have developed a challenging DDoS dataset named THU-DDoS2024 and conducted experiments across four classification algorithms. This novel evaluation method and the newly generated dataset are integrated into an evaluation framework named E-DDoS. Leveraging E-DDoS, the “Intelligent Classification of High-Speed Network Traffic (ICNT)”, Grand Challenge was initiated. This event aims to motivate both academic and industrial sectors to delve into high-speed traffic classification tasks, thereby enhancing the applicability of research outputs to real-world applications.
Kaiwen Chi, Xiaohui Xie, Yannan Hu, Dongyang Zhao, Yuming Xie, Yong Cui 0001
ICNP7
2024 Ptu: Pre-Trained Model for Network Traffic Understanding
abstract
Network traffic understanding is crucial to providing high-quality network services and protecting network security. However, due to the growing complexity of networks and the rising proportion of encrypted traffic, existing methods for network traffic understanding face severe challenges. Traditional approaches rely on manually designed features or require a large amount of labeled data, while pre-trained models offer new possibilities. Nevertheless, existing pre-trained models have the following limitations: (1) Their inputs only contain features from the packet content, neglecting temporal information about network dynamics. (2) Their pre-training targets only focus on static characteristics of the data stream without understanding the process of the network transmission. This paper presents the Pre-trained model for network Traffic Understanding (PTU), an innovative model that employs self-supervised pre-training to address the challenges of network traffic understanding. In PTU, we design a traffic representation scheme that integrates static packet content and network dynamics into a unified input space. Furthermore, we propose a pre-training method that includes four tailored pre-training targets. This approach enables PTU to capture both static and dynamic characteristics of network traffic from massive amounts of unlabeled data, thereby achieving enhanced performance in downstream tasks through fine-tuning. Extensive experiments confirm PTU's state-of-the-art (SOTA) performance. In traffic classification tasks, PTU achieves an F1 score of over$\mathbf{0. 9 9}$and secures a more than$\mathbf{1 0 \%}$improvement in accuracy in the most challenging task of encrypted application classification.
Lingfeng Peng, Xiaohui Xie, Sijiang Huang, Ziyi Wang 0002, Yong Cui 0001
ICNP5
2024 Netmamba: Efficient Network Traffic Classification Via Pre-Training Unidirectional Mamba
abstract
Network traffic classification is a crucial research area aiming to enhance service quality, streamline network management, and bolster cybersecurity. To address the growing complexity of transmission encryption techniques, various machine learning and deep learning methods have been proposed. However, existing approaches face two main challenges. Firstly, they struggle with model inefficiency due to the quadratic complexity of the widely used Transformer architecture. Secondly, they suffer from inadequate traffic representation because of discarding important byte information while retaining unwanted biases. To address these challenges, we propose NetMamba, an efficient linear-time state space model equipped with a comprehensive traffic representation scheme. We adopt a specially selected and improved unidirectional Mamba architecture for the networking field, instead of the Transformer, to address efficiency issues. In addition, we design a traffic representation scheme to extract valid information from massive traffic data while removing biased information. Evaluation experiments on six public datasets encompassing three main classification tasks showcase NetMamba's superior classification performance compared to state-of-the-art baselines. It achieves an accuracy rate of nearly$99 \%$(some over$99 \%$) in all tasks. Additionally, NetMamba demonstrates excellent efficiency, improving inference speed by up to 60 times while maintaining comparably low memory usage. Furthermore, NetMamba exhibits superior few-shot learning abilities, achieving better classification performance with fewer labeled data. To the best of our knowledge, NetMamba is the first model to tailor the Mamba architecture for networking.
Tongze Wang, Xiaohui Xie, Wenduo Wang, Chuyi Wang, Youjian Zhao, Yong Cui 0001
ICNP6
2024 Iphicles: Tuning Parameters of Data Center Networks with Differentiable Performance Model
abstract
Tuning parameters in Data Center Networks (DCN) has long been a nuisance and one of the reasons service providers are reluctant to deploy new mechanisms in their production environments. Despite the excessive time and resources devoted to finding better configurations, a "one-size-fits-all" solution remains elusive. Neither manual configuration by experts nor black-box optimization can address the challenges of network heterogeneity and dynamics. One essential factor impeding efficient and stable parameter optimization is the need to explore in real environments, which has a long convergence time alongside the risk of performance degradation. To address this problem, we build a twin performance model of the physical DCN that approximates the mapping from parameters to Quality of Service (QoS) metrics for fast and safe performance inference and present a DCN configuration framework called Iphicles. Leveraging gradients provided by differentiable performance models built with Graph Neural Networks (GNN), Iphicles can automatically recommend better parameters efficiently and stably. Experimental results based on extensive simulation demonstrate that in complex scenarios with mixed and dynamic traffic, Iphicles can deliver parameters that lead to evident improvements in flow completion time (FCT) for both mice and elephant flows simultaneously, with minimum convergence time while maintaining performance stability during the optimization process.
Sijiang Huang, Mowei Wang, Yashe Liu, Zhenhua Liu 0008, Yong Cui 0001
IWQoS5
2024 NetSentry: Scalable Volumetric DDoS Detection with Programmable Switches
abstract
Distributed Denial of Service (DDoS) attack is a critical and persistent threat to the Internet. Recent DDoS detection schemes based on emerging programmable switches can achieve higher processing throughput and improve detection accuracy. However, with limited data plane memory, such schemes are not suitable for handling a large number of concurrent flows. Prior arts that attempt to increase memory efficiency have failed to do so without the expense of cost and accuracy. In this paper, we propose NetSentry, the first programmable switch based dynamic pooled testing DDoS detector. NetSentry detects DDoS in a pooled testing manner, where multiple flows are grouped to share the same storage unit on the data plane. NetSentry designs an elastic flow aggregation mechanism to dynamically adjust the detection granularity. Further, to achieve accurate DDoS detection for aggregated flows, NetSentry implements frequency domain DDoS detection on programmable switches. Evaluations of NetSentry’s hardware prototype show that NetSentry can achieve better accuracy while saving up to 91% of the data plane memory required to store flow features compared to the state-of-the-art programmable switch-based flow classification scheme.
Junchen Pan, Kunpeng He, Lei Zhang 0157, Zhuotao Liu, Xinggong Zhang, Yong Cui 0001
IWQoS6
2024 TransPortal: Keeping Applications Enjoying Tailored Transport Effortlessly
abstract
Since its advent, TCP has been the overwhelming choice at the transport layer. Developers utilize transport entities of TCP provided by the operating system to implement their network applications. Recently, due to the limitations of TCP and challenges in extending it, noteworthy protocols, such as TLS and QUIC, are emerging. Applications are motivated to try these modern transport protocols for security, performance, flexibility, and extensibility purposes. Unlike TCP, which adheres to the POSIX standards, most entities of these new protocols provide complex project-specific interfaces, making integrating them difficult and may introduce security risks.In this paper, we propose TransPortal, a framework enabling applications to utilize modern transport services flexibly and safely. TransPortal decouples the implementation of applications from underlying entities, and executes transport entities in separate transport agents, which can be replaced without application modification and even runtime. In addition, TransPortal monitors the execution of agents and provides error handling, protecting applications from instability or misuse of new transport entities. We implement a prototype of TransPortal and explore it with several popular transport entities. Extensive evaluation demonstrates that TransPortal provides applications convenience and safety while incurring acceptable overhead.
Xutong Zuo, Xiaohui Xie, Yong Cui 0001
IWQoS5
2024 NegotiaToR: Towards A Simple Yet Effective On-demand Reconfigurable Datacenter Network
abstract
Recent advances in fast optical switching technology show promise in meeting the high goodput and low latency requirements of datacenter networks (DCN). We present NegotiaToR, a simple network architecture for optical reconfigurable DCNs that utilizes on-demand scheduling to handle dynamic traffic. In NegotiaToR, racks exchange scheduling messages through an in-band control plane and distributedly calculate non-conflicting paths from binary traffic demand information. Optimized for incasts, it also provides opportunities to bypass scheduling delays. NegotiaToR is compatible with prevalent flat topologies, and is tailored towards a minimalist design for on-demand reconfigurable DCNs, enhancing practicality. Through large-scale simulations, we show that NegotiaToR achieves both small mice flow completion time (FCT) and high goodput on two representative flat topologies, especially under heavy loads. Particularly, the FCT of mice flows is one to two orders of magnitude better than the state-of-the-art traffic-oblivious reconfigurable DCN design.
Cong Liang 0005, Xiangli Song, Mowei Wang, Yashe Liu, Zhenhua Liu 0008, Shizhen Zhao, Yong Cui 0001
SIGCOMM8
2024 FIGRET: Fine-Grained Robustness-Enhanced Traffic Engineering
abstract
Traffic Engineering (TE) is critical for improving network performance and reliability. A key challenge in TE is the management of sudden traffic bursts. Existing TE schemes either do not handle traffic bursts or uniformly guard against traffic bursts, thereby facing difficulties in achieving a balance between normal-case performance and burst-case performance. To address this issue, we introduce FIGRET, a Fine-Grained Robustness-Enhanced TE scheme. FIGRET offers a novel approach to TE by providing varying levels of robustness enhancements, customized according to the distinct traffic characteristics of various source-destination pairs. By leveraging a burst-aware loss function and deep learning techniques, FIGRET is capable of generating high-quality TE solutions efficiently. Our evaluations of real-world production networks, including Wide Area Networks and data centers, demonstrate that FIGRET significantly outperforms existing TE schemes. Compared to the TE scheme currently deployed in Google's Jupiter data center networks, FIGRET achieves a 9%-34% reduction in average Maximum Link Utilization and improves solution speed by 35×-1800×. Against DOTE, a state-of-the-art deep learning-based TE method, FIGRET substantially lowers the occurrence of significant congestion events triggered by traffic bursts by 41%-53.9% in topologies with high traffic dynamics.
Ximeng Liu, Shizhen Zhao, Yong Cui 0001, Xinbing Wang
SIGCOMM3
2024 BOOM: Bottleneck-Aware Opportunistic Multicast Strategy for Cooperative Maritime Sensing
abstract
With the advancements in sensing technologies, maritime sensing has become indispensable in various domains, including logistics, weather forecasting, and marine ranching. However, transmitting large volumes of sensing data faces many challenges in the maritime environment. First, the transmissions purely depend on satellite links often costly and suffer from long propagation latency. On the other hand, traditional unicast transmission results in data duplication, wasting valuable marine communication resources. With the increasing density of sensing devices, the communication distance between maritime sensors has become closer, enabling the deployment of maritime opportunistic networks consisting of device-to-device links. Rather than using unicast transmission over satellite links, employing multicast with opportunistic routing enables simultaneous data transmission to multiple destinations and saves communication resources. Even though the multicast method can avoid redundancy, conducting multicast without considering the maritime characteristics (i.e., the dynamics and the distribution of sensors) may lead to inefficient data delivery. Through real-world experiments, we observe that devices located on the edges of the network have a relatively low receiving rate compared with internal ones and tend to be the bottleneck of the overall multicast progress. Based on this observation, we propose BOOM, a bottleneck-aware opportunistic multicast strategy aiming at reducing multicast latency, taking into account the influence of the bottleneck node and broadcasting rate. Prominently, within maritime scenarios challenged by extreme conditions, such as storms, typhoons, and tsunamis, BOOM’s emphasis encompasses the adaptability of multicast strategies, which necessitates dynamic adjustments in response to equipment failures and shifts in network topology. Through mathematical analysis, we prove the formation of opportunistic multicast is an NP-hard problem and further design a heuristic algorithm based on the convex-hull method to reduce the computational cost in strategy generation. We compare BOOM with four other algorithms using real-world maritime vessel trajectories in various scenarios. The simulation result illustrates that the BOOM achieves a significant reduction in transmission latency, which reduces 36% when sensors are sparsely located in water areas, and the reduction could reach up to 59% when sensors are more dense. Furthermore, in extreme environmental testing conditions, BOOM continues to outperform other algorithms in terms of completion time, with performance improvements of up to 39% and 49% in sparse and dense topology environments, respectively.
Xiao Chen 0002, Chao Zhu 0002, Guanju Shi, Xiang Gao 0013, Yong Cui 0001
IEEE Internet Things J.7
2024 Coordination of networking and computing: toward new information infrastructure and new services mode
Xiaoyun Wang 0005, Tao Sun 0010, Yong Cui 0001, Rajkumar Buyya, Deke Guo, Qun Huang 0001, Hassnaa Moustafa, Chen Tian 0001, Shangguang Wang
Frontiers Inf. Technol. Electron. Eng.3
2024 xNet: Modeling Network Performance With Graph Neural Networks
abstract
Today’s network is notorious for its complexity and uncertainty. Network operators often rely on network models for efficient network planning, operation, and optimization. The network model is responsible for understanding the complex relationships between network performance metrics (e.g., delay and jitter) and network characteristics (e.g., traffic and configuration). However, we still lack a systematic approach to developing accurate and lightweight network models that are aware of the impact of network configurations (i.e., expressiveness) and provide fine-grained flow-level temporal predictions (i.e., granularity). In this paper, we propose xNet, a data-driven network modeling framework based on graph neural networks (GNN). It is worth noting that xNet is not a dedicated network model designed for a specific network scenario with constraint considerations. On the contrary, xNet provides a general approach to modeling the network characteristics of concern with relation graph representations and configurable GNN blocks. xNet learns the state transition functions between time steps and rolls them out to obtain the full fine-grained prediction trajectory. We implement and instantiate xNet with three use cases. The experimental results show that xNet can accurately predict different performance metrics (i.e. temporal and steady-state QoS) in different scenarios, with performance comparable to state-of-the-art domain-specific models. Compared with traditional packet-level simulators, xNet achieves a speed improvement of more than two orders of magnitude, demonstrating its promising application in real-time optimization of network configurations.
Sijiang Huang, Yunze Wei, Lingfeng Peng, Mowei Wang, Linbo Hui, Zongpeng Du, Zhenhua Liu 0008, Yong Cui 0001
IEEE/ACM Trans. Netw.9
2024 Infrared and visible image fusion in a rolling guided filtering framework based on deep feature extraction
abstract
Abstract To preserve rich detail information and high contrast, a novel image fusion algorithm is proposed based on rolling-guided filtering combined with deep feature extraction. Firstly, input images are filtered to acquire various scales decomposed images using rolling guided filtering. Subsequently, PCANet is introduced to extract weight maps to guide base layer fusion. For the others layer, saliency maps of input images are extracted by a saliency measure. Then, the saliency maps are optimized by guided filtering to guide the detail layer fusion. Finally, the final fusion result are reconstructed by all fusion layers. The experimental fusion results demonstrate that fusion algorithm in this study obtains following advantages of rich detail information, high contrast, and complete edge information preservation in the subjective evaluation and better results in the objective evaluation index. In particular, the proposed method is 16.9% ahead of the best comparison result in the SD objective evaluation index.
Wei Cheng 0008, Liming Cheng, Yong Cui 0001
Wirel. Networks4
2023 Datacenter Network Deserves Better Traffic Models
abstract
Traffic modeling of Datacenter Network (DCN) today is over-simplified, deviating from the ground truth. Adopted by numerous researchers, the common practice relies on the assumptions of traffic homogeneity and independence for ease of use. Based on our investigation of a real-world traffic dataset, we disprove these assumptions and point out the severe fidelity issue of the common practice that could invalidate many motivations and conclusions from influential research works. In this paper, we present Encore, a DCN traic modeling framework for ine-grained traic modeling and high-fidelity synthetic traffic generation. Leveraging machine learning techniques, Encore effectively extracts and preserves essential distribution and sequential features from raw traic. Preliminary experiments demonstrate that the traic generated by Encore not only restores the key features of real traffic but also achieves high consistency when used to evaluate network performance. We envision further expanding Encore to full-process traffic modeling and generation, and expect these critical improvements in traffic models can facilitate the DCN performance evaluation and optimization.
Sijiang Huang, Lingfeng Peng, Mowei Wang, Yashe Liu, Zhenhua Liu 0008, Xin Wang 0001, Yong Cui 0001
HotNets7
2023 Edge-Assisted Adaptive Configuration for Serverless-Based Video Analytics
abstract
The growth of video volumes and increased DNN capabilities have led to a growing desire for video analytics, which demands intensive computation resources. Traditional resource provisioning strategies, such as configuring a cluster per peak utilization, lead to low resource efficiency. Serverless computing is a promising way to avoid wasteful resource provisioning since video analytics regularly encounters bursty input workloads and finegrained video content dynamics. For serverless-based video analytics, the application configuration (frame rate, detection model, and computation resources) will impact several metrics, such as computation cost and analytics accuracy. In this paper, we investigate the joint configuration adjustment problem for video knobs and computation resources provided by the serverless platform. We propose an algorithm that can efficiently adapt configurations for video streams to address two key challenges in serverless-based video analytics systems, including the complex relationships between the configurations and the key performance metrics, and the dynamically best configuration. Our algorithm is developed based on Markov approximation to minimize the computation cost within an accuracy constraint. We have developed a prototype over AWS Lambda and conducted extensive experiments with real-world video streams. The results show that our algorithm can greatly reduce the computation cost under the constraint of target accuracy.
Ziyi Wang 0002, Songyu Zhang, Zhixiong Wu, Yong Cui 0001
ICDCS6
2023 Poster: NegotiaToR: A Simple On-Demand Reconfigurable Data Center Network
abstract
Optical switching technology has developed fast in recent years, which has the potential to provide high good put as well as low latency in reconfigurable data center networks (RDCN). However, existing optical RDCN proposals fail to balance good put, latency, and design complexity well. In this paper, we introduce NegotiaToR, a simple optical RDCN architecture. NegotiaToR utilizes on-demand scheduling to ensure high performance and reduces possible deployment complexity with an in-band control protocol to do the schedule distributedly. Our preliminary evaluation shows that NegotiaToR outperforms the state-of-the-art optical RDCN proposal under similar complexity in both goodput and flow completion time (FCT).
Cong Liang 0005, Xiangli Song, Mowei Wang, Yashe Liu, Zhenhua Liu 0008, Yong Cui 0001
ICNP8
2023 ZGaming: Zero-Latency 3D Cloud Gaming by Image Prediction
abstract
In cloud gaming, interactive latency is one of the most important factors in users' experience. Although the interactive latency can be reduced through typical network infrastructures like edge caching and congestion control, the interactive latency of current cloud-gaming platforms is still far from users' satisfaction.
Jiangkai Wu, Yu Guan 0005, Qi Mao 0002, Yong Cui 0001, Zongming Guo, Xinggong Zhang
SIGCOMM4
2023 Edge-Assisted Real-Time Video Analytics With Spatial-Temporal Redundancy Suppression
abstract
Driven by plummeting camera prices and advances of video inference algorithms, video cameras are deployed ubiquitously and organizations usually rely on live video analytics to retrieve key information, such as the locations and identities of target objects. However, analyzing real-time video poses severe challenges to today’s network and computation systems. To balance accuracy, bandwidth usage, and latency, we present EVA, an edge-assisted real-time video analytics framework, which coordinates computationally weak cameras with more powerful edge servers to enable video analytics under the accuracy and latency requirements of applications. EVA treats the region where a target object is located as a fine-grained transmission unit and exploits the redundancies in both spatial and temporal domains to reduce the bandwidth usage. Based on the framework, we design an adaptive offloading algorithm, which coordinates the recognition process between the camera and the server. To adapt to complex environments, we then design a threshold adjustment algorithm to tune the confidence threshold dynamically. Experiments on real-world video feeds show that compared to several recent baselines on multiple video genres, EVA maintains high accuracy while reducing bandwidth usage by up to 90%.
Ziyi Wang 0002, Zhizhen Zhang, Yishuo Zhang, Wei Cheng 0008, Wendong Wang 0003, Yong Cui 0001
IEEE Internet Things J.8
2022 Locality Matters! Traffic Demand Modeling in Datacenter Networks
abstract
Understanding and modeling traffic demand characteristics in datacenter networks is of great importance for datacenter network optimization. However, prior traffic models are over-simplified and insufficient in capturing the complex locality properties of traffic demand. We analyze real-world traffic traces and discover strong dependency between the spatial attributes (source, destination) and non-spatial attributes (interarrival time, flow size) of traffic demand. We propose Lomas to model the joint distribution of multi-dimensional traffic demand attributes and generate synthetic traces. Lomas is a novel extension of hierarchical Bayes model that can represent the relationships among these attributes as a dependency graph. We validate Lomas by showing its ability to recreate the flow-level traffic demand patterns of real-world traffic traces. Our approach can be easily adapted to different datacenters with heterogeneous traffic demand patterns, making it a convenient tool for practitioners to utilize.
Mowei Wang, Yong Cui 0001
APNet3
2022 AI-enabled Multi-modal Network Anomaly Association: A Deep Self/Semi-Supervised Learning Approach
abstract
In nowadays large-scale networks, it is challenging for network operation and maintenance systems to analyze the reported massive network anomaly information. To handle this problem, we proposed a deep multi-modal learning approach called multi-modal anomaly root cause analysis, which enables network operation and maintenance systems to automatically and effectively associate the related network anomalies that appear from different modalities or aspects, and then locate the root causes. As a self/semi-supervised approach, our proposal is capable of realizing self-learning, self-adapting, and does not rely on a large number of manual annotations. According to the experimental results in a real large-scale network, without any annotations, our approach achieves up to 14% accuracy improvement in terms of multi-modal network anomaly association and root cause locating compared to the classical association rule mining algorithm Apriori, while its performance turns even much better when a few of labeled training samples are provided. The experiment also well proves the versatility and self-adaptability of our approach, which means our learning-based approach is able to not only achieve fast convergence but also automatically adapt itself to network changes.
Yinan Tang, Yabo Zhang, Zhifeng Yin, Jianxi Deng, Yong Cui 0001
ICC6
2022 To Punctuality and Beyond: Meeting Application Deadlines with DTP
abstract
Many applications have deadline requirements for their data delivery, such as real-time video, multiplayer gaming, and cloud AR/VR. However, the current transport layers' APIs are too primitive to accomplish that. Therefore, today's applications are forced to build their customized and complex deadline-aware data delivery mechanisms. In this work, we design Deadline-aware Transport Protocol (DTP) to provide deliver-before-deadline service over the wild Internet. To fulfill the diverse and sometimes conflicting requirements over the fluctuating network, we design the Active-Drop-at-Sender scheduler and adaptive redundancy. We build DTP by extending QUIC, and then develop two applications that utilize DTP. Extensive evaluations demonstrate that DTP is easy to use and can bring significant performance improvement (1.2x to 5x) compared to vanilla QUIC.
Yong Cui 0001, Feng Qian 0001, Kai Zheng 0003
ICNP3
2022 xNet: Improving Expressiveness and Granularity for Network Modeling with Graph Neural Networks
abstract
Today’s network is notorious for its complexity and uncertainty. Network operators often rely on network models to achieve efficient network planning, operation, and optimization. The network model is responsible for understanding the complex relationships between the network performance metrics (e.g., latency) and the network characteristics (e.g., traffic). However, we still lack a systematic approach to developing accurate and lightweight network models that are aware of the impact of network configurations (i.e., expressiveness) and provide fine-grained flow-level temporal predictions (i.e., granularity).In this paper, we propose xNet, a data-driven network modeling framework based on graph neural networks (GNN). Unlike the previous proposals, xNet is not a dedicated network model designed for specific network scenarios with constraint considerations. On the contrary, xNet provides a general approach to modeling the network characteristics of concern with relation graph representations and configurable GNN blocks. xNet learns the state transition function between time steps and rolls it out to obtain the full fine-grained prediction trajectory. We implement and instantiate xNet with three use cases. The experiment results show that xNet can accurately predict different performance metrics while achieving over two orders of magnitude of speedup compared with the conventional packet-level simulator.
Mowei Wang, Linbo Hui, Yong Cui 0001, Ru Liang, Zhenhua Li 0001
INFOCOM3
2022 Learning Buffer Management Policies for Shared Memory Switches
abstract
Today’s network switches often use on-chip shared memory to improve buffer efficiency and absorb bursty traffic. Current buffer management practices usually rely on simple heuristics and have unrealistic assumptions about the traffic pattern, since developing a buffer management policy suited for every scenario is infeasible. We show that modern machine learning techniques can be of essential help to learn efficient policies automatically.In this paper, we propose Neural Dynamic Threshold (NDT) that uses deep reinforcement learning (RL) to learn buffer management policies without human instructions except for a high-level objective. To tackle the high complexity and scale of the buffer management problem, we develop two domain-specific techniques upon off-the-shelf deep RL solutions. First, we design a scalable RL model by leveraging the permutation symmetry of the switch ports. Second, we use a two-level control mechanism to achieve efficient training and decision-making. The buffer allocation is directly controlled by a low-level heuristic during the decision interval, while the RL agent only decides the high-level control factor according to the traffic density. Testbed and simulation experiments demonstrate that NDT generalizes well and outperforms hand-tuned heuristic policies even on workloads for which it was not explicitly trained.
Mowei Wang, Sijiang Huang, Yong Cui 0001, Wendong Wang 0003, Zhenhua Li 0001
INFOCOM3
2022 Deadline-aware Multipath Transmission for Streaming Blocks
abstract
Interactive applications have deadline requirements, e.g. video conferencing and online gaming. Compared with a single path, which may be less stable or bandwidth insufficient, using multiple network paths simultaneously (e.g., WiFi and cellular network) can leverage the ability of multiple paths to service for the deadline. However, existing multipath schedulers usually ignore the deadline and the influence from subsequent blocks to the current scheduling decision when multiple blocks exist at the sender. In this paper, we propose DAMS, a Deadline-Aware Multipath Scheduler aiming to deliver more blocks with heterogeneous attributes before their deadlines. DAMS carefully schedules the sending order of blocks and balances its allocation on multiple paths to reduce the waste of bandwidth resources with the consideration of the block’s deadline. We implement DAMS with the inspiration of MPQUIC in user space. The extensive experimental results show that DAMS brings 41%-63% performance improvement on average compared with existing multipath solutions.
Xutong Zuo, Yong Cui 0001, Xin Wang 0001
INFOCOM2
2022 Adaptive Bitrate with User-level QoE Preference for Video Streaming
abstract
Recent years have witnessed tremendous growth of video streaming applications. To describe users’ expectations of videos, QoE was proposed, which is critical for content providers. Current video delivery systems optimize QoE with ABR algorithms. However, ABR is usually designed for an abstract "average user" without considering that QoE varies with users. In this paper, to investigate the difference in user preferences, we conduct a user study with 90 subjects and find that the average user can not represent all users. This observation inspires us to propose Ruyi, a video streaming system that incorporates preference awareness into the QoE model and the ABR algorithm. Ruyi profiles QoE preference of users and introduces preference-aware weights over different quality metrics into the QoE model. Based on this QoE model, Ruyi’s ABR is designed to directly predict the influence on metrics after taking different actions. With these predicted metrics, Ruyi chooses the bitrate that maximizes user-specific QoE once the preference is given. Consequently, Ruyi is scalable to different user preferences without re-training the learned models for each user. Simulation results show that Ruyi increases QoE for all users with up to 65.22% improvement. Testbed experimental results show that Ruyi has the highest ratings from subjects.
Xutong Zuo, Mowei Wang, Yong Cui 0001
INFOCOM4
2022 ROG: A High Performance and Robust Distributed Training System for Robotic IoT
abstract
Critical robotic tasks such as rescue and disaster response are more prevalently leveraging ML (Machine Learning) models deployed on a team of wireless robots, on which data parallel (DP) training over Internet of Things of these robots (robotic IoT) can harness the distributed hardware resources to adapt their models to changing environments as soon as possible. Unfortunately, due to the need for DP synchronization across all robots, the instability in wireless networks (i.e., fluctuating bandwidth due to occlusion and varying communication distance) often leads to severe stall of robots, which affects the training accuracy within a tight time budget and wastes energy stalling. Existing methods to cope with the instability of datacenter networks are incapable of handling such straggler effect. That is because they are conducting model-granulated transmission scheduling, which is much more coarse-grained than the granularity of transient network instability in real-world robotic IoT networks, making a previously reached schedule mismatch with the varying bandwidth during transmission. We present ROG, the first ROw-Granulated distributed training system optimized for ML training over unstable wireless networks. ROG confines the granularity of transmission and synchronization to each row of a layer’s parameters and schedules the transmission of each row adaptively to the fluctuating bandwidth. In this way the ML training process can update partial and the most important gradients of a stale robot to avoid triggering stalls, while provably guaranteeing convergence. The evaluation shows that, given the same training time, ROG achieved about 4.9%~6.5% training accuracy gain compared with the baselines and saved 20.4%~50.7% of the energy to achieve the same training accuracy.
Xiuxian Guan, Zekai Sun, Shengliang Deng, Xusheng Chen, Shixiong Zhao, Zongyuan Zhang, Tianyang Duan, Chenshu Wu, Yong Cui 0001, Libo Zhang 0001, Rui Wang 0007, Heming Cui
MICRO10
2022 Bandwidth-Efficient Multi-video Prefetching for Short Video Streaming
abstract
Applications that allow sharing of user-created short videos exploded in popularity in recent years. A typical short video application allows a user to swipe away the current video being watched and start watching the next video in a video queue. Such user interface causes significant bandwidth waste if users frequently swipe a video away before finishing watching. Solutions to reduce bandwidth waste without impairing the Quality of Experience (QoE) are needed. Solving the problem requires adaptively prefetching of short video chunks, which is challenging as the download strategy needs to match unknown user viewing behavior and network conditions. In our work, we first formulate the problem of adaptive multi-video prefetching in short video streaming. Then, to facilitate the integration and comparison of researchers' algorithms towards solving the problem, we design and implement a discrete-event simulator, which we release as open source. Finally, based on the organization of the Short Video Streaming Grand Challenge at ACM Multimedia 2022, we analyze and summarize the algorithms of the contestants, with the hope of promoting the research community towards addressing this problem.
Xutong Zuo, Yishu Li, Mohan Xu, Wei Tsang Ooi, Jiangchuan Liu, Junchen Jiang, Xinggong Zhang, Kai Zheng 0003, Yong Cui 0001
ACM Multimedia9
2022 Less is More: Service Profit Maximization in Geo-Distributed Clouds
abstract
Nowadays cloud providers purchase a good deal of bandwidth from Internet service providers to satisfy the growing requests from corporate customers for the exclusive use of inter-datacenter bandwidth. For exclusive bandwidth services, neither maximizing the revenue nor minimizing the cost can bring the maximal profit to cloud providers. The diversity of bandwidth prices and the random arrival time of user requests further increase the difficulty in economically scheduling the services to meet user requests from cloud providers. In this article, we propose to help cloud providers maximize their service profits by properly selecting user requests to serve rather than satisfying them all. We formulate the problem of service profit maximization and prove its NP-hardness. To handle offline request submission, we propose a solution that maximizes the service profit by alternately maximizing the service revenue and minimizing the service cost. To maximize service profit under online request submission, we propose an online scheduling algorithm that carefully handles the risk of not being able to pay off the incremental service cost and makes scheduling decisions in real time. Our extensive evaluations demonstrate that our solutions can achieve more than 1.6x the service profits of existing solutions.
Yong Cui 0001, Xin Wang 0001, Minming Li
IEEE Trans. Cloud Comput.2
2022 Traffic-Aware Buffer Management in Shared Memory Switches
abstract
Switch buffer serves an important role in the modern internet. To achieve efficiency, today’s switches often use on-chip shared memory. Shared memory switches rely on buffer management policies to allocate buffer among ports. To avoid waste of buffer resources or excessive buffer occupation by a few ports, existing policies tend to maximize overall buffer utilization and pursue queue length fairness. However, blind pursuit of utilization and misleading fairness definition based on queue length lead to buffer occupation with no benefit to throughput but extends queuing delay and undermines burst absorption of other ports. With analysis of current dynamic threshold policies, we demonstrate that meaningless buffer occupation can potentially impair the absorption capability of shared buffer, whereas none of the existing policies have addressed this problem. We contend that a buffer management policy should proactively detect port traffic and adjust buffer allocation accordingly. In this paper, we propose Traffic-aware Dynamic Threshold (TDT) policy. On the basis of the classic dynamic threshold policy, TDT proactively raises or lowers port threshold to absorb burst traffic or evacuate meaningless buffer occupation. We present detailed designs of port control state transition and state decision module that detect real-time traffic and change port thresholds accordingly. Simulation and DPDK-based real testbed demonstrate that TDT simultaneously optimizes for throughput, loss and delay, and reduces up to 50% flow completion time.
Sijiang Huang, Mowei Wang, Yong Cui 0001
IEEE/ACM Trans. Netw.3
2022 MultiLive: Adaptive Bitrate Control for Low-Delay Multi-Party Interactive Live Streaming
abstract
In multi-party interactive live streaming, each user can act as both the sender and the receiver of a live video stream. Designing adaptive bitrate (ABR) algorithm for such applications poses three challenges: (i) due to the interaction requirement among the users, the playback buffer has to be kept small to reduce the end-to-end delay; (ii) the algorithm needs to decide what is the bitrate to receive and what is the set of bitrates tosend; (iii) the delay and quality requirements between each pair of users may differ, for instance, depending on whether the pair is interacting directly with each other. To address these challenges, we first develop a quality of experience (QoE) model for multi-party live streaming applications. Based on this model, we designMultiLive, an adaptive bitrate control algorithm for the multi-party scenario. MultiLive models the many-to-many ABR selection problem as a non-linear programming problem. Solving the non-linear programming equation yields the target bitrate for each pair of sender-receiver. To alleviate system errors during the modeling and measurement process, we update the target bitrate through the buffer feedback adjustment. To address the throughput limitation of the uplink, we cluster the ideal streams into a few groups, and aggregate these streams through scalable video coding for transmissions. We also deploy the algorithm on a commercial live streaming platform that provides such services for more than 2300 users. The experimental results show that MultiLive outperforms the fixed bitrate algorithm, with 2-$5\times $improvement in average QoE. Furthermore, the end-to-end delay is reduced to around 100 ms, much lower than the 400 ms threshold recommended for video conferencing.
Ziyi Wang 0002, Yong Cui 0001, Xin Wang 0001, Wei Tsang Ooi, Yi Li 0015
IEEE/ACM Trans. Netw.2
2022 DeepCC: Bridging the Gap Between Congestion Control and Applications via Multiobjective Optimization
abstract
The increasingly complicated and diverse applications have distinct network performance demands, e.g., some desire high throughput while others require low latency. Traditional congestion controls (CC) have no perception of these demands. Consequently, literatures have explored the objective-specific algorithms, which are based on either offline training or online learning, to adapt to certain application demands. However, once generated, such algorithms are tailored to a specific performance objective function. Newly emerged performance demands in a changeable network environment require either expensive retraining (in the case of offline training), or manually redesigning a new objective function (in the case of online learning). To address this problem, we propose a novel architecture, DeepCC. It generates a CC agent that is generically applicable to a wide range of application requirements and network conditions. The key idea of DeepCC is to leverage both offline deep reinforcement learning and online fine-tuning. In the offline phase, instead of training towards a specific objective function, DeepCC trains its deep neural network model using multi-objective optimization. With the trained model, DeepCC offers near Pareto optimal policies w.r.t different user-specified trade-offs between throughput, delay, and loss rate without any redesigning or retraining. In addition, a quick online fine-tuning phase further helps DeepCC achieve the application-specific demands under dynamic network conditions. The simulation and real-world experiments show that DeepCC outperforms state-of-the-art schemes in a wide range of settings. DeepCC gains a higher target completion ratio of application requirements up to 67.4% than that of other schemes, even in an untrained environment.
Lei Zhang 0157, Yong Cui 0001, Mowei Wang, Kewei Zhu, Yibo Zhu 0001, Yong Jiang 0001
IEEE/ACM Trans. Netw.2
2021 Deadline-Aware Transmission Control for Real-Time Video Streaming
abstract
The deadline requirements of real-time applications rapidly increase in recent years (e.g., cloud gaming, cloud VR, online conferencing). Due to diverse network conditions, meeting deadline requirements for these applications has become one of the research hotspots. However, the current schemes focus on providing high bitrate instead of meeting deadline requirements. In this paper, we propose D3T, a flexible deadline-aware transmission mechanism that aims to improve user quality of experience (QoE) for real-time video streaming. To fulfill the diverse deadline requirements over fluctuating network conditions, D3T uses a deadline-aware scheduler to select the high priority frame before the deadline. To reduce congestion and retransmission delay, we leverage a deep reinforcement learning algorithm to make decisions of sending rate and FEC (forward error correction) redundancy ratio based on observed network status and frame information. We evaluate D3T via trace-driven simulator spanning diverse network environments, video contents and QoE metrics. D3T significantly improves the frame completion rate by reducing the bandwidth waste before the deadline. In the considered scenarios, D3T outperforms previously approaches with the improvements in average QoE of 57%.
Lei Zhang 0157, Yong Cui 0001, Junchen Pan, Yong Jiang 0001
ICNP2
2021 Poster: EasyTrans: Enable Fast Iteration of Transport Protocol
abstract
The main iteration goal of transport protocols is to optimize the performance of specific modules. In this poster, we propose a framework named EasyTrans, that enables fast iteration of transport protocol modules. With EasyTrans, developers can focus on the modules they want to iterate and no longer need to deal with other unnecessary parts of the transport protocol. Through different module calling modes, EasyTrans enables high performance even if the modules use algorithms that require sophisticated computation such as machine learning. We implement EasyTrans based on QUIC. Evaluation results show that the overhead of EasyTrans is slight.
Kai Zheng 0003, Yong Cui 0001
ICNP5
2021 Traffic-aware Buffer Management in Shared Memory Switches
abstract
Switch buffer serves an important role in modern internet. To achieve efficiency, today's switches often use on-chip shared memory. Shared memory switches rely on buffer management policies to allocate buffer among ports. To avoid waste of buffer resources or a few ports occupy too much buffer, existing policies tend to maximize overall buffer utilization and pursue queue length fairness. However, blind pursuit of utilization and misleading fairness definition based on queue length leads to buffer occupation with no benefit to throughput but extends queuing delay and undermines burst absorption of other ports. We contend that a buffer management policy should proactively detect port traffic and adjust buffer allocation accordingly. In this paper, we propose Traffic-aware Dynamic Threshold (TDT) policy. On the basis of classic dynamic threshold policy, TDT proactively raise or lower port threshold to absorb burst traffic or evacuate meaningless buffer occupation. We present detailed designs of port control state transition and state decision module that detect real time traffic and change port thresholds accordingly. Simulation and DPDK-based real testbed demonstrate that TDT simultaneously optimizes for throughput, loss and delay, and reduces up to 50% flow completion time.
Sijiang Huang, Mowei Wang, Yong Cui 0001
INFOCOM3
2021 NFD: Using Behavior Models to Develop Cross-Platform Network Functions
abstract
NFV ecosystem is flourishing and more and more NF platforms appear, but this makes NF vendors difficult to deliver NFs rapidly to diverse platforms. We propose an NF development framework named NFD for cross-platform NF development. NFD's main idea is to decouple the functional logic from the platform logic -it provides a platform-independent language to program NFs' behavior models, and a compiler with interfaces to develop platform-specific plugins. By enabling a plugin on the compiler, various NF models would be compiled to executables integrated with the target platform. We prototype NFD, build 14 NFs, and support 6 platforms (standard Linux, OpenNetVM, GPU, SGX, DPDK, OpenNF). Our evaluation shows that NFD can save development workload for cross-platform NFs and output valid and performant NFs.
Hongyi Huang, Wenfei Wu, Yongchao He, Bangwen Deng, Ying Zhang 0022, Yongqiang Xiong, Guo Chen 0001, Yong Cui 0001, Peng Cheng 0005
INFOCOM8
2021 The ACM Multimedia 2021 Meet Deadline Requirements Grand Challenge
abstract
Delay-sensitive multimedia streaming applications require their data to be delivered before a deadline to be useful. The data transmitted by these applications can usually be partitioned into blocks with different priorities, assigned based on the impact of a block on the Quality of Experience (QoE) if it misses its delivery deadline. Meet their deadline requirements is challenging due to the dynamics of the network and these applications' high demand on network resources. To encourage the research community to address this challenge, we organize the "Meet Deadline Requirements" Grand Challenge at ACM Multimedia 2021. This grand challenge provides a simulation platform onto which the participants can implement their block scheduler and bandwidth estimator and then benchmark against each other using a common set of application traces and network traces.
Junjie Deng, Mowei Wang, Yong Cui 0001, Wei Tsang Ooi, Jiangchuan Liu, Xinyu Zhang 0003, Kai Zheng 0003, Yi Li 0015
ACM Multimedia4
2021 S2Net: Preserving Privacy in Smart Home Routers
abstract
At present, wireless home routers are becoming increasingly smart. While these smart routers provide rich functionalities to users, they also raise security concerns. Although the existing end-to-end encryption techniques can be applied to protect personal data, such rich functionalities become unavailable due to the encrypted payloads. On the other hand, if the smart home routers are allowed to process and store the personal data of users, once compromised, the users' sensitive data will be exposed. As a consequence, users face a difficult trade-off between the benefits of the rich functionalities and potential privacy risks. To deal with this dilemma, we propose a novel system named Secure and Smart Network (S2Net) for home routers. For S2Net, we propose a secure OS that can distinguish and manage multiple sessions belonging to different users. The secure OS and all the router applications are placed in the secure world using the ARM TrustZone technology. In S2Net, we also confine the router applications in sandboxes provided by the proposed secure OS to prevent data leakage. As a result, S2Net can provide rich functionalities for users while preserving strong privacy for home routers. In addition, we develop a crypto-worker model that provides an abstraction layer of cryptographic tasks performed by a heterogeneous multi-core system. The other important role of crypto-worker is to parallelize the computations in order to resolve the high computation cost of cryptographic functions. We report the system design of S2Net and the details of our implementation. Experimental results with benchmarks and real applications demonstrate that our implementation is capable of achieving high performance in terms of throughput while mitigating the overhead of S2Net design.
SeungSeob Lee, Kun Tan 0002, Yunxin Liu 0001, Yong Cui 0001
IEEE Trans. Dependable Secur. Comput.6
2020 Reinforcement Learning Based Congestion Control in a Real Environment
abstract
Congestion control plays an important role in the Internet to handle real-world network traffic. It has been dominated by hand-crafted heuristics for decades. Recently, reinforcement learning shows great potentials to automatically learn optimal or near-optimal control policies to enhance the performance of congestion control. However, existing solutions train agents in either simulators or emulators, which cannot fully reflect the real-world environment and degrade the performance of network communication. In order to eliminate the performance degradation caused by training in the simulated environment, we first highlight the necessity and challenges to train a learningbased agent in real-world networks. Then we propose a framework, ARC, for learning congestion control policies in a real environment based on asynchronous execution and demonstrate its effectiveness in accelerating the training. We evaluate our scheme on the real testbed and compare it with state-of-the-art congestion control schemes. Experimental results demonstrate that our schemes can achieve higher throughput and lower latency in comparison with existing schemes.
Lei Zhang 0157, Kewei Zhu, Junchen Pan, Yong Jiang 0001, Yong Cui 0001
ICCCN6
2020 MultiLive: Adaptive Bitrate Control for Low-delay Multi-party Interactive Live Streaming
abstract
In multi-party interactive live streaming, each user can act as both the sender and the receiver of a live video stream. Designing adaptive bitrate (ABR) algorithm for such applications poses three challenges: (i) due to the interaction requirement among the users, the playback buffer has to be kept small to reduce the end-to-end delay; (ii) the algorithm needs to decide what is the bitrate to receive and what is the set of bitrates to send; (iii) the delay and quality requirements between each pair of users may differ, for instance, depending on whether the pair is interacting directly with each other. To address these challenges, we first develop a quality of experience (QoE) model for multi-party live streaming applications. Based on this model, we design MultiLive, an adaptive bitrate control algorithm for the multi-party scenario. MultiLive models the many-to-many ABR selection problem as a non-linear programming problem. Solving the non-linear programming equation yields the target bitrate for each pair of sender-receiver. To alleviate system errors during the modeling and measurement process, we update the target bitrate through the buffer feedback adjustment. To address the throughput limitation of the uplink, we cluster the ideal streams into a few groups, and aggregate these streams through scalable video coding for transmissions. We conduct extensive trace-driven simulations to evaluate the algorithm. The experimental results show that MultiLive outperforms the fixed bitrate algorithm, with 2-5× improvement in average QoE. Furthermore, the end-to-end delay is reduced to around 100 ms, much lower than the 400 ms threshold recommended for video conferencing.
Ziyi Wang 0002, Yong Cui 0001, Xin Wang 0001, Wei Tsang Ooi, Yi Li 0015
INFOCOM2
2020 Delay-Sensitive Computation Partitioning for Mobile Augmented Reality Applications
abstract
Good user experiences in Mobile Augmented Reality (MAR) applications require timely processing and rendering of virtual objects on user devices. Today's wearable AR devices are limited in computation, storage, and battery lifetime. Edge computing, where edge devices are employed to offload part or all computation tasks, allows an acceleration of computation without incurring excessive network latency. In this paper, we use acyclic data flow graphs to model the computation and data flow in MAR applications and aim to minimize the makespan of processing input frames. Due to task dependencies and variable resource availability, makespan minimization is proven to be NP-hard in general. We design DPA, a polynomial-time heuristic algorithm for this problem. For special data flow graphs including chain or star, the algorithm can provide optimal solutions or solutions with a constant approximation ratio. The effectiveness of DPA has been evaluated using extensive simulations with realistic workloads and resource availability measured from a prototype implementation.
Chaokun Zhang, Rong Zheng 0001, Yong Cui 0001, Chenhe Li
IWQoS3
2020 MPTCP+: Enhancing Adaptive HTTP Video Streaming over Multipath
abstract
This paper presents a systematic study on adaptive streaming over MPTCP. We start from realworld experiments with Dynamic Adaptive Streaming over HTTP (DASH) and analysis on its performance over MPTCP. We show that DASH can greatly benefit from the improved aggregated throughput by MPTCP; yet the inter-path throughput difference and the intra-path throughput fluctuation have noticeable (negative) impact, too. Without a proper design of path selection and adaptation in MPTCP, they can easily confuse the adaptation logic of DASH, resulting in low bitrates or frequent rebuffering even if high-bandwidth paths are available. We present MPTCP+, an extended multipath TCP solution to offer high quality and smooth playback for adaptive HTTP streaming. MPTCP+ incorporates a path use decision algorithm that smartly disables/enables a path to minimize the inter-path difference, and a novel congestion control algorithm that smooths congestion window evolution with multiple paths. We have implemented MPTCP+ in the MPTCP Linux kernel, with minimum change on the server-side MPTCP module only. It is fully compatible with the existing MPTCP clients and requires no change on the upper-layer protocols, too. Our experiments suggest that MPTCP+ increases the quality of experience (QoE) of DASH by up to 50%.
Jia Zhao 0006, Jiangchuan Liu, Cong Zhang 0002, Yong Cui 0001, Yong Jiang 0001, Wei Gong 0001
IWQoS4
2020 A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU Clusters
Yibo Zhu 0001, Chang Lan, Bairen Yi, Yong Cui 0001, Chuanxiong Guo
OSDI5
2020 Furion: Engineering High-Quality Immersive Virtual Reality on Today's Mobile Devices
abstract
Despite the growing market penetration, today's high-end virtual reality (VR) systems remain tethered, which not only limits users' VR experience but also creates a safety hazard. In this paper, we perform a systematic design study of the “elephant in the room” facing the VR industry - is it feasible to enable high-quality VR apps on untethered mobile devices such as smartphones? Our quantitative, performance-driven design study makes two contributions. First, we show that the QoE achievable for high-quality VR applications on today's mobile hardware and wireless networks via local rendering or offloading is about 10X away from the acceptable QoE, yet waiting for future mobile hardware or next-generation wireless networks (e.g., 5G) is unlikely to help, because of power limitation and the higher CPU utilization needed for processing packets under higher data rate. Second, we present Furion, a VR framework that enables high-quality, immersive mobile VR on today's mobile devices and wireless networks. Furion exploits a key insight about the VR workload that foreground interactions and background environment have contrasting predictability and rendering workload, and employs a split renderer architecture running on both the phone and the server. Supplemented with video compression, use of panoramic frames, parallel decoding on multiple cores on the phone, and view-based bitrate adaptation we demonstrate Furion can support high-quality VR apps on today's smartphones over WiFi, with under 14 ms latency and 60 FPS (the phone display refresh rate).
Zeqi Lai, Y. Charlie Hu, Yong Cui 0001, Linhui Sun, Ningwei Dai, Hung-Sheng Lee
IEEE Trans. Mob. Comput.3
2020 HyCloud: Tweaking Hybrid Cloud Storage Services for Cost-Efficient Filesystem Hosting
abstract
Today's cloud storage infrastructures typically provide two distinct types of services for hosting files: object storage like Amazon S3 and filesystem storage like Amazon EFS. In practice, a cloud storage user often desires the advantages of both-efficient filesystem operations with a low unit storage price. An intuitive approach to achieving this goal is to combine the two types of services, e.g., by hosting large files in S3 and small files together with directory structures in EFS. Unfortunately, our benchmark experiments indicate that the clients' download performance for large files becomes a severe system bottleneck. In this article, we attempt to address the bottleneck with little overhead by carefully tweaking the usages of S3 and EFS. Guided by two key observations, we design and implement an open-source system called HyCloud. It automatically invokes the data APIs of S3 and EFS on behalf of users, and intelligently schedules the data transfer among S3, EFS and the clients in a distributed manner. Real-world evaluations demonstrate that the unit storage price of HyCloud is close to that of S3, and the filesystem operations are executed as quickly as in EFS in most times (sometimes even more quickly than in EFS).
Jinlong E, Yong Cui 0001, Zhenhua Li 0001, Mingkang Ruan, Ennan Zhai
IEEE/ACM Trans. Netw.2
2019 DTP: Deadline-aware Transport Protocol
abstract
More and more applications have deadline requirements for their data delivery such as 360° video, cloud VR gaming and autonomous driving. Those applications usually are band-width hungry. Fortunately, the data of those applications can be split into multiple blocks with different priorities making it possible to reduce the bandwidth consumption by prioritizing some blocks over others. However, the existing transport layer is too primitive to accomplish that. So those applications are forced to build their own customized and complex wheels. In this work, we propose Deadline-aware Transport Protocol (DTP) to provide deliver-before-deadline service. The application expresses the deadline and metadata of the data to DTP. Then DTP tries to meet the requirement by scheduling blocks. Compared to existing protocols, DTP provides meaningful service and reduces the burden of the application developer.
Feng Qian 0001, Yong Cui 0001, Yuming Hu
APNet3
2019 Software-Defined Wide Area Network (SD-WAN): Architecture, Advances and Opportunities
abstract
Emerging applications and operational scenarios raise strict requirements for long-distance data transmission, driving network operators to design wide area networks from a new perspective. Software-defined wide area network, i.e., SD-WAN, has been regarded as the promising architecture of next-generation wide area network. To demystify software-defined wide area network, we revisit the status and challenges of legacy wide area network. We briefly introduce the architecture of software-defined wide area network. In the order from bottom to top, we survey the representative advances in each layer of software-defined wide area network. As SD-WAN based multi-objective networking has been widely discussed to provide high-quality and complicated services, we explore the opportunities and challenges brought by new techniques and network protocols.
Yong Cui 0001, Baochun Li
ICCCN2
2019 SpeedyBox: Low-Latency NFV Service Chains with Cross-NF Runtime Consolidation
abstract
Software-based service chains in Network Function Virtualization (NFV) typically suffers high processing latency. This latency grows as chain lengths increase and possibly violates application requirements. Previous efforts focus on reducing latency while maintaining the perspective of each NF being an independent, isolated module. This results in processing redundancy that could eventually become the performance bottleneck. In this paper, we propose a low-latency NFV framework called SpeedyBox, that innovatively enables cross-NF runtime optimizations in a service chain to eliminate processing redundancy. SpeedyBox builds a fast data path for flows at runtime by consolidating the aggregate actions across diverse network functions (NFs) in a service chain. In SpeedyBox, each NF is instrumented with a stateful Local Match-Action Table (MAT), and leverages our easy-to-use APIs to record its per-flow behavior in the Local MAT. Next, SpeedyBox uses a Global MAT to build the fast data path by consolidating actions from each Local MAT, while providing the ability to express the stateful NF behaviors with an Event Table. We have implemented a prototype of SpeedyBox on the BESS and OpenNetVM NFV platforms. Our trace-driven evaluation on common NFs shows that SpeedyBox achieves significant latency reduction under real world scenarios.
Yong Cui 0001, Wenfei Wu, Jiahan Gu, K. K. Ramakrishnan, Yongchao He, Xuehai Qian
ICDCS2
2019 Towards Maximal Service Profit in Geo-Distributed Clouds
abstract
With the proliferation of globally-distributed services and the quick growth of user requests for inter-datacenter bandwidth, cloud providers have to lease a good deal of bandwidth from Internet service providers to satisfy the user demands. Neither maximizing the service revenue nor minimizing the service cost can bring the maximal service profit to cloud providers. The diversity of user requests and the large unit of inter-datacenter bandwidth further increase the difficulty of scheduling user requests. In this paper, we propose a cloud operational model to help cloud providers to make more service profit by properly selecting requests to serve rather than serving all user requests. We formulate the problem of service profit maximization and prove its NP-hardness. Considering the complicated coupling between maximizing revenue and minimizing cost, we propose a framework, Metis, for the efficient scheduling of user requests over inter-datacenter networks to maximize the service profit for cloud providers. Metis is formed with the alternate operations of two algorithms derived from randomized rounding techniques and Chernoff-Hoeffding bound. We prove that they can provide the guarantees on approximation ratios. Our extensive evaluations demonstrate that Metis can achieve more than 1.3x the service profits of existing solutions.
Yong Cui 0001, Xin Wang 0001, Minming Li
ICDCS2
2019 HyCloud: Tweaking Hybrid Cloud Storage Services for Cost-Efficient Filesystem Hosting
abstract
Today's cloud storage infrastructures typically provide two distinct types of services for hosting files: object storage like Amazon S3 and filesystem storage like Amazon EFS. The former supports simple, flat object operations with a low unit storage price, while the latter supports complex, hierarchical filesystem operations with a high unit storage price. In practice, however, a cloud storage user often desires the advantages of both-efficient filesystem operations with a low unit storage price. An intuitive approach to achieving this goal is to combine the two types of services, e.g., by hosting large files in S3 and small files together with directory structures in EFS. Unfortunately, our benchmark experiments indicate that the clients' download performance for large files becomes a severe system bottleneck. In this paper, we attempt to address the bottleneck with little overhead by carefully tweaking the usages of S3 and EFS. This attempt is enabled by two key observations. First, since S3 and EFS have the same unit network-traffic price and the data transfer between S3 and EFS is free of charge, we can employ EFS as a relay for the clients' quickly downloading large files. Second, noticing that significant similarity exists between the files hosted at the cloud and its users, in most times we can convert large-size file downloads into small-size file synchronizations (through delta encoding and data compression). Guided by the observations, we design and implement an open-source system called HyCloud. It automatically invokes the data APIs of S3 and EFS on behalf of users, and handles the data transfer among S3, EFS and the clients. Real-world evaluations demonstrate that the unit storage price of HyCloud is close to that of S3, and the filesystem operations are executed as quickly as in EFS in most times (sometimes even more quickly than in EFS).
Jinlong E, Yong Cui 0001, Mingkang Ruan, Zhenhua Li 0001, Ennan Zhai
INFOCOM2
2019 Trigger relationship aware mobile traffic classification
abstract
Network traffic classification is important to network operators to ensure visibility of traffic. Network management, monitoring, and other services are built upon such classification results for improving quality of service. Compared with traffic classification in non-mobile setting, classification in mobile settings focuses on applications and has become increasingly important. Traditionally, a rule-based method is deployed in a deep packet inspector (DPI) engine for traffic classification. However, with the explosive growth in application usage, the complicated relationships including the use of content delivery networks (CDN) and sharing behaviors among applications make such methods less effective. The traffic may be identified wrongly when one application is connected to another application's server.
Heyi Tang, Yong Cui 0001, Xiaowei Yang 0001
IWQoS2
2019 Characterizing and Detecting Malicious Accounts in Privacy-Centric Mobile Social Networks: A Case Study
abstract
Malicious accounts are one of the biggest threats to the security and privacy of online social networks (OSNs). In this work, we study a new type of OSN, called privacy-centric mobile social network (PC-MSN), such as KakaoTalk and LINE, which has attracted billions of users recently. The design of PC-MSN is inspired to protect their users' privacy from strangers: (1) a stranger is not easy to send a friend request to a user who does not want to make friends with strangers; and (2) strangers cannot view a user's post. Such a design mitigates the security issue of malicious accounts. At the same time, it also brings the battleground between attackers and defenders to an earlier stage, i.e., making friendship, than the one studied in previous works. Also, previous defense proposals mostly rely on certain assumptions on the attacker, which may not be robust in the new PC-MSNs. As a result, previous malicious accounts detection approaches are less effective on a PC-MSN.
Zenghua Xia, Chang Liu 0021, Neil Zhenqiang Gong, Qi Li 0002, Yong Cui 0001, Dawn Song
KDD5
2019 The ACM Multimedia 2019 Live Video Streaming Grand Challenge
abstract
Live video streaming delivery over Dynamic Adaptive Video Streaming (DASH) is challenging as it requires low end-to-end latency, is more prone to stall, and the receiver has to decide online which representation at which bitrate to download and whether to adjust the playback speed to control the latency. To encourage the research community to come together to address this challenge, we organize the Live Video Streaming Grand Challenge at ACM Multimedia 2019. This grand challenge provides a simulation platform onto which the participants can implement their adaptive bitrate (ABR) logic and latency control algorithm, and then benchmark against each other using a common set of video traces and network traces. The ABR algorithms are evaluated using a common Quality-of- Experience (QoE) model that accounts for playback bitrate, latency constraint, frame-skipping penalty, and rebuffering penalty.
Gang Yi, Abdelhak Bentaleb, Yi Li 0015, Kai Zheng 0003, Jiangchuan Liu, Wei Tsang Ooi, Yong Cui 0001
ACM Multimedia9
2019 TailCutter: Wisely Cutting Tail Latency in Cloud CDNs Under Cost Constraints
abstract
Cloud computing platforms enable applications to offer low-latency services to users by deploying data storage in multiple geo-distributed data centers. In this paper, through benchmark measurements on Amazon AWS and Microsoft Azure together with an analysis of a large-scale dataset collected from a major cloud CDN provider, we identify the high tail latency problem in cloud CDNs, which can substantially undermine the efficacy of cloud CDNs. One crucial idea to reduce the tail latency is to send requests in parallel to multiple clouds in cloud CDNs. However, since application providers often have a budget for using cloud services, deciding how many chunks to download from each cloud and when to download chunks in a cost-efficient manner still remain as open problems in our concerned scenario. To address the problem, we present TailCutter, a workload scheduling framework that aims at optimizing the tail latency while meeting cost constraints given by application providers. Specifically, we formulate the tail latency minimization (TLM) problem in cloud CDNs and design the receding horizon control based maximum tail minimization algorithm (RHC-based MTMA) to efficiently solve the TLM problem in practice. We implement TailCutter across multiple data centers of Amazon AWS and Microsoft Azure. Extensive evaluations using a large-scale real-world data trace (collected from a major ISP) illustrate that TailCutter can reduce up to 58.9% of the 100th-percentile user-perceived latency, as compared with alternative solutions under the cost constraint.
Yong Cui 0001, Ningwei Dai, Zeqi Lai, Minming Li, Zhenhua Li 0001, Yuming Hu, Kui Ren 0001, Yuchi Chen
IEEE/ACM Trans. Netw.1
2019 Wireless Network Instabilities in the Wild: Measurement, Applications (Non)Resilience, and OS Remedy
abstract
While the bandwidth and latency improvement of both WiFi and cellular data networks in the past decades are plenty evident, the extent of signal strength fluctuation and network disruptions (unexpected switching or disconnections) experienced by mobile users in today's network deployment remains less clear. This paper makes three contributions. First, we conduct the first extensive measurement of network disruptions and significant signal strength fluctuations (together denoted as network instabilities) experienced by 2000 smartphones in the wild. Our results show that network disruptions and signal strength fluctuations remains prevalent as we moved into the 4G era. Second, we study how well popular mobile apps today handle such network instabilities. Our results show that even some of the most popular mobile apps do not implement any disruption-tolerant mechanisms. Third, we present Janus, an intelligent interface management framework that exploits the multiple interfaces on a handset to transparently handle network disruptions and satisfy apps' performance requirement. We have implemented a prototype of Janus and our evaluation using a set of popular apps shows that Janus can: 1) transparently and efficiently handle network disruptions; 2) reduce video stalls by 2.9 times and increase 31% of the time of good voice quality; 3) reduce traffic size by 26.4% and energy consumption by 16.3% compared to naive solutions.
Yong Cui 0001, Zeqi Lai, Y. Charlie Hu, Kun Tan 0002, Minglong Dai, Kai Zheng 0003, Yi Li 0015
IEEE/ACM Trans. Netw.1
2019 Cost-Efficient Scheduling of Bulk Transfers in Inter-Datacenter WANs
abstract
With the quick growth of traffic between data centers, inefficient transfer scheduling in inter-datacenter networks can lead to a huge waste of bandwidth thus significant bandwidth cost. Previous work have explored different ways, such as software-defined WANs and dynamic pricing mechanisms, to overcome the inefficiency of inter-datacenter networks. However, there is a big challenge in addressing the fundamental conflicts between the deadline-aware transfer scheduling and minimizing the bandwidth cost. Unlike existing efforts that schedule inter-datacenter transfers under fixed link capacities, wherein some deadlines are violated and the service quality is degraded, we aim to finish all the transfers on time with as little bandwidth as possible to minimize the bandwidth cost. We take into account the variation of bandwidth price and the deadline requirements of services, and formulate the problem of cost-efficient scheduling of bulk transfers with deadline guarantee, which is shown to be NP-hard. Benefitting from the relax-and-round method, we propose a progressively-descending algorithm (PDA) to schedule bulk transfers and meet the above goals with a guaranteed approximation ratio. We apply our algorithm in a bulk transfer scheduler, Butler, and build a small-scale testbed to evaluate its efficiency. Both large-scale simulation and testbed experiment results validate the ability of our scheme on cutting down the bandwidth cost. Compared with existing approaches, it reduces up to 60% bandwidth cost and increases the network utilization by up to 140%.
Yong Cui 0001, Xin Wang 0001, Minming Li, Shihan Xiao, Chuming Li
IEEE/ACM Trans. Netw.2
2018 Task Scheduling with Optimized Transmission Time in Collaborative Cloud-Edge Learning
abstract
Deep learning has been applied in many recent advanced applications in the field of transportation, finance and medicine. These applications require significant computation resources and large-scale training samples. Cloud becomes a natural choice for conducting these learning tasks due to its abundant resources. However, deeper penetration of deep learning techniques in mission critical applications, like driverless car, calls for stricter time requirement to guarantee its interaction and larger amount of dataset for training to guarantee its accuracy, which cannot be easily satisfied by the cloud and makes the network transmission become the bottleneck. Edge learning emerges to be a promising direction to reduce data transmission time by processing and compressing the raw data at the edge of the network, while brings the concern of accuracy reduction at the meantime. To balance this tradeoff under cloud-edge architecture, we study a task scheduling problem for reducing weighted transmission time which takes learning accuracy into consideration. We also propose efficient scheduling algorithms which are able to achieve up to 50% reduction in makespan with extensive trace-driven simulations.
Yutao Huang, Yifei Zhu 0001, Xiaoyi Fan 0001, Xiaoqiang Ma, Fangxin Wang 0001, Jiangchuan Liu, Ziyi Wang 0002, Yong Cui 0001
ICCCN8
2018 I Can Hear More: Pushing the Limit of Ultrasound Sensing on Off-the-Shelf Mobile Devices
abstract
Recent years have seen various acoustic applications on mobile devices, e.g. range finding, gesture recognition, and device-to-device data transport, which use near-ultrasound signals at frequencies around 18-24 kHz. Due to the fixed low sound sample rate and hardware limitation, the highest detectable sound frequency on commercial-off-the-shelf (COTS) mobile devices is capped at 24 kHz, presenting a daunting barrier that prevents high-frequency ultrasounds from benefiting acoustic applications. To bridge this gap, we present iChemo, a technology that enables COTS mobile devices to sense high-frequency ultrasound signals. Specifically, we demonstrate how to detect the power spectral density (PSD) of a high-frequency ultrasound signal by customizing the coprime sampling algorithm on COTS devices. Through our prototype and evaluation on extensive mobile devices, we demonstrate that iChemo can sense the PSD of ultrasound at frequency of 60 kHz, which is over twice of the current sensible frequency threshold.
Yuchi Chen, Wei Gong 0001, Jiangchuan Liu, Yong Cui 0001
INFOCOM4
2018 Building Generic Scalable Middlebox Services Over Encrypted Protocols
abstract
The trends of the increasing middleboxes make the middle network more and more complex. Today, many middleboxes work on application layer and offer significant network services by the plain-text traffic, such as firewalling, intrusion detecting and application layer gateways. At the same time, more and more network applications are encrypting their data transmission to protect security and privacy. It is becoming a critical task and hot topic to continue providing application-layer middlebox services in the encrypted Internet, however, the state of the art is far from being able to be deployed in the real network. In this paper, we propose a practical architecture, named PlainBox, to enable session key sharing between the communication client and the middleboxes in the network path. It employs Attribute-Based Encryption (ABE) in the key sharing protocol to support multiple chaining middleboxes efficiently and securely. We develop a prototype system and apply it to popular security protocols such as TLS and SSH. We have tested our prototype system in a lab testbed as well as real-world websites. Our result shows PlainBox introduces very little overhead and the performance is practically deployable.
Cong Liu 0029, Yong Cui 0001, Kun Tan 0002, Quan Fan, Kui Ren 0001
INFOCOM2
2018 Improving Quality of Experience for Mobile Broadcasters in Personalized Live Video Streaming
abstract
Ensuring high video quality of experience (QoE) on the broadcaster side is critical for interactive live streaming. However, measurements on multiple live streaming platforms show that they all suffer from broadcaster-side video quality degradation in the presence of transient bandwidth fluctuations. This paper presents Greedy Variable Bitrate (GVBR), a suite of solutions that optimizes the QoE through an approriate keyframe interval that trades cross-frame compression for lowered inter-frame interdependency, a simple-yet-efficient frame dropping strategy to prevent excessive frame drops, and a bitrate adaptation strategy customized for broadcasters who have shallow buffer. We compare GVBR with state-of-art algorithms in different network conditions, and find that GVBR can cut video interruption incidents by 90%, while achieving comparable bitrate.
Qingmei Ren, Yong Cui 0001, Wenfei Wu, Changfeng Chen, Yuchi Chen, Jiangchuan Liu, Hongyi Huang
IWQoS2
2018 STMS: Improving MPTCP Throughput Under Heterogeneous Networks
Yong Cui 0001, Xin Wang 0001, Yuming Hu, Minglong Dai, Fanzhao Wang, Kai Zheng 0003
USENIX ATC2
2018 Is cloud storage ready? Performance comparison of representative IP-based storage systems
Zhonghong Ou, Meina Song, Zhen-Huan Hwang, Antti Ylä-Jääski, Ren Wang 0001, Yong Cui 0001, Pan Hui 0001
J. Syst. Softw.6
2018 Dependency- and similarity-aware caching for HTTP adaptive streaming
Cong Zhang 0002, Jiangchuan Liu, Fei Chen 0010, Yong Cui 0001, Edith C. H. Ngai, Yueming Hu 0001
Multim. Tools Appl.4
2018 SDN-Based Big Data Caching in ISP Networks
abstract
Cooperative cache has become a promising technique to optimize the traffic by caching big data in networks. However, controlling distributed cache nodes to update cached contents synergistically is still challenging in designing cooperative cache systems. This paper proposes an SDN-based Cooperative Cache Network (SCCN) for ISP networks, aiming to minimize the content transmission latency while reducing the inter-ISP traffic. Based on the proposed increment recording mechanism, the SCCN Controller can timely capture the change of content popularity, and place the most popular contents on the appropriate SCCN Switches. We formulate the optimal content placement as a specific multi-commodity facility location problem and prove its NP-hardness. We propose a Relaxation Algorithm (RA) based on relaxation-rounding technique to solve the problem, which can achieve an approximation ratio of 1/2 in the worst case. To solve large scale problems for big data efficiently, we further design a Heuristic Algorithm (HA), which can find a near-optimal solution with three orders of magnitude speedup compared to RA. Specifically, HA can achieve a desirable tradeoff between the transmission delay and the Internet traffic. We implement a prototype based on Open vSwitch to demonstrate the feasibility of SCCN. Extensive trace-based simulation results show the effectiveness of SCCN under various network conditions.
Yong Cui 0001, Minming Li, Qingmei Ren, Xuejun Cai
IEEE Trans. Big Data1
2018 Diamond: Nesting the Data Center Network With Wireless Rings in 3-D Space
abstract
The introduction of wireless transmissions into the data center has shown to be promising in improving cost effectiveness of data center networks (DCNs). For high transmission flexibility and performance, a fundamental challenge is to increase the wireless availability and enable fully hybrid and seamless transmissions over both wired and wireless DCN components. Rather than limiting the number of wireless radios by the size of top-of-rack switches, we propose a novel DCN architecture, Diamond, which nests the wired DCN with radios equipped on all servers. To harvest the gain allowed by the rich reconfigurable wireless resources, we propose the low-cost deployment of scalable 3-D ring reflection spaces (RRSs) which are interconnected with streamlined wired herringbone to enable large number of concurrent wireless transmissions through high-performance multi-reflection of radio signals over metal. To increase the number of concurrent wireless transmissions within each RRS, we propose a precise reflection method to reduce the wireless interference. We build a 60-GHz-based testbed to demonstrate the function and transmission ability of our proposed architecture. We further perform extensive simulations to show the significant performance gain of diamond, in supporting up to five times higher server-to-server capacity, enabling network-wide load balancing, and ensuring high fault tolerance.
Yong Cui 0001, Shihan Xiao, Xin Wang 0001, Shenghui Yan, Chao Zhu 0002, Xiang-Yang Li 0001, Ning Ge 0001
IEEE/ACM Trans. Netw.1
2018 Truthful Online Auction Toward Maximized Instance Utilization in the Cloud
Yifei Zhu 0001, Silvery D. Fu, Jiangchuan Liu, Yong Cui 0001
IEEE/ACM Trans. Netw.4
2018 CoCloud: Enabling Efficient Cross-Cloud File Collaboration Based on Inefficient Web APIs
abstract
Cloud storage services such as Dropbox have been widely used for file collaboration among multiple users. However, this desirable functionality is yet restricted to the “walled-garden” of each service. At present, the only feasible approach to cross-cloud file collaboration seems to be using web APIs, whose performance is known to be highly unstable and unpredictable. Now that using inefficient web APIs is inevitable, in this paper we attempt to achieve sound user-perceived performance for cross-cloud file collaboration. This attempt is enabled by two key observations from real-world measurements. First, for each cloud, we are always able to deploy one or several nearby (client) proxies which can efficiently access the web APIs. Second, during file collaboration, significant similarity exists among different versions of a file. This can be exploited to substantially reduce inter-proxy traffic and thus shorten the data sync time. Guided by the observations, we design and implement an open-source prototype system called CoCloud. Currently, it supports file collaboration among four popular cloud storage services in the US and China. Its performance is well acceptable to users under representative workloads, even approaching or exceeding that of intra-cloud collaboration in many cases.
Jinlong E, Yong Cui 0001, Peng Wang 0037, Zhenhua Li 0001, Chaokun Zhang
IEEE Trans. Parallel Distributed Syst.2
2018 On the Synchronization Bottleneck of OpenStack Swift-Like Cloud Storage Systems
abstract
As one type of the most popular cloud storage services, OpenStack Swift and its follow-up systems replicate each object across multiple storage nodes and leverageobject sync protocolsto achieve high reliability andeventual consistency. The performance of object sync protocols heavily relies on two key parameters:$r$(number of replicas for each object) and$n$(number of objects hosted by each storage node). In existing tutorials and demos, the configurations are usually$r=3$and$n<1,000$by default, and the sync process seems to perform well. However, we discover in data-intensive scenarios, e.g., when$r>3$and$n\gg 1,000$, the sync process is significantly delayed and produces massive network overhead, referred to as thesync bottleneck problem. By reviewing the source code of OpenStack Swift, we find that its object sync protocol utilizes a fairly simple and network-intensive approach to check the consistency among replicas of objects. Hence in a sync round, the number of exchanged hash values per node is$\Theta (n\times r)$. To tackle the problem, we propose a lightweight and practical object sync protocol,LightSync, which not only remarkably reduces the sync overhead, but also preserves high reliability and eventual consistency. LightSync derives this capability from three novel building blocks: 1)Hashing of Hashes, which aggregates all the$h$hash values of each data partition into a single but representative hash value with the Merkle tree; 2)Circular Hash Checking, which checks the consistency of different partition replicas by only sending the aggregated hash value to the clockwise neighbor; and 3)Failed Neighbor Handling, which properly detects and handles node failures with moderate overhead to effectively strengthen the robustness of LightSync. The design of LightSync offers provable guarantee on reducing the per-node network overhead from$\Theta (n\times r)$to$\Theta (\frac{n}{h})$. Furthermore, we have implemented LightSync as an open-source patch and adopted it to OpenStack Swift, thus reducing the sync delay by up to 879$\times$and the network overhead by up to 47.5$\times$.
Mingkang Ruan, Thierry Titcheu Chekam, Ennan Zhai, Zhenhua Li 0001, Yao Liu 0001, Jinlong E, Yong Cui 0001, Hong Xu 0001
IEEE Trans. Parallel Distributed Syst.7
2017 Generic application layer protocol translation for IPv4/IPv6 transition
abstract
The exhaustion of IPv4 addresses has led to the transition to IPv6 becoming a major task for the Internet. Providing IPv6 users with accessibility to IPv4 services has been one of the most challenging tasks during the IPv4/IPv6 transition. However, some protocols contain IP addresses in the application layer, resulting in incompatibilities in traversing the network-layer translators. In this paper, we propose a Generic Application Layer Translator (GALT), which is designed to bridge the gap between IPv6 clients and IPv4 cross-layer services for the IPv4/IPv6 translation scenario. To ensure correctness and efficiency, we propose a protocol description language that enables GALT to be aware of protocol semantics. We develop the GALT system and apply it to widely used protocols with cross-layer issues, including HTTP and SIP. Our evaluation shows that GALT correctly solves the cross-layer problem in IPv4/IPv6 translation, with efficient performance for practical services.
Cong Liu 0029, Yong Cui 0001, Chaokun Zhang
ICC2
2017 GreenWay: Joint VM Placement and Topology Adaption for Green Data Center Networking
abstract
Energy consumption has become a key issue for running large-scale data center networks (DCN) nowadays. Previous studies mainly focus on energy saving through reducing the number of active servers or network switches with traffic consolidation. However, since this mechanism benefits from the routing flexibility, both the gained energy savings and flow performance are limited by the conventional static network topology. Recent advances in DCN architecture propose to implement an adaptive network topology with reconfigurable optical/wireless links (i.e., topology-adaptive DCNs), which has shown a great potential to improve the transmission performance of existing static wired network. In this paper, we propose GreenWay, an energy-efficient solution to jointly optimize the virtual machine (VM) placement and flow transmissions under the new paradigm of topology adaption. Based on the VM traffic demands, it can construct a proper run-time network topology to lower both the energy cost and communication cost among VMs. We first formulate the energy optimization problem in topology-adaptive DCNs. After showing the its NP-hardness, we then develop heuristic algorithms to address the VM placement and topology adaption effectively. Extensive trace-based simulations show that GreenWay consumes much less energy than other state-of-the- art solutions while ensuring better flow performance. We finally implement an OpenvSwitch- based testbed and demonstrate the efficiency of our solution.
Shenghui Yan, Shihan Xiao, Yuchi Chen, Yong Cui 0001, Jiangchuan Liu
ICCCN4
2017 Truthful Online Auction for Cloud Instance Subletting
abstract
Despite that IaaS users are busy scaling up/out their cloud instances to meet the ever-increasing demands, the dynamics of their demands, as well as the coarse-grained billing options offered by leading cloud providers, have led to substantial instance underutilization in both temporal and spatial domains. This paper theoretically examines an instance subletting service, where underutilized instances are leased to others within user-specified periods. Serving as a secondary market that complements the existing instance market of IaaS providers,we specifically identify the theoretical challenges in instance subletting services, and design an online auction mechanism tomake allocation and pricing decisions for the instances to besublet. Our mechanism guarantees truthfulness and individualrationality with the best possible competitive ratio. Extensivetrace-driven simulations show that our proposed mechanismachieves significant performance gains in both cost and socialwelfare.
Yifei Zhu 0001, Silvery D. Fu, Jiangchuan Liu, Yong Cui 0001
ICDCS4
2017 Wireless network instabilities in the wild: Prevalence, App (non)resilience, and OS remedy
abstract
While the bandwidth and latency improvement of both WiFi and cellular data networks in the past decade are plenty evident, the extent of signal strength fluctuation and network disruptions (unexpected switching or disconnections) experienced by mobile users in today's network deployment remains less clear. This paper makes three contributions. First, we conduct the first extensive measurement of network disruptions and signal strength fluctuations (together denoted as instabilities) experienced by 2000 smartphones in the wild. Our results show that network disruptions and signal strength fluctuations remain prevalent as we moved into the 4G era. Second, we study how well popular mobile apps today handle such network instabilities. Our results show that even some of the most popular mobile apps do not implement any disruption-tolerant mechanisms. Third, we present JANUS, an intelligent interface management framework that exploits the multiple interfaces on a handset to transparently handle network disruptions and improve apps' QoE. We have implemented JANUS on Android and our evaluation using a set of popular apps shows that Janus can (1) transparently and efficiently handle network disruptions, (2) reduce video stalls by 2.9 times and increase 31% of the time of good voice quality compared to naive solutions.
Zeqi Lai, Yong Cui 0001, Y. Charlie Hu, Kun Tan 0002, Minglong Dai, Kai Zheng 0003
ICNP2
2017 CoCloud: Enabling efficient cross-cloud file collaboration based on inefficient web APIs
abstract
Cloud storage services such as Dropbox have been widely used for file collaboration among multiple users. However, this desirable functionality is yet restricted to the “walled-garden” of each service. At present, the only effective approach to cross-cloud file collaboration seems to be using web APIs, whose performance is known to be highly unstable and unpredictable. Now that using inefficient web APIs is inevitable, in this paper we attempt to achieve sound user-perceived performance for cross-cloud file collaboration. This attempt is enabled by two key observations from real-world measurements. First, for each cloud, we are always able to deploy one or several nearby (client) proxies which can efficiently access the web APIs. Second, during file collaboration, significant similarity exists among different versions of a file. This can be exploited to substantially reduce inter-proxy traffic and thus shorten the data sync time. Guided by the observations, we design and implement an open-source prototype system called CoCloud. Currently, it supports file collaboration among four popular cloud storage services in the US and China. Its performance is well acceptable to users under representative workloads, even approaching or exceeding intra-cloud performance in many cases.
Jinlong E, Yong Cui 0001, Peng Wang 0037, Zhenhua Li 0001, Chaokun Zhang
INFOCOM2
2017 Furion: Engineering High-Quality Immersive Virtual Reality on Today's Mobile Devices
abstract
In this paper, we perform a systematic design study of the "elephant in the room" facing the VR industry -- is it feasible to enable high-quality VR apps on untethered mobile devices such as smartphones? Our quantitative, performance-driven design study makes two contributions. First, we show that the QoE achievable for high-quality VR applications on today's mobile hardware and wireless networks via local rendering or offloading is about 10X away from the acceptable QoE, yet waiting for future mobile hardware or next-generation wireless networks (e.g. 5G) is unlikely to help, because of power limitation and the higher CPU utilization needed for processing packets under higher data rate. Second, we present Furion, a VR framework that enables high-quality, immersive mobile VR on today's mobile devices and wireless networks. Furion exploits a key insight about the VR workload that foreground interactions and background environment have contrasting predictability and rendering workload, and employs a split renderer architecture running on both the phone and the server. Supplemented with video compression, use of panoramic frames, and parallel decoding on multiple cores on the phone, we demonstrate Furion can support high-quality VR apps on today's smartphones over WiFi, with under 14ms latency and 60 FPS (the phone display refresh rate).
Zeqi Lai, Y. Charlie Hu, Yong Cui 0001, Linhui Sun, Ningwei Dai
MobiCom3
2017 QuickSync: Improving Synchronization Efficiency for Mobile Cloud Storage Services
abstract
Mobile cloud storage services have gained phenomenal success in recent few years. In this paper, we identify, analyze, and address the synchronization (sync) inefficiency problem of modern mobile cloud storage services. Our measurement results demonstrate that existing commercial sync services fail to make full use of available bandwidth, and generate a large amount of unnecessary sync traffic in certain circumstances even though the incremental sync is implemented. For example, a minor document editing process in Dropbox may result in sync traffic 10 times that of the modification. These issues are caused by the inherent limitations of the sync protocol and the distributed architecture. Based on our findings, we propose QuickSync, a system with three novel techniques to improve the sync efficiency for mobile cloud storage services, and build the system on two commercial sync services. Our experimental results using representative workloads show that QuickSync is able to reduce up to 73.1 percent sync time in our experiment settings.
Yong Cui 0001, Zeqi Lai, Xin Wang 0001, Ningwei Dai
IEEE Trans. Mob. Comput.1
2017 Performance-Aware Energy Optimization on Mobile Devices in Cellular Network
abstract
In cellular networks, it is important to conserve energy while at the same time satisfying different user performance requirements. In this paper, we first propose a comprehensive metric to capture the user performance cost due to task delay, deadline violation, different application profiles, and user preferences. We prove that finding the energy-optimal scheduling solution while meeting the requirements on the performance cost is NP-hard. Then, we design an adaptive online scheduling algorithm PerES to minimize the total energy cost on data transmissions subject to user performance constraints. We prove that PerES can make the energy consumption arbitrarily close to that of the optimal scheduling solution. Further, we develop offline algorithms to serve as the evaluation benchmark for PerES. The evaluation results demonstrate that PerES achieves average 2.5 times faster convergence speed compared to state-of-art static methods, and also higher performance than peers under various test conditions. Using 821 million traffic flows collected from a commercial cellular carrier, we verify our scheme could achieve on average 32-56 percent energy savings over the total transmission energy with different levels of user experience.
Yong Cui 0001, Shihan Xiao, Xin Wang 0001, Zeqi Lai, Minming Li, Hongyi Wang 0004
IEEE Trans. Mob. Comput.1
2017 Software Defined Cooperative Offloading for Mobile Cloudlets
abstract
Device to Device communication enables the deployment of mobile cloudlets in LTE-advanced networks. The distributed nature of mobile users and dynamic task arrivals makes it challenging to schedule tasks fairly among multiple devices. Leveraging the idea of software defined networking, we propose a software defined cooperative offloading model (SDCOM), where the SDCOM controller is deployed at the PDN gateway and schedules tasks in a centralized manner to save the energy of mobile devices and reduce the traffic on access links. We formulate the minimum-energy task scheduling problem as a 0-1 knapsack problem and prove its NP-hardness. To compute the optimal solution as a benchmark, we design the conditioned optimal algorithm based on the aggregated analysis of energy consumption. The greedy algorithm with a polynominal-time complexity is proposed to solve large-scale problems efficiently. To address the problem without predicting future information on task arrivals, we further design an online task scheduling algorithm (OTS). It can minimize the energy consumption arbitrarily close to the optimal solution by appropriately setting the tradeoff coefficient. Moreover, we extend OTS to design a proportional fair online task scheduling algorithm to achieve the fair energy consumption among mobile devices. Extensive trace-based simulations demonstrate the effectiveness of SDCOM for a variety of typical mobile devices and applications.
Yong Cui 0001, Kui Ren 0001, Minming Li, Zongpeng Li, Qingmei Ren
IEEE/ACM Trans. Netw.1
2017 Traffic-Aware Virtual Machine Migration in Topology-Adaptive DCN
abstract
Virtual machine (VM) migration is a key technique for network resource optimization in modern data center networks. Previous work generally focuses on how to place the VMs efficiently in a static network topology by migrating the VMs with large traffic demands to close servers. As the flow demands between VMs change, however, a great cost will be paid for the VM migration. In this paper, we propose a new paradigm for VM migration by dynamically constructing adaptive topologies based on the VM demands to lower the cost of both VM migration and communication. We formulate the traffic-aware VM migration problem in an adaptive topology and show its NP-hardness. For periodic traffic, we develop a novel progressive-decompose-rounding algorithm to schedule VM migration in polynomial time with a proved approximation ratio. For highly dynamic flows, we design an online decision-maker (ODM) algorithm with proved performance bound. Extensive trace-based simulations show that PDR and ODM can achieve about four times flow throughput among VMs with less than a quarter of the migration cost compared to other state-of-art VM migration solutions. We finally implement an OpenvSwitch-based testbed and demonstrate the efficiency of our solutions.
Yong Cui 0001, Shihan Xiao, Xin Wang 0001, Shenghui Yan
IEEE/ACM Trans. Netw.1
2017 SPABox: Safeguarding Privacy During Deep Packet Inspection at a MiddleBox
abstract
Widely used over the Internet to encrypt traffic, HTTPS provides secure and private data communication between clients and servers. However, to cope with rapidly changing and sophisticated security attacks, network operators often deploy middleboxes to perform deep packet inspection (DPI) to detect attacks and potential security breaches, using techniques ranging from simple keyword matching to more advanced machine learning and data mining analysis. But this creates a problem: how can middleboxes, which employ DPI, work over HTTPS connections with encrypted traffic while preserving privacy? In this paper, we present SPABox, a middlebox-based system that supports both keyword-based and data analysis-based DPI functions over encrypted traffic. SPABox preserves privacy by using a novel protocol with a limited connection setup overhead. We implement SPABox on a standard server and show that SPABox is practical for both long-lived and short-lived connection. Compared with the state-of-the-art Blindbox system, SPABox is more than five orders of magnitude faster and requires seven orders of magnitude less bandwidth for connection setup while SPABox can achieve a higher security level.
Chaowen Guan, Kui Ren 0001, Yong Cui 0001, Chunming Qiao
IEEE/ACM Trans. Netw.4
2016 Inter-player Delay Optimization in Multiplayer Cloud Gaming
abstract
Novel cloud computing technology makes multiplayer cloud gaming a reality, where players play games that do not run on local devices, but on servers in the cloud. Nevertheless, the necessary communication between player's local device and cloud server increases the response delay of the gaming session. Besides the degrade of responsiveness, the inter-player delay, which is the difference of response delays perceived by players who are interacting with each other, can significantly affect the fairness of the multiplayer game. In this paper, we first introduce the Inter-player Delay Optimization (IDO) problem that aims at minimizing this inter-player delay, while preserving good-enough absolute response delay experienced by players. We further propose an efficient heuristic algorithm to solve the IDO problem. The evaluation result of a comprehensive simulation using a large-scale real-world data trace shows that IDO can reduce up to about 30% of maximum inter-player delay among interacting players comparing with the other existing solution.
Yuchi Chen, Jiangchuan Liu, Yong Cui 0001
CLOUD3
2016 Enabling Ciphertext Deduplication for Secure Cloud Storage and Access Control
abstract
To secure cloud storage and enforce access control, data encryption has become essential, given the ever increasing cyber threat everywhere. Attribute-based Encryption (ABE) crypto systems are widely considered as a promising solution under such a context for its security strength, scalability and control flexibility. One major challenge, however, for applying ABE-based techniques in real world applications is its high overhead in various aspects. In this research, we are particularly concerned with the storage size expansion in existing ABE schemes. This combined with the vast-size nature of the cloud data poses an enormous challenge to the effective usage of the cloud data storage space and affects the utility of data deduplication. Normally, data deduplication is carried out based on identifying similar and even identical contents both within and between data files, however, these patterns will be destroyed after performing data encryption using any semantically secure encryption scheme including ABE. In this research, we focus on ciphertexts deduplication under ABE, which to our best knowledge is the first of such an effort. Our fundamental observation stems from the structure of ABE ciphertexts and the possible similarities among different access structures. We show how to design a secure ciphertext deduplication scheme based on a classical CP-ABE scheme by innovatively modifying the construction with a recursive algorithm, eliminating the duplicated secrets and adding additional randomness to some certain ciphertext. We then give a detailed analysis on the proposed scheme with respect to both efficiency and security. To thoroughly assess the performance of the proposed scheme, we also implement a prototype system and conduct comprehensive experiments, which shows that our ciphertext reduplication scheme could reduce up to 80% computation and storage cost in the best case.
Heyi Tang, Yong Cui 0001, Chaowen Guan, Jian Weng 0001, Kui Ren 0001
AsiaCCS2
2016 EDASH: Energy-Aware QoE Optimization for Adaptive Video Delivery over LTE Networks
abstract
Dynamic adaptive streaming over HTTP (DASH) has emerged as a popular Internet video service, constituting a growing fraction of LTE network traffic today. We identify the root causes of DASH performance problems in bit-ate stability, energy consumption of User Equipment (UE) and efficiency of bandwidth utilization, from both the users' and network operators' perspectives. Unlike the existing researches that separately studied two important performance metrics in DASH, i.e., Quality of Experience (QoE) and UEs' energy consumption. We propose an energy-aware DASH delivery framework over LTE networks (EDASH), jointly optimizing the network throughput, users' QoE and UEs' energy efficiency. We formulate the bandwidth allocation problem as a nonlinear integer program, and design the EDASH Online Allocation algorithm (EOA). EOA assigns bandwidth based on channel conditions and buffer occupancy of UEs to achieve efficient video delivery among multiple users. Furthermore, we present the detailed design and implementation of EDASH using Apache HTTP server and Android smartphones. Both simulation and experiment results reveal that our scheme can improve the network throughput while striking a better balance between users' QoE and UEs' energy consumption.
Yong Cui 0001, Zongpeng Li, Yayun Bao, Lanshan Zhang
ICCCN2
2016 Traffic-aware virtual machine migration in topology-adaptive DCN
abstract
Virtual machine (VM) migration is a key technique for network resource optimization in modern data center networks (DCNs). Previous work generally focuses on how to place the VMs efficiently in a static network topology by migrating the VMs with large traffic demands to close servers. When the VM demands change, however, a great cost will be paid on the VM migration. With the advance of software-defined network (SDN), recent studies have shown great potential to implement an adaptive network topology at a low cost. Taking advantage of the topology adaptability, in this paper, we propose a new paradigm for VM migration by dynamically constructing a topology based on the VM demands to lower the cost of both VM migration and communication. We formulate the traffic-aware VM migration problem in an adaptive topology and show its NP-hardness. Then we develop a novel progressive-decompose-rounding (PDR) algorithm to solve this problem in polynomial time with a proved approximation ratio. Extensive trace-based simulations show that PDR can achieve higher flow throughput among VMs with only a quarter of the migration cost compared to other state-of-art VM migration solutions. We finally implement an OpenvSwitch-based testbed and demonstrate the efficiency of our solution.
Shihan Xiao, Yong Cui 0001, Xin Wang 0001, Shenghui Yan
ICNP2
2016 TECH: A Thermal-Aware and Cost Efficient Mechanism for Colocation Demand Response
abstract
Data centers are promising participants in emergency demand response (EDR) programs, in which the power grids incentivize large energy consumers to reduce energy consumption in emergency to avoid potential huge financial losses. However, in multi-tenant colocation data centers, tenants manage their own servers and often sign fixed energy contracts with data center operators, thus having no incentives to contribute to EDR. To solve this problem, several studies have investigated various market-based mechanisms to incentivize tenants to reduce their server energy consumption for EDR. Nonetheless, these purely market-based studies are severely limited in one or both of the following key aspects. (1) Lack of coordination of cooling system: Due to thermal unawareness, the existing mechanisms leave the supplied cooling air temperature at an unnecessarily low level to avoid server overheating, resulting in cooling energy inefficiency (2) Violation of cost efficiency: The mechanism must be implemented in a cost efficient way such that operators do not lose financial interest, which, however, is violated by many of the existing mechanisms. This work proposes a novel thermal-aware and cost efficient mechanism, called TECH, which coordinate tenants' energy reduction in concert with the cooling system control to enable colocation EDR in a cost efficient way.
Fan Wu 0006, Shaolei Ren, Xiaofeng Gao 0001, Guihai Chen, Yong Cui 0001
ICPP6
2016 On the synchronization bottleneck of OpenStack Swift-like cloud storage systems
abstract
As one type of the most popular cloud storage services, OpenStack Swift and its follow-up systems replicate each data object across multiple storage nodes and leverage object sync protocols to achieve high availability and eventual consistency. The performance of object sync protocols heavily relies on two key parameters: r (number of replicas for each object) and η (number of objects hosted by each storage node). In existing tutorials and demos, the configurations are usually r = 3 and n3 and n ≫ 1000, the object sync process is significantly delayed and produces massive network overhead. This phenomenon is referred to as the sync bottleneck problem. Then, to explore the root cause, we review the source code of OpenStack Swift and find that its object sync protocol utilizes a fairly simple and network-intensive approach to check the consistency among replicas of objects. In particular, each storage node is required to periodically multicast the hash values of all its hosted objects to all the other replica nodes. Thus in a sync round, the number of exchanged hash values per node is Θ(n×r). Further, to tackle the problem, we propose a lightweight object sync protocol called LightSync. It remarkably reduces the sync overhead by using two novel building blocks: 1) Hashing of Hashes, which aggregates all the h hash values of each data partition into a single but representative hash value with the Merkle tree; 2) Circular Hash Checking, which checks the consistency of different partition replicas by only sending the aggregated hash value to the clockwise neighbor. Its design provably reduces the per-node network overhead from Θ(n×r) to Θ(n/h). In addition, we have implemented LightSync as an open-source patch and adopted it to OpenStack Swift, thus reducing sync delay by up to 28.8× and network overhead by up to 14.2×.
Thierry Titcheu Chekam, Ennan Zhai, Zhenhua Li 0001, Yong Cui 0001, Kui Ren 0001
INFOCOM4
2016 TailCutter: Wisely cutting tail latency in cloud CDN under cost constraints
abstract
Cloud computing platforms enable applications to offer low latency access to user data by offering storage services in several geographically distributed data centers. In this paper, we identify the high tail latency problem in cloud CDN via analyzing a large-scale dataset collected from 783,944 users in a major cloud CDN. We find that the data downloading latency in cloud CDN is highly variable, which may significantly degrade the user experience of applications. To address the problem, we present TailCutter, a workload scheduling mechanism that aims at optimizing the tail latency while meeting the cost constraint given by application providers. We further design the Maximum Tail Minimization Algorithm (MTMA) working in TailCutter mechanism to optimally solve the Tail Latency Minimization (TLM) problem in polynomial time. We implement TailCutter across data centers of Amazon S3 and Microsoft Azure. Our extensive evaluation using large-scale real world data traces shows that TailCutter can reduce up to 68% 99th percentile user-perceived latency in comparison with alternative solutions under cost constraints.
Zeqi Lai, Yong Cui 0001, Minming Li, Zhenhua Li 0001, Ningwei Dai, Yuchi Chen
INFOCOM2
2016 Indoor Tracking Using Crowdsourced Maps
abstract
Using crowdsourced visual and inertial sensor data for indoor mapping has attracted much attention in recent years. Nevertheless, the opportunities and challenges of indoor tracking using crowdsourced maps have not been fully explored. In this work, we aim at tackling the challenges due to incomplete obstacle information in crowdsourced indoor maps, especially at the initialization stage of crowdsourcing. We propose a novel solution for particle-filtering-based indoor tracking, using the crowdsourced maps derived from image-based 3D point clouds. Our solution enhances particle filtering with density-based collision detection and history-based particle regeneration. Evaluation with real user traces demonstrates that our solution outperforms the state-of-the-art. In particular, it reduces the average distance error of indoor tracking by 47% when using crowdsourced 3D point clouds.
Yu Xiao 0001, Zhonghong Ou, Yong Cui 0001, Antti Ylä-Jääski
IPSN4
2016 Multi-Resource Partial-Ordered Task Scheduling in cloud computing
abstract
In this paper, we investigate the scheduling problem with multi-resource allocation in cloud computing environments. In contrast to existing work that focuses on flow-level scheduling, which treats flows in isolation, we consider dependency among subtasks of applications that imposes a partial order relationship in execution. We formulate the problem of Multi-Resource Partial-Ordered Task Scheduling (MR-POTS) to minimize the makespan. In the first stage, the proposed Dominant Resource Priority (DRP) algorithm decides the collection of subtasks for resource allocation by taking into account the partial order relationship and characteristics of subtasks. In the second stage, the proposed Maximum Utilization Allocation (MUA) algorithm partitions multiple resources among selected subtasks with the objective to maximize the overall utilization. Both theoretical analysis and experimental evaluation demonstrate the proposed algorithms can approximately achieve the minimal makespan with high resource utilization. Specifically, a reduction of 50% in makespan can be achieved compared with existing scheduling schemes.
Chaokun Zhang, Yong Cui 0001, Rong Zheng 0001, Jinlong E
IWQoS2
2016 Diamond: Nesting the Data Center Network with Wireless Rings in 3D Space
Yong Cui 0001, Shihan Xiao, Xin Wang 0001, Chao Zhu 0002, Xiang-Yang Li 0001, Ning Ge 0001
NSDI1
2016 Throughput Optimization via Association Control in Wireless LANs
Heyi Tang, Zhonghong Ou, Yong Cui 0001
Mob. Networks Appl.5
2015 Demand-Aware Load Balancing in Wireless LANs Using Association Control
abstract
The densely deployed Access Points (APs) have overlapping coverage areas. In the hotspot area, a user can usually receive signals of more than ten APs. The signal-based association in IEEE 802.11 may result in significant unbalanced loads among APs. Moreover, diverse user demands on bandwidth further exacerbate the load unbalance. Some APs are too overloaded when multiple high-demand users gather on them due to the strongest signal, while the others are light-loaded with a few low-demand users associated to them. The severe load unbalance degrades the user performance. In this paper, by adding different user bandwidth demands as new constraints, we formulate the joint AP association and bandwidth allocation problem. We comprehensively analyze the solution space of optimal bandwidth allocation while satisfying the time-based fairness among users. We develop a 1/2-approximation algorithm to solve the problem. Our extensive trace- driven evaluations show that our algorithm achieves better load balance. As a result, it greatly improves the aggregated throughput and provides better user fairness than conventional association schemes.
Yong Cui 0001, Heyi Tang, Shihan Xiao
GLOBECOM2
2015 Joint Media Streaming Optimization of Energy and Rebuffering Time in Cellular Networks
abstract
Streaming services are gaining popularity and have contributed a tremendous fraction of today's cellular network traffic. Both playback fluency and battery endurance are significant performance metrics for mobile streaming services. However, because of the unpredictable network condition and the loose coupling between upper layer streaming protocols and underlying network configurations, jointly optimizing rebuffering time and energy consumption for mobile streaming services remains a significant challenge. In this paper, we propose a novel framework that effectively addresses the above limitations and optimizes video transmission in cellular networks. We design two complementary algorithms, Rebuffering Time Minimization Algorithm (RTMA) and Energy Minimization Algorithm (EMA) in this framework, to achieve smoothed playback and energy-efficiency on demand over multi-user scenarios. Our algorithms integrate cross-layer parameters to schedule video delivery. Specifically, RTMA aims at achieving the minimum rebuffering time with limited energy and EMA tries to obtain the minimum energy consumption while meeting the rebuffering time constraint. Extensive simulation demonstrates that RTMA is able to reduce at least 68% rebuffering time and EMA can achieve more than 27% energy reduction compared with other state-of-the-art solutions.
Zeqi Lai, Yong Cui 0001, Yayun Bao, Jiangchuan Liu, Yingchao Zhao 0001, Xiao Ma 0009
ICPP2
2015 Dynamic flow consolidation for energy savings in green DCNs
abstract
Energy consumption of data center has become an important challenge due to high electric cost and carbon dioxide emissions. Previous work has mainly focused on saving energy cost of servers, though the energy consumption of data center networks (DCNs), consisting of networking equipments like switches, also takes a significant part of the overall energy consumption. In this paper, we propose ProCons, an energy saving mechanism that dynamically consolidates traffic flows onto a small set of networking equipments in order to shut down idle ones for energy saving. Different from previous works that assume the traffic demands to be stable, ProCons takes into account the variance of traffic demand over time, and predicts future demand based on historical statistics. The traffic flows are then scheduled based on the predicted future demands and the capacity of each link. We evaluate ProCons with real life traces collected from data centers using a flow-level simulator. Our experimental results show that using ProCons, 40% of energy savings for DCNs can be gained while maintaining the good performance of flow transmission.
Chao Zhu 0002, Yu Xiao 0001, Yong Cui 0001, Shihan Xiao, Antti Ylä-Jääski
IPCCC3
2015 QuickSync: Improving Synchronization Efficiency for Mobile Cloud Storage Services
abstract
Mobile cloud storage services have gained phenomenal success in recent few years. In this paper, we identify, analyze and address the synchronization (sync) inefficiency problem of modern mobile cloud storage services. Our measurement results demonstrate that existing commercial sync services fail to make full use of available bandwidth, and generate a large amount of unnecessary sync traffic in certain circumstance even though the incremental sync is implemented. These issues are caused by the inherent limitations of the sync protocol and the distributed architecture. Based on our findings, we propose QuickSync, a system with three novel techniques to improve the sync efficiency for mobile cloud storage services, and build the system on two commercial sync services. Our experimental results using representative workloads show that QuickSync is able to reduce up to 52.9% sync time in our experiment settings.
Yong Cui 0001, Zeqi Lai, Xin Wang 0001, Ningwei Dai, Congcong Miao
MobiCom1
2015 FMTCP: A Fountain Code-Based Multipath Transmission Control Protocol
abstract
Ideally, the throughput of a Multipath TCP (MPTCP) connection should be as high as that of multiple disjoint single-path TCP flows. In reality, the throughput of MPTCP is far lower than expected. In this paper, we conduct an extensive simulation-based study on this phenomenon, and the results indicate that a subflow experiencing high delay and loss severely affects the performance of other subflows, thus becoming the bottleneck of the MPTCP connection and significantly degrading the aggregate goodput. To tackle this problem, we propose Fountain code-based Multipath TCP (FMTCP), which effectively mitigates the negative impact of the heterogeneity of different paths. FMTCP takes advantage of the random nature of the fountain code to flexibly transmit encoded symbols from the same or different data blocks over different subflows. Moreover, we design a data allocation algorithm based on the expected packet arriving time and decoding demand to coordinate the transmissions of different subflows. Quantitative analyses are provided to show the benefit of FMTCP. We also evaluate the performance of FMTCP through ns-2 simulations and demonstrate that FMTCP outperforms IETF-MPTCP, a typical MPTCP approach, when the paths have diverse loss and delay in terms of higher total goodput, lower delay, and jitter. In addition, FMTCP achieves high stability under abrupt changes of path quality.
Yong Cui 0001, Xin Wang 0001, Hongyi Wang 0004
IEEE/ACM Trans. Netw.1
2015 Cooperative Coverage Extension for Relay-Union Networks
abstract
Multi-hop coverage extension can be utilized as a feasible approach to facilitating uncovered users to get Internet service in public area WLANs. In this paper we introduce a relay-union network (RUN), which refers to a public area WLAN in which users often wander in the same area and have the ability to provide data forwarding services for others. We develop a RUN framework to model the cost of providing forwarding services and the utility obtained by gaining services. The objective of the RUN is to maximize the total Quality of Cooperation (QoC) of users in the RUN. Two optimal bandwidth allocation schemes are proposed for both free and dynamic bandwidth demand models. To make our scheme more pragmatic, we then consider a more practical scenario in which the bandwidth capacity of the relays and the minimum demand of the clients are bounded. We prove that the problems under both the single relay and the multi-relay scenario are NP-hard. Three heuristic algorithms are proposed to deal with bandwidth allocation and relay-client association. We also propose a distributed signaling protocol and divide the centralized MRMC algorithm into three distributed ones to better adapt for real network environment. Finally, extensive simulations demonstrate that our RUN framework can significantly improve the efficiency of cooperation in the long term.
Yong Cui 0001, Xiao Ma 0009, Xiuzhen Cheng, Minming Li, Jiangchuan Liu, Tianze Ma, Yihua Guo, Biao Chen 0002
IEEE Trans. Parallel Distributed Syst.1
2014 Localized routing optimization for multi-access Mobile Nodes in PMIPv6
abstract
Proxy Mobile IPv6(PMIPv6) [1], as one of the most promising technologies for the next generation IP network, has attracted much attention in academia and industry. As an extension of PMIPv6, localized routing optimization can improve the performance of local communication greatly. However, recent study on localized routing optimizations only focus on single-access scenario, while the multi-access scenario has not been widely studied yet. In this paper, (1) we analyze the local communication scenarios for multi-access Mobile Node(MN), and propose two localized routing optimization schemes in the scenarios that MNs have interfaces attached to the same Mobile Access Gateway(MAG). (2) We then propose a localized routing selection algorithm to enable LMA to choose the best transmission path between MNs. (3) Besides, we also provide a performance analysis for the localized routing optimization schemes, and the results show that our scheme can greatly reduce the transmission cost and traffic load.
Yong Cui 0001
ICCCN4
2014 Towards a system theoretic approach to wireless network capacity in finite time and space
abstract
In asymptotic regimes, both in time and space (network size), the derivation of network capacity results is grossly simplified by brushing aside queueing behavior in nonJackson networks. This simplifying double-limit model, however, lends itself to conservative numerical results in finite regimes. To properly account for queueing behavior beyond a simple calculus based on average rates, we advocate a system theoretic methodology for the capacity problem in finite time and space regimes. This methodology also accounts for spatial correlations arising in networks with CSMA/CA scheduling and it delivers rigorous closed-form capacity results in terms of probability distributions. Unlike numerous existing asymptotic results, subject to anecdotal practical concerns, our transient results can be used in practical settings, e.g., to compute the time scales at which multi-hop routing is more advantageous than single-hop routing.
Florin Ciucu, Ramin Khalili, Yuming Jiang 0001, Yong Cui 0001
INFOCOM5
2014 Performance-aware energy optimization on mobile devices in cellular network
abstract
In cellular networks, it is important to conserve energy while at the same time ensuring users to have good transmission experiences. The energy cost can result from tail energy due to the radio resource control strategies designed in cellular networks and data transmission. Existing efforts generally consider one of the energy issues, and also ignore the adverse impact on user transmission performance due to energy conservation. In addition, many existing algorithms are based on prediction and knowledge on future traffic, which are hard to apply in a practical wireless system with dynamic user traffic and channel condition. The goal of this work is to design an efficient online scheduling algorithm to minimize energy consumption both due to tail energy and transmissions while meeting user performance expectation. We prove the problem to be NP-hard, and design a practical online scheduling algorithm PerES to minimize the total energy cost of multiple mobile applications subject to user performance constraints. We propose a comprehensive performance cost metric to capture the impacts due to task delay, deadline violation, different application profiles and user preferences. We prove that our proposed scheduling algorithm can make the energy consumption arbitrarily close to that of the optimal scheduling solution. The evaluation results demonstrate the effectiveness of our scheme and its higher performance than peers. Moreover, by supporting dynamic performance requirement by mobile users, PerES can achieve 2 times faster convergence to both the performance degradation bound and optimal energy conversation bound than those of traditional static methods. Using 821 million traffic flows collected from a commercial cellular carrier, we verify our scheme could achieve on average 32%-56% energy savings with different levels of user experience.
Yong Cui 0001, Shihan Xiao, Xin Wang 0001, Minming Li, Hongyi Wang 0004, Zeqi Lai
INFOCOM1
2014 Energy-traffic tradeoff cooperative offloading for mobile cloud computing
abstract
This paper presents a quantitative study on the energy-traffic tradeoff problem from the perspective of entire Wireless Local Area Network (WLAN). We propose a novel Energy-Efficient Cooperative Offloading Model (E2COM) for energy-traffic tradeoff, which can ensure the fairness of energy consumption of mobile devices and reduce the computation repetition and eliminate the Internet data traffic redundancy through cooperative execution and sharing computation results. We design an Online Task Scheduling Algorithm (OTS) based on a pricing mechanism and Lyapunov optimization to address the problem without predicting future information on task arrivals, transmission rates and so on. OTS can achieve a desirable tradeoff between the energy consumption and Internet data traffic by appropriately setting the tradeoff coefficient. Simulation results demonstrate that E2COM is more efficient than no offloading and cloud offloading for a variety of typical mobile devices, applications and link qualities in WLAN.
Yong Cui 0001, Minming Li, Jiezhong Qiu, Rajkumar Buyya
IWQoS2
2014 Tolerating path heterogeneity in multipath TCP with bounded receive buffers
Ming Li 0035, Andrey Lukyanenko, Sasu Tarkoma, Yong Cui 0001, Antti Ylä-Jääski
Comput. Networks4
2014 Policy-based flow control for multi-homed mobile terminals with IEEE 802.11u standard
Yong Cui 0001, Xiao Ma 0009, Jiangchuan Liu, Yuri Ismailov
Comput. Commun.1
2014 Robust and Adaptive Scheduling of Sequential Periodic Sensing for Cognitive Radios
abstract
Spectrum sensing is a crucial element of dynamic spectrum access (DSA) as it enables cognitive radios (CRs) to opportunistically access the under-utilized spectrum. Existing efforts on sensing have not adequately addressed sensing scheduling over time for better detection performance. In this work, we consider sequential periodic sensing of an in-band channel. We focus primarily on finding the appropriate sensing frequency during an SU's active data transmission on a licensed channel. Detection schemes addressing channel state change and anomalous data are designed specifically to facilitate short-term sensing adaptation to the variations in sensed data. In addition, long-term adaptation is also considered so that the evolving sensing environment can be reflected in the sensing schedule as well. Simulation results demonstrate that our design guarantees better conformity to the spectrum access policies by significantly reducing the delay in change detection while ensuring better sensing accuracy.
Qiang Liu 0007, Xin Wang 0001, Yong Cui 0001
IEEE J. Sel. Areas Commun.3
2014 Reliable Multicast in Data Center Networks
abstract
Multicast benefits data center group communication in both saving network traffic and improving application throughput. Reliable packet delivery is required in data center multicast for data-intensive computations. However, existing reliable multicast solutions for the Internet are not suitable for the data center environment, especially with regard to keeping multicast throughput from degrading upon packet loss, which is norm instead of exception in data centers. We present RDCM, a novel reliable multicast protocol for data center network. The key idea of RDCM is to minimize the impact of packet loss on the multicast throughput, by leveraging the rich link resource in data centers. A multicast-tree-aware backup overlay is explicitly built on group members for peer-to-peer packet repair. The backup overlay is organized in such a way that it causes little individual repair burden, control overhead, as well as overall repair traffic. RDCM also realizes a window-based congestion control to adapt its sending rate to the traffic status in the network. Simulation results in typical data center networks show that RDCM can achieve higher application throughput and less traffic footprint than other representative reliable multicast protocols. We have implemented RDCM as a user-level library on Windows platform. The experiments on our test bed show that RDCM handles packet loss without obvious throughput degradation during high-speed data transmission, gracefully respond to link failure and receiver failure, and causes less than 10% CPU overhead to data center servers.
Dan Li 0001, Mingwei Xu 0001, Ying Liu 0024, Yong Cui 0001, Guihai Chen
IEEE Trans. Computers5
2014 Modeling Energy Consumption of Data Transmission Over Wi-Fi
abstract
Wireless data transmission consumes a significant part of the overall energy consumption of smartphones, due to the popularity of Internet applications. In this paper, we investigate the energy consumption characteristics of data transmission over Wi-Fi, focusing on the effect of Internet flow characteristics and network environment. We present deterministic models that describe the energy consumption of Wi-Fi data transmission with traffic burstiness, network performance metrics like throughput and retransmission rate, and parameters of the power saving mechanisms in use. Our models are practical because their inputs are easily available on mobile platforms without modifying low-level software or hardware components. We demonstrate the practice of model-based energy profiling on Maemo, Symbian, and Android phones, and evaluate the accuracy with physical power measurement of applications including file transfer, web browsing, video streaming, and instant messaging. Our experimental results show that our models are of adequate accuracy for energy profiling and are easy to apply.
Yu Xiao 0001, Yong Cui 0001, Petri Savolainen, Matti Siekkinen, Antti Ylä-Jääski, Sasu Tarkoma
IEEE Trans. Mob. Comput.2
2014 AP Association for Proportional Fairness in Multirate WLANs
abstract
In this paper, we investigate the problem of achieving proportional fairness via access point (AP) association in multirate WLANs. This problem is formulated as a nonlinear programming with an objective function of maximizing the total user bandwidth utilities in the whole network. Such a formulation jointly considers fairness and AP selection. We first propose a centralized algorithm Non-Linear Approximation Optimization for Proportional Fairness (NLAO-PF) to derive the user-AP association via relaxation. Since the relaxation may cause a large integrality gap, a compensation function is introduced to ensure that our algorithm can achieve at least half of the optimal in the worst case. This algorithm is assumed to be adopted periodically for resource management. To handle the case of dynamic user membership, we propose a distributed heuristic Best Performance First (BPF) based on a novel performance revenue function, which provides an AP selection criterion for newcomers. When an existing user leaves the network, the transmission times of other users associated with the same AP can be redistributed easily based on NLAO-PF. Extensive simulation study has been performed to validate our design and to compare the performance of our algorithms to those of the state of the art.
Wei Li 0059, Shengling Wang 0001, Yong Cui 0001, Xiuzhen Cheng, Ran Xin, Mznah Al-Rodhaan, Abdullah Al-Dhelaan
IEEE/ACM Trans. Netw.3
2014 Safe and Practical Energy-Efficient Detour Routing in IP Networks
abstract
The Internet is generally not energy-efficient since all network devices are running all the time and only a small fraction of consumed power is actually related to traffic forwarding. Existing studies try to detour around links and nodes during traffic forwarding to save powers for energy-efficient routing. However, energy-efficient routing in traditional IP networks is not well addressed. The most challenges within an energy-efficient routing scheme in IP networks lie in safety and practicality. The scheme should ensure routing stability and loop- and congestion-free packet forwarding, while not requiring modifications in the traditional IP forwarding diagram and shortest-path routing protocols. In this paper, we propose a novel energy-efficient routing approach called safe and practical energy-efficient detour routing (SPEED) for power savings in IP networks. We provide theoretical insight into energy-efficient routing and prove that determining if energy-efficient routing exists is NP-complete. We develop a heuristic in SPEED to maximize pruned links in computing energy-efficient routings. Extensive experimental results show that SPEED significantly saves power consumptions without incurring network congestions using real network topologies and traffic matrices.
Qi Li 0002, Mingwei Xu 0001, Yuan Yang 0001, Lixin Gao 0001, Yong Cui 0001
IEEE/ACM Trans. Netw.5
2014 A Model Approach to the Estimation of Peer-to-Peer Traffic Matrices
abstract
Peer-to-Peer (P2P) applications have witnessed an increasing popularity in recent years, which brings new challenges to network management and traffic engineering (TE). As basic input information, P2P traffic matrices are of significant importance for TE. Because of the excessively high cost of direct measurement, many studies aim to model and estimate general traffic matrices, but few focus on P2P traffic matrices. In this paper, we propose a model to estimate P2P traffic matrices in operational networks. Important factors are considered, including the number of peers, the localization ratio of P2P traffic, and the network distance. Here, the distance can be measured with AS hop counts or geographic distance. To validate our model, we evaluate its performance using traffic traces collected from both the real P2P video-on-demand (VoD) and file-sharing applications. Evaluation results show that the proposed model outperforms the other two typical models for the estimation of the general traffic matrices in several metrics, including spatial and temporal estimation errors, stability in the cases of oscillating and dynamic flows, and estimation bias. To the best of our knowledge, this is the first research on P2P traffic matrices estimation. P2P traffic matrices, derived from the model, can be applied to P2P traffic optimization and other TE fields.
Ke Xu 0002, Meng Shen 0001, Yong Cui 0001, Mingjiang Ye, Yifeng Zhong
IEEE Trans. Parallel Distributed Syst.3
2014 Path diversified multi-QoS optimization in multi-channel wireless mesh networks
Xiaoyuan Guo, Feng Wang 0001, Jiangchuan Liu, Yong Cui 0001
Wirel. Networks4
2013 Scheduling of sequential periodic sensing for cognitive radios
abstract
Spectrum sensing enables cognitive radios (CRs) to opportunistically access the under-utilized spectrum. Existing efforts on sensing have not adequately addressed sensing scheduling over time for better detection performance. In this work, we consider sequential periodic sensing of an in-band channel. We focus primarily on finding the appropriate sensing frequency during an SU's active data transmission on a licensed channel. Change and outlier detection schemes are designed specifically to facilitate short-term sensing adaptation to the variations in sensed data. Simulation results demonstrate that our design guarantees better conformity to the spectrum access policies by significantly reducing the delay in change detection while ensuring better sensing accuracy.
Qiang Liu 0007, Xin Wang 0001, Yong Cui 0001
INFOCOM3
2013 Tolerating path heterogeneity in multipath TCP with bounded receive buffers
abstract
No abstract available.
Ming Li 0035, Andrey Lukyanenko, Sasu Tarkoma, Yong Cui 0001, Antti Ylä-Jääski
SIGMETRICS4
2013 Spectrum Assignment and Sharing for Delay Minimization in Multi-Hop Multi-Flow CRNs
abstract
This paper investigates the problem of spectrum assignment and sharing to minimize the total delay of multiple concurrent flows in multi-hop cognitive radio networks. We first analyze the expected per-hop delay, which incorporates the sensing delay and transmission delay characterizing the PU activities and spectrum capacities. Then we formulate a minimum delay optimization problem with interference constraints, and propose an approximation algorithm termed MCC to solve the problem. According to our theoretical analysis, MCC has a bounded performance ratio and a low computational complexity. Finally, we exploit the minimum potential delay fairness in spectrum sharing to mitigate the inter-flow contentions. Extensive simulation study has been performed to validate our design and to compare the performance of our algorithms with that of the state-of-the-art.
Wei Li 0059, Xiuzhen Cheng, Yong Cui 0001, Wendong Wang 0003
IEEE J. Sel. Areas Commun.4
2013 Data Centers as Software Defined Networks: Traffic Redundancy Elimination with Wireless Cards at Routers
abstract
We propose a novel architecture of data center networks (DCN), which adds wireless network card to both servers and routers. Existing traffic redundancy elimination (TRE) mechanisms reduce link loads and increase network capacity in several environments by removing strings that have appeared in earlier packets through encoding and decoding them several hops downstream. This article is the first to explore TRE mechanisms in large-scale DCNs and the first to exploit cooperative TRE among servers. Moreover, it also achieves the `logically centralized' control over the physically distributed states in emerging software defined networks (SDN) paradigm, by sharing information among servers and routers in data centers with wireless cards. We first formulate the TREDaCeN (TRE in Data Center Networks) problem and reduce the cycle cover problem to prove that finding an optimal caching task assignment for TREDaCeN problem is NP-hard. We further describe an offline TREDaCeN algorithm which is proved to have good approximation ratio. We then discuss efficient online zero-delay and semi-distributed implementations of TREDaCeN supported by physical proximity of servers and routers, enabling status updates in a single wireless transmission, using an efficient prioritized schedule. We also address online cache replacement and consistency of information in servers and routers with and without delay. Our framework is tested on different parameters and shows superior performance in comparison to other mechanisms (imported directly to this setting). Our results show the robustness and the trade-off between the `logically centralized' implementation and the overhead on handling inconsistency of distributed information in DCN.
Yong Cui 0001, Shihan Xiao, Chunpeng Liao, Ivan Stojmenovic, Minming Li
IEEE J. Sel. Areas Commun.1
2013 A Survey of Energy Efficient Wireless Transmission and Modeling in Mobile Cloud Computing
Yong Cui 0001, Xiao Ma 0009, Hongyi Wang 0004, Ivan Stojmenovic, Jiangchuan Liu
Mob. Networks Appl.1
2013 Load-balanced AP association in multi-hop wireless mesh networks
Yong Cui 0001, Tianze Ma, Jiangchuan Liu, Sajal K. Das 0001
J. Supercomput.1
2013 Dynamic Scheduling for Wireless Data Center Networks
abstract
Unbalanced traffic demands of different data center applications are an important issue in designing data center networks (DCN). In this paper, we present our exploratory investigation on a hybrid DCN solution of utilizing wireless transmissions in DCNs. Our work aims to solve the congestion problem caused by a few hot nodes to improve the global performance. We model the wireless transmissions in DCN by considering both the wireless interference and the adaptive transmission rate. Besides, both throughput and job completion time are considered to measure the impact of wireless transmissions on the global performance. Based on the model, we formulate the problem of channel allocation as an optimization problem. We also design an approximation algorithm with an approximation bound of 1/2 and a genetic algorithm to address the scheduling problem. A series of simulations are performed to evaluate and demonstrate the effectiveness of our wireless DCN scheme.
Yong Cui 0001, Hongyi Wang 0004, Xiuzhen Cheng, Dan Li 0001, Antti Ylä-Jääski
IEEE Trans. Parallel Distributed Syst.1
2013 Guest Editors' Introduction: Special Issue on Cloud Computing
abstract
The articles in this special section focus on the topic of cloud computing, technologies, applications, and new areas of technological innovation.
Vojislav B. Misic, Rajkumar Buyya, Dejan S. Milojicic, Yong Cui 0001
IEEE Trans. Parallel Distributed Syst.4
2012 FMTCP: A Fountain Code-Based Multipath Transmission Control Protocol
abstract
Ideally, the throughput of a Multipath TCP (MPTCP) connection should be as high as that of multiple disjoint single-path TCP flows. In reality, the throughput of MPTCP is far lower than expected. This is fundamentally caused by the fact that a sub flow with high delay and loss affects the performance of other sub flows, and thus becomes the bottleneck of the MPTCP connection and significantly degrades the aggregate good put. To tackle this problem, we propose Fountain code-based Multipath TCP (FMTCP), which effectively mitigates the negative impact of the heterogeneity of different paths. FMTCP takes advantage of the random nature of the fountain code to flexibly transmit encoded symbols from the same or different data blocks over different sub flows. Moreover, we design a data allocation algorithm based on the expected packet arriving time and decoding demand to coordinate the transmissions of different sub flows. Quantitative analyses are provided to show the benefit of FMTCP. We also evaluate the performance of FMTCP through ns-2 simulations and demonstrate that FMTCP can outperform IETF-MPTCP, a typical MPTCP approach, when the paths have diverse loss and delay in terms of higher total good put, lower delay and jitter. In addition, FMTCP achieves much more stable performance under abrupt changes of path quality.
Yong Cui 0001, Xin Wang 0001, Hongyi Wang 0004, Guangjin Pan
ICDCS1
2012 Dynamic region-based mobile multicast
abstract
Abstract Traditional mobile multicast schemes have higher multicast tree reconfiguration cost or multicast packet delivery cost. Two costs are very critical because the former affects the service disruption time during handoff while the latter affects the packet delivery delay. Although the range‐based mobile multicast (RBMoM) scheme and its similar schemes offer the trade‐off between two costs to some extent, most of them do not determine the size of service region, which is critical to the network performance. Hence, we propose a dynamic region‐based mobile multicast (DRBMoM) to dynamically determine the optimal service region for reducing the multicast tree reconfiguration and multicast packet delivery costs. DRBMoM provides two versions: (i) the per‐user version, named DRBMoM‐U, and (ii) the aggregate‐users version, named DRBMoM‐A. Two versions have different applicability, which are the complementary technologies for pursuing efficient mobile multicast. Though having different data information and operations, two versions have the same method for finding the optimal service region. To that aim, DRBMoM models the users' mobility with arbitrary movement directional probabilities in 2‐D mesh network using Markov Chain, and predicts the behaviors of foreign agents' (FAs') joining in a multicast group. DRBMoM derives a cost function to formulate the average multicast tree reconfiguration cost and the average multicast packet delivery cost, which is a function of service region. DRBMoM finds the optimal service region that can minimize the cost function. The simulation tests some key parameters of DRBMoM. In addition, the simulation and numerical analyses show the cost in DRBMoM is about 22∼50% of that in RBMoM. At last, the applicability and computational complexity of DRBMoM and its similar scheme are analyzed. Copyright © 2010 John Wiley & Sons, Ltd.
Shengling Wang 0001, Yong Cui 0001, Sajal K. Das 0001
Wirel. Commun. Mob. Comput.2
2011 Partially overlapping channel assignment based on "node orthogonality" for 802.11 wireless networks
abstract
In this study, we investigate the problem of partially overlapping channel assignment to improve the performance of 802.11 wireless networks. We first derive a novel interference model that takes into account both the adjacent channel separation and the physical distance of the two nodes employing adjacent channels. This model defines “node orthogonality”, which states that two nodes over adjacent channels are orthogonal if they are physically sufficiently separated. We propose an approximate algorithm MICA to minimize the total interference for throughput maximization. Extensive simulation study has been performed to validate our design and to compare the performances of our algorithm with those of the state-of-the-art.
Yong Cui 0001, Wei Li 0059, Xiuzhen Cheng
INFOCOM1
2011 Multi-hop access pricing in public area WLANs
abstract
Public area WLANs stand for WLANs deployed in public areas such as classroom and office buildings to provide Internet connections. Nevertheless, such a service may not be always available because of limited AP coverage, poor signal strength, or password authentication. Multi-hop access is a feasible approach to facilitate users without direct AP accesses to resort to other online users as relays for data forwarding. This paper employs credit-exchange for multi-hop access in public area WLANs to encourage users to cooperate, and proposes a complete pricing framework. We first investigate a revenue model to define the profit of a relay. Next we point out that cutoff bandwidth allocation is a crucial issue in pricing strategy. Optimal bandwidth allocation schemes are then proposed for two bandwidth demand models. Following that we consider a more practical scenario where the relay's bandwidth capacity and the client's bandwidth demand are bounded, and propose two heuristic algorithms SRMC and MRMC to compute bandwidth allocation and/or relay-client association. Extensive simulation study has been performed to validate our design.
Yong Cui 0001, Tianze Ma, Xiuzhen Cheng
INFOCOM1
2011 Channel allocation in wireless data center networks
abstract
Unbalanced traffic demands of different data center applications are an important issue in designing Data center networks (DCNs). In this paper, we present our exploratory investigation of utilizing wireless transmissions in DCNs. Our work aims to solve the congestion problem caused by a few hot nodes to improve the global performance. We model the wireless transmissions in a DCN by considering both the wireless interference and the adaptive transmission rate. Moreover, both throughput and job completion time are taken into account to evaluate the impact of wireless transmissions on the global performance. Based on this model, we formulate the channel allocation in wireless DCNs as an optimization problem and design a genetic algorithm (GA) based approach to address it. To demonstrate the effectiveness of wireless transmissions as well as our GA-based algorithm in a wireless DCN, extensive simulation study is carried out and the results validate our design.
Yong Cui 0001, Hongyi Wang 0004, Xiuzhen Cheng
INFOCOM1
2011 Sparse target counting and localization in sensor networks based on compressive sensing
abstract
In this paper, we propose a novel compressive sensing (CS) based approach for sparse target counting and positioning in wireless sensor networks. While this is not the first work on applying CS to count and localize targets, it is the first to rigorously justify the validity of the problem formulation. Moreover, we propose a novel greedy matching pursuit algorithm (GMP) that complements the well-known signal recovery algorithms in CS theory and prove that GMP can accurately recover a sparse signal with a high probability. We also propose a framework for counting and positioning targets from multiple categories, a novel problem that has never been addressed before. Finally, we perform a comprehensive set of simulations whose results demonstrate the superiority of our approach over the existing CS and non-CS based techniques.
Bowu Zhang, Xiuzhen Cheng, Nan Zhang 0004, Yong Cui 0001, Yingshu Li 0001, Qilian Liang
INFOCOM4
2011 Load Balancing Access Point Association Schemes for IEEE 802.11 Wireless Networks
Yuan Le, Liran Ma, Xiuzhen Cheng, Yong Cui 0001, Mznah Al-Rodhaan, Abdullah Al-Dhelaan
WASA5
2011 Impact of user selfishness in construction action on the streaming quality of overlay multicast
Dan Li 0001, Yong Cui 0001, Jiangchuan Liu, Ke Xu 0002
Comput. Networks3
2011 Distributed dynamic mobile multicast
Yong Cui 0001, Shengling Wang 0001, Sajal K. Das 0001
J. Parallel Distributed Comput.1
2011 Defending Against Distance Cheating in Link-Weighted Application-Layer Multicast
abstract
Application-layer multicast (ALM) has recently emerged as a promising solution for diverse group-oriented applications. Unlike dedicated routers in IP multicast, the autonomous end-hosts are generally unreliable and even selfish. A strategic host might cheat about its private information to affect protocol execution and, in turn, to improve its individual benefit. Specifically, in a link-weighted ALM protocol where the hosts measure the distances from their neighbors and accordingly construct the ALM topology, a selfish end-host can easily intercept the measurement message and exaggerate the distances to other nodes, so as to reduce the probability of being a relay. Such distance cheating, rarely happening in IP multicast, can significantly impact the efficiency and stability of the ALM topology. To defend against this kind of cheating, we present a Vickrey–Clarke–Groves (VCG)-based cheat-proof mechanism in this paper. We demonstrate a practical mapping from the utility, payment, and welfare of a VCG mechanism to the link-weighted ALM context. Based on this, we further discuss practical issues for implementing the cheat-proof mechanism—specifically, a trustworthy distributed algorithm for payment computation. Performance analyses show that the overheads of the computation, storage, and communication of our implementation are controlled at low levels, and extensive simulations further testify the implementation's effectiveness. Although there are other similar studies in this area, the contribution of our cheat-proof mechanism and its implementation primarily lies in two aspects. On one hand, we first explicitly solve the distance cheating problem in link-weighted ALM since its proposal by mapping the VCG mechanism to link-weighted ALM context. On the other hand, our distributed implementation can not only effectively defend against distance cheating, but can also avoid the potential cheating behaviors when selfish ALM nodes fulfill the cheat-proof mechanism itself.
Dan Li 0001, Jiangchuan Liu, Yong Cui 0001, Ke Xu 0002
IEEE/ACM Trans. Netw.4
2011 Mobility in IPv6: Whether and How to Hierarchize the Network?
abstract
Mobile IPv6 (MIPv6) offers a basic solution to support mobility in IPv6 networks. Although Hierarchical MIPv6 (HMIPv6) has been designed to enhance the performance of MIPv6 by hierarchizing the network, it does not always outperform MIPv6. In fact, two solutions have different application scopes. Existing work studies the impact of various parameters on the performance of MIPv6 and HMIPv6, but without analyzing their application scopes. In this paper, we propose a model to analyze the application scopes of MIPv6 and HMIPv6, through which an Optimal Choice of Mobility Management (OCMM) scheme is designed. Different from the existing work that either propose new mobility management schemes or enhance existing mobility management schemes, OCMM chooses the better alternative between MIPv6 and HMIPv6 according to the mobility and service characteristics of users, addressing whether to hierarchize the network. Besides that, OCMM chooses the best mobility anchor point and regional size when HMIPv6 is adopted, addressing how to hierarchize the network. Simulation results demonstrate the impact of key parameters on the application scopes of MIPv6 and HMIPv6 as well as the optimal regional size of HMIPv6. Finally, we show that OCMM outperforms MIPv6 and HMIPv6 in terms of total cost including average registration and packet delivery costs.
Shengling Wang 0001, Yong Cui 0001, Sajal K. Das 0001, Wei Li 0059
IEEE Trans. Parallel Distributed Syst.2
2011 Achieving Proportional Fairness via AP Power Control in Multi-Rate WLANs
abstract
In this paper, we consider how to achieve proportional fairness in multi-rate 802.11 WLANs by investigating an integrated problem of power control and AP Association in order to provide an effective tradeoff between network throughput and fairness. Since jointly considering power control and AP association for proportional fairness is NP-hard, we propose a centralized heuristic approach. By introducing a new concept of AP utility, we establish the relationship between the network utility and the AP utility according to proportional fairness. This relationship is exploited to design an algorithm PCAP to optimize the network utility by increasing the average and decreasing the variance of the AP utility. Extensive simulation study is performed and the results demonstrate that PCAP yields a significant improvement in terms of throughput, fairness, and power consumption compared to other popular power control algorithms.
Wei Li 0059, Yong Cui 0001, Xiuzhen Cheng, Mznah Al-Rodhaan, Abdullah Al-Dhelaan
IEEE Trans. Wirel. Commun.2
2010 IP Fast Reroute: NotVia with Early Decapsulation
abstract
Network survivability is an important topic for the Internet. To improve the performance of the Internet during failure, IP Fast Reroute (IPFRR) mechanisms are proposed to establish backup routes for failure-affected packets. NotVia, a most prominent one, provides 100% protection coverage for single-node failures. However, it brings in nontrivial computing and memory pressure to routers with special NotVia addresses, in which only some are necessary for a specific router. Besides, the protection path of NotVia is 20% longer than the optimal path on average. In this paper, we propose early decapsulated NotVia (ED-NotVia) handling the aforementioned problems and thus making NotVia more practical. We first analyze the properties of necessary NotVia addresses to any specific node. Then we develop a heuristic Nec-NotVia Algorithm for a node to find the necessary NotVia addresses and compute routes for them, where unnecessary addresses are eliminated. Based on this elimination, early decapsulation is imported to optimize the protection path with marginal overhead. We evaluate our algorithm and demonstrate the effectiveness of ED-NotVia using topologies from Rocketfuel and Brite. The results show that 1) only 5% to 20% of SPT(Shortest Path Tree)-related NotVia addresses (1.23% to 6.41% of all the NotVia addresses) in an AS are necessary for a node; 2) by computing the routes for 15% to 40% SPT-related NotVia addresses, ED-NotVia provides 98% protection coverage; and 3) the protection path stretch ratio of ED-NotVia is only 1.03 on average as compared to 1.20 for NotVia.
Qing Li 0006, Mingwei Xu 0001, Qi Li 0002, Dan Wang 0002, Yong Cui 0001
GLOBECOM5
2010 PET: Prefixing, Encapsulation and Translation for IPv4-IPv6 Coexistence
abstract
IPv6 transition problem has become one of the key factors which are holding up the development of the next generation Internet. Aiming to solve IPv6 transition problem, several translation and tunneling techniques have been proposed, satisfying the demand of IPv4-IPv6 interconnection and traversing respectively. However, translation techniques can't convert the semantic between IPv4 and IPv6 protocol perfectly, and they have serious limitations in operation complexity and scalability. Researchers tried to decompose, simplify these problems and improve translation techniques accordingly, but they've come to little achievement since these problems result from the very nature of translation. We propose a novel approach of choosing appropriate translation spot to solve these problems in a different angle, and hence make effective use of translation technique. Then we propose a framework for IPv4-IPv6 coexistence called PET, which integrates tunneling and translation to support both traversing and IPv4-IPv6 interconnection, and uses them properly to constitute communication models in different scenarios. Moreover, we put forward PET signaling method to achieve automatic translation spot election and translation context advertisement, as a complement to the framework.
Peng Wu 0007, Yong Cui 0001, Mingwei Xu 0001, Xing Li 0001, Chris Metz 0001, Shengling Wang 0001
GLOBECOM2
2010 Segment Level Authentication: Combating internet source spoofing
abstract
This paper presents SLA (Segment Level Authentication), a transport segment level solution designed to prevent both of the intra-domain and inter-domain source spoofing. SLA is based on public key cryptography authentication. It enables intermediate network nodes the ability to validate the packet authenticity by verifying authentication information carried in packets. Although public key cryptography is computationally intensive and induces the traffic overhead, SLA leverages FPGA (Field Programmable Gate Array) based ECC (Elliptic Curve Cryptography) hardware cryptography accelerator to decrease the computation and traffic overhead. SLA provides incremental deployment and offers incentives for both of hosts and ASes. We find that the SLA is feasible for Gigabit links and can effectively mitigate source spoofing in both of intra-domain and inter-domain networks.
Ming Li 0035, Matti Siekkinen, Sasu Tarkoma, Antti Ylä-Jääski, Yong Cui 0001
ISCC5
2010 Approximate Optimization for Proportional Fair AP Association in Multi-rate WLANs
Wei Li 0059, Yong Cui 0001, Shengling Wang 0001, Xiuzhen Cheng
WASA2
2010 Supporting multiple metrics in QoS-aware BGP
Yong Cui 0001, Youjian Zhao, Turgay Korkmaz, Tielei Zhang
Sci. China Inf. Sci.1
2009 Probabilistic routing for multiple flows in wireless multi-hop networks
abstract
Maximizing network throughput is one of the main objectives in wireless multi-hop networks. However, small link bandwidth and severe interference become two great obstacles in network throughput improvement. Single route always encounters great congestion, especially for multiple flows with the same source and destination. Multi-path routing can solve the bandwidth shortage issue to some extent, but conventional routing algorithms of this type have some limitations during route selection and using, which makes bandwidth improvement and interference reduction unfeasible at the same time. In this paper we propose a new metric called effective bandwidth to describe the real bandwidth that a flow can get through a specific route. Based on this metric, we present a probabilistic multi-path routing protocol for multiple flows, which can improve the network throughput by selecting route with bigger effective bandwidth using higher probability. The simulation results show that probabilistic multi-path routing has high superiority over single-path routing in improving network throughput.
Yong Cui 0001, Sasu Tarkoma, Antti Ylä-Jääski
LCN1
2009 Defending Against Buffer Map Cheating in DONet-Like P2P Streaming
abstract
Data-driven overlay network (DONet)-like P2P system is especially suitable to support live stream applications, since its data structure can tolerate node dynamics quite well. However, optimal streaming demands the cooperation of individual nodes. If selfish nodes cheat about their buffer maps to reduce the forwarding burden, the overall streaming quality would be negatively affected. To defend against this kind of cheating, we design a trustworthy service-differentiation based incentive mechanism with low complexity in this paper. The mechanism is composed of the service-differentiation algorithm and the contribution-evaluation algorithm. Compared with other studies in this area, the primary characteristic of our mechanism lies in two aspects. Firstly, the contribution of each node is evaluated considering the features of live streaming, not just by the transferring bytes. Secondly, the potential cheating behavior of overlay nodes during the fulfillment of incentive algorithms can be avoided, which is usually not considered by other similar studies. Extensive simulations suggest that the algorithms are indeed effective for defending against buffer map cheating in DONet-like P2P streaming.
Dan Li 0001, Yong Cui 0001
IEEE Trans. Multim.3
2008 Evading User-Specific Offensive Web Pages via Large-Scale Collaborations
abstract
Web pages polluted by unhealthy contents (e.g. pornography or violence) have offended many users and become a social headache. This paper presents a collaborative rating system and a light-weight algorithm to detect polluted pages and thus improve user experience of web browsing. It mainly tackles two challenges. First, the system should cater to web users' different tastes and judging standards on which polluted pages they like or dislike. Second, the system should be resilient to dishonest ratings and collusions. The model and the algorithm are evaluated by simulations which show that they can work well.
Mingwei Xu 0001, Xue Zhi Jiang, Yong Cui 0001
ICC4
2008 Intelligent Mobility Support for IPv6
abstract
Hierarchical MIPv6 (HMIPv6) is proposed to improve the system performance of Mobile IPv6 (MIPv6). However, HMIPv6 cannot outperform MIPv6 in all scenarios because of its double-registration when a user roams across regions and the longer packet delivery latency. Therefore, to select a proper mobility management scheme between MIPv6 and HMIPv6 becomes an interesting issue, for its potentials in enhancing the capacity and scalability of the system. In this paper, we develop an analytical model to analyze the applicability of MIPv6 and HMIPv6. Based on this model, we design an Intelligent Mobility Support (IMS) scheme that selects the better alternative between MIPv6 and HMIPv6 for a user according to its changing mobility and service characteristics. When HMIPv6 is adopted, IMS chooses the best mobility anchor point and regional size to optimize the system performance. Numerical results illustrate the impact of some key parameters on the applicability of MIPv6 and HMIPv6. Finally, it is demonstrated that IMS outperforms MIPv6 and HMIPv6.
Shengling Wang 0001, Yong Cui 0001, Sajal K. Das 0001
LCN2
2008 A Study of Path Protection in Self-Healing Routing
Qi Li 0002, Mingwei Xu 0001, Lingtao Pan, Yong Cui 0001
Networking4
2008 A Density Adaptive Routing Protocol for Large-Scale Ad Hoc Networks
abstract
Position-based routing protocols use location information to refine the traditional packet flooding method in mobile ad hoc networks. They mainly focus on densely and evenly distributed network scenarios, but their performances degrade quickly as networks become sparse. To overcome these shortcomings, we introduce the network density concept in the routing process, and propose a density adaptive routing protocol (DAR) to optimize the packet forwarding process. DAR algorithm utilizes the local network density to determine the packet forwarding zone; in dense areas, it narrows the forwarding range to reduce the total number of participants in flooding; in sparse area, it enlarges the forwarding scope to enclose enough nodes for packet relaying. Compared to the existing protocols, DAR yields lower flooding overhead and higher route discovery rate. Theoretical analysis proves that it works well in both densely and sparsely distributed network scenarios, and is more scalable than traditional routing protocols. Simulation results are presented in this paper to demonstrate the effectiveness of the proposed protocol.
Zhizhou Li, Yaxiong Zhao, Yong Cui 0001
WCNC3
2008 IETF softwire unicast and multicast framework for IPv6 transition
Yong Cui 0001, Mingwei Xu 0001, Xing Li 0001
Sci. China Ser. F Inf. Sci.1
2007 Truthful Streaming in Selfish DONet
abstract
Data-driven overlay network (DONet) is especially suitable for live stream because it can tolerant node dynamics well. However, optimal streaming demands the cooperation of individual nodes. If selfish nodes in DONet cheat about their buffer maps to reduce the forwarding burden, the overall streaming quality might be negatively affected. To defend this kind of cheating behavior, we design a trustworthy service- differentiation based incentive mechanism with low complexity in this paper. The mechanism is composed of the service- differentiation algorithm and the contribution-evaluation algorithm. Compared with other studies in this area, the primary characteristic of our mechanism lies in two aspects. Firstly, the contribution of each node is evaluated considering the characteristic of live stream, not just by the transferring bytes. Secondly, the potential cheating behavior of overlay nodes during the fulfillment of incentive algorithms can be defended, which is usually not considered by other studies.
Dan Li 0001, Yong Cui 0001
ICC3
2007 QoS-Aware Streaming in Overlay Multicast Considering the Selfishness in Construction Action
abstract
Most existing overlay multicast proposals have assumed that the nodes are cooperative and thus focus on the global topology optimization. However, a unique and important characteristic of overlay nodes is that, as application-layer agents, they can be selfish with their own interests. To achieve better quality-of-service (QoS) or to minimize forwarding overhead, an overlay node can behave selfishly in the information collection or in the overlay construction. While the former has recently been investigated, the impact of selfishness in the construction action remains unclear. In this paper, we present the first systematic study on the impact of selfishness in both tree and mesh overlay construction. Our investigation considers multiple QoS measures for streaming applications, including stream latency, resolution, and continuity. Our contribution is twofold: first, we analyze how for selfish overlay nodes to choose a construction-action policy to optimize their individual multi-metric QoS. Second, we demonstrate that the selfishness-aware policy for the construction action is consistent with the QoS optimization for the global multicast session, but not vice versa. The implication is significant: A globally optimal overlay construction itself can be vulnerable to individual selfishness; but, following our directions, we can design an overlay that is both globally optimal and selfish-resistant.
Dan Li 0001, Yong Cui 0001, Jiangchuan Liu
INFOCOM3
2006 Autonomic Interference Avoidance with Extended Shortest Path Algorithm
Yong Cui 0001, Hao Che, Constantino M. Lagoa, ZhiMei Zheng
ATC1
2006 Segment-sending Schedule in Data-driven Overlay Network
abstract
Data-driven overlay network is suitable for live-event streaming, because it can provide relatively-continuous streaming even in dynamic environment. In terms of improving streaming quality, prior work covered membership management, buffer map exchange, segment requesting schedule, etc. In this paper, we address the problem of segment-sending schedule on the segment-providing node, which may also affect the streaming quality. The schedule methods we discuss include FIFO schedule, lower-sequence favored schedule, and higher-sequence favored schedule. Simulation results show that if users care playing continuity much more than playing delay, the higher-sequence favored schedule brings the best streaming quality; however, if users care playing delay much more than playing continuity, lower-sequence favored schedule is the preferred choice. Through this work, we find another way to improve streaming quality in data-driven overlay network.
Dan Li 0001, Yong Cui 0001, Ke Xu 0002
ICC2
2006 An improved Wu-Manber multiple patterns matching algorithm
abstract
NIDS is a powerful tool to defense the malicious attacks over the Internet. For the purpose of detecting the attack online, NIDS must inspect the payload of the packets very fast to expose the malicious code. String matching is a very important module in NIDS. To raise the performance of the string matching algorithm, we introduce an improved Wu-Manber algorithm QWM in this article. It combined the method of QS algorithm and used the mismatch information during the patterns matching, reached the top shift distance, improved the performance. We compared QWM with Aho-Corasick, Commentz-Walter and Wu-Manber algorithms. The experiment results shows on large alphabet such as English text and Chinese text, QWM algorithm has better performance, is faster than the other three algorithms. It can be used in various fields, such as network content analysis, intrusion detection, and text retrieval.
Donghong Yang, Ke Xu 0002, Yong Cui 0001
IPCCC3
2006 Forwarding IPv4 Traffics in Pure IPv6 Backbone with Stateless Address Mapping
abstract
As the IPv6 networks rapidly deployed, the transition problem has been changed. Because the transition of applications is a steady process with a far longer duration in comparison to the transition of infrastructure, providers cannot give up the duty of serving IPv4 traffics for the time being. On the other hand, operating a dual-stack backbone is highly costing for large-scale deployment of new networks. In this paper, a new technique of forwarding IPv4 traffics in IPv6-only backbone is developed. Rather than the ever-existing tunneling approaches, the proposed one is automatic and stateless for end nodes, without explicit tunneling. The route entries for delivering IPv4 traffics are well aggregated on the borders of IPv6 autonomous systems. This makes the technique suitable for large-scale deployment in an inter-domain networking environment.
Maoke Chen, Xing Li 0001, Ang Li 0002, Yong Cui 0001
NOMS4
2006 An Adaptive Latency-Energy Balance Approach of MAC Layer in Wireless Sensor Networks
Jinniu Chen, Mingwei Xu 0001, Yong Cui 0001
WASA3
2006 Stability of ALM Tree with selfish receivers: A simulation study
Dan Li 0001, Yong Cui 0001
Comput. Commun.3
2005 Impact of receiver cheating on the stability of ALM tree
abstract
Application layer multicast (ALM) is an effective supplement to IP multicast, but it has the potential trouble of trust on end systems. For instance, multicast receivers may cheat in order to obtain a better position in the multicast tree. Receiver cheating may transform the multicast tree, and lead to its instability. We establish the cheating model of ALM receivers and analyze the stability of ALM tree when receiver cheating occurs. Simulation results show that receiver cheating has considerably negative effects on the stability of ALM tree. This discovery brings forward an issue in ALM study, that is, we should take receiver cheating into consideration to maintain a stable ALM tree when designing ALM protocols.
Dan Li 0001, Yong Cui 0001, Ke Xu 0002
GLOBECOM2
2005 Precomputation for intra-domain QoS routing
Yong Cui 0001, Ke Xu 0002
Comput. Networks1
2004 Simple quality-of-service path first protocol and modeling analysis
abstract
QoS (quality-of-service) control is one of the most important mechanisms in the next-generation Internet, where QoS routing (QoSR) is a promising solution. We propose a multi-constrained intradomain QoS routing protocol SQOSPF. The advantages of this protocol include easy implementation, multi-constrained QoS support, high-speed convergence and multiple QoSR algorithms support. Stochastic Petri net is employed to model SQOSPF and analyze impacts of update threshold and routing holding time upon the load of networks and routers. Extensive simulations show that choosing appropriate update threshold and routing holding time can excessively reduce the extra load and keep routing performance at the same time.
Shen Lin 0004, Mingwei Xu 0001, Ke Xu 0002, Yong Cui 0001, Youjian Zhao
ICC4
2003 Precomputation for finding paths with two additive weights
abstract
As the most challenging problems of the upcoming next-generation networks, 2-constrained quality of service routing (QoSR) is NP-complete problem, for which we propose a novel precomputation algorithm, LEFPA. This algorithm converts two additive weights to a single metric with linear energy functions (LEFs) and pre-computes QoS routing table with multiple (B) LEFs to further enhance its scalability. We first analyze the performance of LEFs and give a method to determine the feasible and unfeasible areas in the metric space for a QoS request. We then introduce the proposed LEFPA, whose computation complexity is O(B(m+nlogn+n)). Furthermore, we use three methods to evaluate the routing performance. Extensive simulations show that our LEFPA has both absolutely and competitively high performance.
Yong Cui 0001, Ke Xu 0002, Mingwei Xu 0001
ICC1
2003 Multi-constrained routing based on simulated annealing
abstract
Multi-constrained quality-of-service routing (QoSR) is to find a feasible path that satisfies multiple constraints simultaneously, as an NPC problem, which is also a big challenge for the upcoming next-generation networks. In this paper, we propose SA/spl I.bar/MCP, a novel heuristic algorithm, by applying simulated annealing to Dijkstra's algorithm. This algorithm first uses a nonlinear energy function to translate multiple QoS weights into a single metric and then seeks to find a feasible path by simulated annealing. The paper outlines simulated annealing algorithm and analyzes the problems met when we apply it to QoSR. Extensive simulations demonstrate that SA/spl I.bar/MCP has good scalability regarding both network size and the number of QoS constraints with high performance. Furthermore, when most QoS requests are feasible, the running time of SA/spl I.bar/MCP is about O(k(m+nlogn)), which is only k times that of the traditional Dijkstra's algorithm, where k is the number of QoS constraints.
Yong Cui 0001, Ke Xu 0002, Zhongchao Yu, Youjian Zhao
ICC1
2003 Precomputation for Multi-constrained QoS Routing in High-speed Networks
abstract
As one of the most challenging problems of the next-generation high-speed networks, quality-of- service routing (QoSR) with multiple (k) constraints is an NP-complete problem. In this paper, we propose a multiconstrained energy function-based precomputation algorithm, MEFPA. It cares each QoS weight to b degrees, and computes a number (B= C/sub b+k-2//sup k-1/) of coefficient vectors uniformly distributed in the k-dimensional QoS metric space to construct B linear energy functions. Using each LEF, it then converts k QoS constraints to a single energy value. At last, it uses Dijkstra's algorithm to create B least energy trees, based on which the QoS routing table is created. We first analyze the performance of energy functions with k constraints, and give the method to determine the feasible and unfeasible areas for QoS requests in the k-dimensional QoS metric space. We then introduce our MEFPA for k-constrained routing with the computation complexity of O(B(m+n+nlogn)). Extensive simulations show that, with few coefficient vectors, this algorithm performs well in both absolute performance and competitive performance. In conclusion, for its high scalability, high performance and simplicity, MEFPA is a promising QoSR algorithm in the next-generation high-speed networks.
Yong Cui 0001, Ke Xu 0002
INFOCOM1
2003 Adjustable multi-constrained routing with a novel evaluation method
abstract
Quality-of-service routing (QoSR) with multiple constraints, which seeks to find a feasible path satisfying multiple constraints simultaneously, is a challenging problem of the next-generation networks. For its NP-complete complexity, we propose an adjustable heuristic based on converting multiple QoS weights to a single metric with energy functions. By applying the breadth-first search (BFS) to Dijkstra's algorithm with adjustable depth, BFS _MCP (BFS for multi-constrained paths) can adjust its time complexity according to the CPU load on a router in real time. Thus, it has an extensive adaptability. Additionally, we propose a novel approach to performance evaluation by generating QoS constraints, named weight-proportion simulation. Generating QoS constraints similar to QoS applications, this method extends the original success ratio, only used in relative performance comparison, to the evaluation of absolute performance. By this method, extensive simulations show that BFS improves the performance greatly. The main contribution of the paper includes a heuristic for multi-constrained routing and a novel approach to performance evaluation.
Yong Cui 0001, Ke Xu 0002
IPCCC1