Hui Wang 0011

dblp:79/2151 · also Jessie Hui Wang · DBLP profile ↗
← Back
90ranked-venue papers
16as first author
53since 2021 · last 2026
0000-0002-7825-4137ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 63 · 11 first-author · 35 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Security and privacy · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AdaptPipe: Mitigating Runtime Bubbles via Granularity-Adaptive Scheduling under Memory Constraints
Yumeng Cui, Hui Wang 0011, Najila Liu, Ling Deng, Chuxuan Zeng, Jilong Wang 0001
INFOCOM2
2026 On the Multiple-Unicast Conjecture: Session Dominance
Yeqiao Hou, Hui Wang 0011, Zongpeng Li
INFOCOM3
2026 A Self-updating Checkpointing and Fast Failure Recovery System for Distributed LLM Training
Leyi Ye, Zhiyi Yao, Boliang Liu, Yuedong Xu 0001, Zeng Chuxuan, Ling Deng, Hui Wang 0011
INFOCOM8
2026 Crack in the Armor: Underlying Infrastructure Threats to RPKI Publication Point Reachability
Yunhao Liu 0001, Hui Wang 0011, Yuedong Xu 0001, Zongpeng Li, Jilong Wang 0001
NDSS2
2026 COACH: Adaptive Robust Human-Robot Collaboration for Efficient Smart Manufacturing
abstract
Modern smart manufacturing pipelines have pervasively collaborated human workers, mobile robots, and industrial Internet of Things (IIoT) in shared workspaces for versatile production tasks. Despite the promising capacity of individual entities, the performance of these IIoT systems largely relies on pipeline coordination, i.e., task dispatching between humans and robots, which is particularly challenging under heterogeneous physical constraints and complex environmental uncertainties. Nonetheless, existing works either rely on traditional operation frameworks that lack scalability for large-scale complex production, or propose customized solutions for fixed agent models, overlooking the evolving nature of IIoT environments. To address these limitations, this paper proposes COACH, a human-robot collaborative manufacturing system that enables robust constraint-aware coordination across humans, robots, and IIoT. Specifically, COACH designs a scalable contextual encoder to represent the evolving relationships among human and robot agents in dynamic heterogeneous graphs. With that, a novel experience-driven task dispatcher is developed, enabling both high-performance and computation-efficient policy generation concerning the status of IIoT. To accommodate changing human fatigue and pipeline scales, COACH further develops a curriculum-enhanced reinforcement learning module for efficient dispatcher adaptation. Extensive evaluations using both synthetic testbeds and real-world manufacturing datasets demonstrate that COACH improves the feasible ratio of manufacturing pipelines by up to 27.4% and achieves up to 13.9% improvement in time efficiency compared to competing baselines across diverse job scales and environmental settings.
Hui Wang 0011, Liekang Zeng, Zhiwen Yu 0001, Yao Zhang 0005, Di Duan, Mu Yuan, Bin Guo 0001, Guoliang Xing
SenSys1
2026 Link Prediction-Based Measurement Strategy for Efficient Topology Completeness Improvements
abstract
The Autonomous System (AS) level topology observed from current measurement infrastructures is far from complete. Although Looking Glass (LG) vantage points (VPs) that support BGP route queries can provide valuable topology information, the query rate limitations of LG VPs imply that blindly using the VPs to conduct more measurements to improve the topology completeness is inefficient, if not infeasible. In this paper, we try to improve the efficiency by designing a link prediction based measurement strategy, whose basic idea is to first predict where unseen AS links are likely to be located and then use the prediction results to guide the measurements toward a more complete AS-level topology. We formulate the prediction of unseen AS links as a matrix completion problem and develop a side-information assisted learning-based matrix completion method. The method exploits a neural network and utilizes carefully chosen AS attributes based on our understanding on Internet peering practices, thereby learning more expressive latent vectors and achieving outstanding prediction performance in our scenario. We then develop a measurement strategy which takes the link prediction results as guidance to achieve efficient topology completeness improvements. The strategy leverages several heuristics to estimate the utilities of different measurements and takes a greedy algorithm to select the most valuable measurements. Experiments show that our link prediction method can achieve a high AUC (Area Under the Receiver Operating Characteristic Curve) of 0.834 and the link-guided measurement strategy can discover 1.82 times more unseen links than those discovered from non-guided measurement strategies with an equal number of measurements.
Shuying Zhuang, Hui Wang 0011, Jilong Wang 0001, Changqing An, Yuedong Xu 0001, Tianhao Wu 0010
IEEE Trans. Netw.2
2025 InternetSim: A Fast and Memory-Efficient Internet-Scale Inter-Domain Routing Simulator
abstract
Existing inter-domain routing simulators often suffer from low simulation accuracy, slow processing speeds, and high memory consumption, hindering their ability to perform large-scale Internet simulations. In this paper, we present InternetSim, a multi-threaded simulator for Internet-scale inter-domain routing. InternetSim enables incremental computation to deal with policy changes by some ASes, thus significantly accelerating multi-iteration context-continuous routing simulations. The memory efficiency is also improved by designing compact data structures for routing tables and out-of-memory errors are prevented by offloading the data to disk when necessary. Our simulator can complete an iteration of Internet-scale simulation within 13 hours (21× speedup compared to C-BGP) and can complete an ''incremental computation'' iteration within 15 seconds. The memory requirement is only 45% of C-BGP (without offloading) and the simulator can work well for Internet-scale simulations on a server with only 64GB if offloading is always enabled.
Jiahong Lai, Hui Wang 0011, Yunhao Liu 0001, Jilong Wang 0001
IMC2
2025 MemFerry: A Fast and Memory Efficient Offload Training Framework with Hybrid GPU Computation
Zhiyi Yao, Zuning Liang, Yuedong Xu 0001, Jin Zhao 0001, Hui Wang 0011
INFOCOM5
2025 Upper bound on the predictability of rating prediction in recommender systems
En Xu, Zhiwen Yu 0001, Hui Wang 0011, Helei Cui, Yunji Liang, Bin Guo 0001
Inf. Process. Manag.4
2025 FingHV: Efficient Sharing and Fine-Grained Scheduling of Virtualized HPU Resources
abstract
While artificial intelligence (AI) technology has advanced in real-world applications, there is a strong motivation to develop hybrid systems where AI algorithms and humans collaborate, promoting more human-centered approaches in AI system design. This has led to the emergence of a novel human-machine computing (HMC) paradigm, which combines human cognitive abilities with machine computational power to create a collaborative computing framework that meets the demands of large-scale, complex tasks and enables human-machine symbiosis. Human processing units (HPUs) are crucial computing resources in HMC-oriented systems, and efficient HPU resource provisioning is key to boosting system performance. However, existing schemes often fail to assign tasks to the most suitable HPUs and optimize HPU utility, as they either cannot quantitatively measure skills or overlook utility concerns during task assignment and scheduling. To address these challenges, this article proposes a fine-grained HPU virtualization (FingHV) approach, which leverages virtualization techniques to improve flexibility, fairness, and utility in the provisioning process. The core idea is to use a tree-based skill model to precisely measure the levels and correlations of multiple skills within individual HPUs, and to apply a mixed time/event-based scheduling policy to maximize HPU utility. Specifically, we begin by proposing a hierarchical multiskill tree to model HPU skills and their correlations. Next, we formulate the HPU virtualization problem and present a fine-grained virtualization method, which includes a quality-driven HPU assignment process and a mixed time/event-based scheduling policy to improve resource-sharing efficiency. Finally, we evaluate FingHV on a synthetic dataset with varying task sizes and a real-world case. The results demonstrate that FingHV improves global matching quality by up to 39.7% and increases HPU utility by 11.2% compared to the baselines.
Hui Wang 0011, Zhiwen Yu 0001, Zhuoli Ren, Yao Zhang 0005, Jiaqi Liu 0002, Liang Wang 0017, Bin Guo 0001
IEEE Trans. Cybern.1
2025 Cost-Efficient FEC Scheme for Time-Sensitive Multi-Hop Transmissions in Overlay Networks
abstract
In pursuit of low latency, real-time communication (RTC) service providers usually use multi-hop overlay links worldwide to bypass congested links, especially for medium- and long-distance transmissions. In such multi-hop long-distance transmission scenarios, utilizing retransmission to recover lost packets can result in increased end-to-end latency. Therefore, Forward Error Correction (FEC) is viewed as a promising way to solve the loss problem. However, for multi-hop overlay transmission, existing FEC schemes either introduce a non-negligible processing delay at each hop or reduce the processing delay at the cost of a high coefficient overhead. In this work, we propose a multi-hop FEC scheme, i.e., FEC-OEM, which considers both processing delay and coefficient overhead. FEC-OEM is designed based on two observations we obtained from measurements. First, coefficient overhead can only be reduced through an implicit transmission way. Therefore, we design a modulation-based recoding module that enables implicit coefficient transmission and hop-by-hop recoding at the same time. Second, using on-the-fly computation is a promising way to reduce processing delay. Accordingly, we design an elimination method to make the modulation-based recoding can be carried out on-the-fly. Real-world experiments demonstrate that FEC-OEM can reduce the processing delay by up to 88% without increasing the coefficient overhead compared to state-of-the-art schemes. We also use FEC-OEM to transmit packets for applications with different loss tolerances, and the results show that FEC-OEM can improve the QoE more effectively than state-of-the-art coding schemes.
Chao Xu 0015, Hui Wang 0011, Jilong Wang 0001, Jun Zhang 0004
IEEE Trans. Mob. Comput.2
2025 Orchestrating Joint Offloading and Scheduling for Low-Latency Edge SLAM
abstract
Visual Simultaneous Localization and Mapping (vSLAM) is a prevailing technology for many emerging robotic applications. Achieving real-time SLAM on mobile robotic systems with limited computational resources is challenging because the complexity of SLAM algorithms increases over time. This restriction can be lifted by offloading computations to edge servers, forming the emerging paradigm ofedge-assisted SLAM. Nevertheless, the exogenous and stochastic input processes affect the dynamics of the edge-assisted SLAM system. Moreover, the requirements of clients on SLAM metrics change over time, exerting implicit and time-varying effects on the system. In this paper, we aim to push the limit beyond existing edge-assist SLAM by proposing a new architecture that can handle the input-driven processes and also satisfy clients’ implicit and time-varying requirements. The key innovations of our work involve a regional feature prediction method for importance-aware local data processing, a configuration adaptation policy that integrates data compression/decompression and task offloading, and an input-dependent learning framework for task scheduling with constraint satisfaction. Extensive experiments prove that our architecture improves pose estimation accuracy and saves up to 47% of communication costs compared with a popular edge-assisted SLAM system, as well as effectively satisfies the clients’ requirements.
Yao Zhang 0005, Yuyi Mao, Hui Wang 0011, Zhiwen Yu 0001, Song Guo 0001, Jun Zhang 0004, Liang Wang 0017, Bin Guo 0001
IEEE Trans. Mob. Comput.3
2025 Joint Semantic Extraction and Resource Optimization in Communication-Efficient UAV Crowd Sensing
abstract
With the integration of IoT and 5G technologies, UAV crowd sensing has emerged as a promising solution to overcome the limitations of traditional Mobile Crowd Sensing (MCS) in terms of sensing coverage. As a result, UAV crowd sensing has been widely adopted across various domains. However, existing UAV crowd sensing methods often overlook the semantic information within sensing data, leading to low transmission efficiency. To address the challenges of semantic extraction and transmission optimization in UAV crowd sensing, this paper decomposes the problem into two sub-problems: semantic feature extraction and task-oriented sensing data transmission optimization. To tackle the semantic feature extraction problem, we propose a semantic communication module based on Multi-Scale Dilated Fusion Attention (MDFA), which aims to balance data compression, classification accuracy, and feature reconstruction under noisy channel conditions. For transmission optimization, we develop a reinforcement learning-based joint optimization strategy that effectively manages UAV mobility, bandwidth allocation, and semantic compression, thereby enhancing transmission efficiency and task performance. Extensive experiments conducted on real-world datasets and simulated environments demonstrate the effectiveness of the proposed method, showing significant improvements in communication efficiency and sensing performance under various conditions.
Erhe Yang, Zhiwen Yu 0001, Yao Zhang 0005, Helei Cui, Zhaoxiang Huang, Hui Wang 0011, Jiaju Ren, Bin Guo 0001
IEEE Trans. Netw. Serv. Manag.6
2024 Collecting Self-reported Semantics of BGP Communities and Investigating Their Consistency with Real-world Usage
abstract
People can extract various kinds of information about the Internet from BGP routes tagged with BGP community values with known semantics. In this paper, we conduct a study on the following three issues related to BGP community semantics. First, we design a method to automatically collect self-reported semantics from the Internet and assemble the collected semantics described in natural language into a structured dictionary. The comparison with prior dictionaries shows many community values are exclusively covered by ours and many of them had been used when prior dictionaries were constructed, which confirms the effectiveness of our method. Second, based on this large-size dictionary, we are able to re-evaluate two recent algorithms designed for categorizing community values with unknown semantics, which is a task that, while easier than inferring the detailed semantics, is also very valuable. Our evaluation uncovers some issues within the algorithms that can contribute to their performance improvement. Third, we investigate the fundamental issue in extracting information using community semantics: whether ISPs' behavior is consistent with the published semantics. Our preliminary best-effort investigation reveals the potential risks of using the semantics of some categories of community values.
Yunhao Liu 0001, Tianhao Wu 0010, Hui Wang 0011, Jilong Wang 0001, Shuying Zhuang
IMC3
2024 Learning Automatic Team Coordination in Human-Machine Partnerships
abstract
As AI-enabled machines become increasingly prevalent, there is a strong impetus to harness the complementary strengths of humans and machines to enhance productivity and reduce costs in collaborative workspaces such as manufacturing and warehouses [1]. However, efficient team coordination remains challenging due to the heterogeneity of team agents and the dynamic nature of human agents. Existing exact methods often rely on assumptions and mathematical models, which struggle to scale and accurately predict time-varying human performance [2]. While offline Reinforcement Learning (RL) demonstrates potential, it is time-consuming and heavily reliant on training data, often limited in practical factory settings [3]. Therefore, a scalable and data-efficient team coordination method that considers the varying capabilities of heterogeneous agents in collaborative systems is urgently needed to facilitate effective human-machine partnerships.
Hui Wang 0011, Youcheng Zhang, Zhiwen Yu 0001, Yao Zhang 0005, Jiaqi Liu 0002, Bin Guo 0001
MSN1
2024 An Efficient FEC Scheme with SLA Consideration for Low Latency Transmissions
abstract
Forward Error Correction (FEC) is the preferred method for recovering lost packets in time-sensitive applications. The key for FEC to recover lost packets successfully is whether the number of redundant packets is sufficient. Due to imperfect loss prediction algorithms, existing FEC schemes, which set the number of redundant packets according to the prediction results, make transport service providers likely to fall into the awkward situation of either failing to recover lost packets or wasting a large amount of bandwidth. In this work, we propose P-FEC, an FEC scheme that can take SLA into consideration and empirically achieve the targeted decoding success rate while minimizing bandwidth waste. P-FEC combines intra- and inter-generation coding to balance decoding success rate and bandwidth waste, in which intra-generation provides quick but conservative recovery and inter-generation coding provides delayed but more efficient loss recovery services. We future profile the loss prediction errors to derive cumulative distribution functions of the errors for diverse network conditions, and then determine the parameters of intra- and inter-generation coding according to these distribution functions. The real-world transmission experiments empirically demonstrate that P-FEC can achieve the targeted decoding success rate while its bandwidth waste is only 1%-16% of the FEC with the code rate that can theoretically guarantee the target success rate. Furthermore, P-FEC can work well with computation light-weighted prediction algorithms although these algorithms have low accuracy, which makes it extremely useful for the transmission environment with limited computing resources.
Chao Xu 0015, Hui Wang 0011, Zongpeng Li, Jilong Wang 0001
NOMS2
2024 A First Look at FEC Code Rate Determination from a Computational Cost Perspective
abstract
Forward Error Correction (FEC) is the preferred method for recovering lost packets in time-sensitive applications. However, setting the code rate to adapt to changing network environment is still a significant but challenging problem. Recently, researchers explored deep learning (DL) methods to solve this problem, but DL-based methods come with increased computation overhead and decision-making latency compared to traditional methods, which may make these methods infeasible or less effective. Thus, real-time network (RTN) providers usually have trouble with these two questions when considering whether to adopt a DL-based method: (1). How much additional computing resources are required by each DL-based method? (2). If resource competition occurs, can these DL-based methods be able to make timely code rate decisions? In this paper, we evaluate existing proposed DL-based methods for code rate determination in terms of computational cost to answer these two questions. Particularly, we measure the computational resources and the time required by these DL-based methods for making one code rate decision. We find that among existing DL-based methods, some methods require several times or even tens of times more computational resources to achieve timely decision-making. Furthermore, we identify that three methods are prone to resource competition with existing FEC schemes, which indicates that RTN providers need to configure their computing resources carefully when selecting these methods. We further analyze the reasons behind this phenomenon and provide some design suggestions for RTN nroviders.
Chao Xu 0015, Hui Wang 0011, Shilin Xie, Jilong Wang 0001
WCNC2
2024 FedDP-SA: Boosting Differentially Private Federated Learning via Local Data Set Splitting
abstract
Federated learning (FL) emerges as an attractive collaborative machine learning framework that enables training of models across decentralized devices by merely exposing model parameters. However, malicious attackers can still hijack communicated parameters to expose clients’ raw samples resulting in privacy leakage. To defend against such attacks, differentially private FL (DPFL) is devised, which incurs negligible computation overhead in protecting privacy by adding noises. Nevertheless, the low model utility and communication efficiency makes DPFL hard to be deployed in the real environment. To overcome these deficiencies, we propose a novel DPFL algorithm called FedDP-SA (namely, federated learning with differential privacy by splitting Local data sets and averaging parameters). Specifically, FedDP-SA splits a local data set into multiple subsets for parameter updating. Then, parameters averaged over all subsets plus differential privacy (DP) noises are returned to the parameter server. FedDP-SA offers dual benefits: 1) enhancing model accuracy by efficiently lowering sensitivity, thereby reducing noise to ensure DP and 2) improving communication efficiency by communicating model parameters with a lower frequency. These advantages are validated through sensitivity analysis and convergence rate analysis. Finally, we conduct comprehensive experiments to verify the performance of FedDP-SA compared with other state-of-the-art baseline algorithms.
Xuezheng Liu, Yipeng Zhou, Di Wu 0001, Miao Hu 0001, Hui Wang 0011, Mohsen Guizani
IEEE Internet Things J.5
2024 hmOS: An Extensible Platform for Task-Oriented Human-Machine Computing
abstract
With rapid advancements in artificial intelligence (AI) technologies, AI-powered machines are increasingly capable of collaborating with humans to enhance decision-making in various human–machine collaboration scenarios, e.g., medical diagnosis, criminal justice, and autonomous driving. As a result, human–machine computing (HMC) has emerged as a promising computing paradigm that integrates the expertise of humans with the reliable data processing capabilities of machines. Using HMC to facilitate the processing of domain-specific tasks has a lot of potential, but is limited in system-level scalability, i.e., there is no one common easy-to-use interface. In this article, we present human-machine operating system(hmOS), an open extensible platform for researchers to experiment with HMC for investigating system-centric human–machine collaboration problems.hmOSsupports flexible human–machine collaboration on the strength of the quality-aware task decomposition and allocation. To achieve that, the underlying system architecture and runtime environment are first developed to build a foundational abstraction for the kernel ofhmOS. Second,hmOSfacilitates flexible human–machine collaboration through a suitability-based task allocation mechanism, quality estimation guided by fuzzy rules, and iterative feedback on result tuning. We implement the newly proposedhmOSin a prototype featuring interactive interfaces. Finally, we conduct extensive and realistic experiments to validate the effectiveness of our platform across diverse tasks, showcasing the broad feasibility ofhmOS.
Hui Wang 0011, Zhiwen Yu 0001, Yao Zhang 0005, Fan Yang 0040, Liang Wang 0017, Jiaqi Liu 0002, Bin Guo 0001
IEEE Trans. Hum. Mach. Syst.1
2024 Feature Matching Data Synthesis for Non-IID Federated Learning
abstract
Federated learning (FL) has emerged as a privacy-preserving paradigm that trains neural networks on edge devices without collecting data at a central server. However, FL encounters an inherent challenge in dealing with non-independent and identically distributed (non-IID) data among devices. To address this challenge, this paper proposes a hard feature matching data synthesis (HFMDS) method to share auxiliary data besides local models. Specifically, synthetic data are generated by learning the essential class-relevant features of real samples and discarding the redundant features, which helps to effectively tackle the non-IID issue. For better privacy preservation, we propose a hard feature augmentation method to transfer real features towards the decision boundary, with which the synthetic data not only improve the model generalization but also erase the information of real features. By integrating the proposed HFMDS method with FL, we present a novel FL framework with data augmentation to relieve data heterogeneity. The theoretical analysis highlights the effectiveness of our proposed data synthesis method in solving the non-IID challenge. Simulation results further demonstrate that our proposed HFMDS-FL algorithm outperforms the baselines in terms of accuracy, privacy preservation, and complexity saving on various benchmark datasets.
Zijian Li 0023, Yuchang Sun 0001, Jiawei Shao, Yuyi Mao, Hui Wang 0011, Jun Zhang 0004
IEEE Trans. Mob. Comput.5
2024 Live Migration of Video Analytics Applications in Edge Computing
abstract
In order to schedule resources efficiently or maintain applications' continuity for mobile customers, edge platforms often need to adaptively migrate the applications on them. However, our measurement shows that existing migration solutions cannot solve the issue of migrating video analytics applications in edge computing because the memory states of video analytics applications have different characteristics from other applications. We conduct a breakdown analysis of the memory states of video analytics applications, and propose to treat three types of states separately with three different techniques,i.e., warm-up, sync, and replay, to minimize the negative influence of migrations on application performance. Based on this idea, we implement a prototype system in which two new components,i.e.,state storeandsidecar, are designed to achieve near-transparent live migration with minimal application code modifications. Evaluation experiments demonstrate that the time of application interruption caused by migrating a video analytics application with our solution is less than 405ms, and our solution does not consume much resources.
Chenghao Rong, Hui Wang 0011, Jilong Wang 0001, Yipeng Zhou, Jun Zhang 0004
IEEE Trans. Mob. Comput.2
2024 Mean-Field Game-Based Task-Offloaded Load Balance for Industrial Mobile Edge Computing Systems Using Software-Defined Networking
abstract
Smart devices (SDs) used in the Industrial Internet of Things can generate computational tasks for processing the data generated during production. However, due to the limited processing power of SDs, it is necessary to transfer these computational tasks to more powerful devices for processing. To this end, we propose a Mobile Edge Computing (MEC) system based on a Software Defined Network (SDN) for SDs to offload their computational tasks. This MEC system includes multiple MEC servers to handle numerous SDs, which leads to load-balancing challenges among these servers. To tackle this problem, we develop a computational offloading model based on mean-field game theory and introduce a mean-field game-based load-balancing algorithm (MFGLB), which reduces processing latency and facilitates task scheduling through Multi-Agent Deep Reinforcement Learning. Each SD in the MEC system is considered a participant in the mean-field game, simplifying the complex stochastic game into a more manageable dual-agent game. We then prove the existence of Nash Equilibrium for this mean-field game. To evaluate the effectiveness of our MFGLB algorithm, we compare its performance with traditional load-balancing algorithms and a stochastic game-based load-balancing algorithm. Our experimental results demonstrate the superiority of MFGLB in reducing processing latency and addressing load imbalances.
Guowen Wu, Hui Wang 0011, Hong Zhang 0046, Yizhou Shen, Shigen Shen, Shui Yu 0001
IEEE Trans. Mob. Comput.2
2023 TinyG: Accurate IP Geolocation Using a Tiny Number of Probers
abstract
IP geolocation is essential for various applications. However, the reliability of IP geolocation databases has been proven to be inadequate. In recent years, the growing number of public probers has offered the potential for more accurate geolocation results through active measurement. The conventional practice is to probe the target IP address using all available probers and feed the measurement results to the active geolocation method. However, this practice is cost-inefficient and may trigger the anti-flood mechanism. Moreover, public probers typically impose user-level limits on the frequency and quantity of measurements. Therefore, it is important to reduce the average number of probers (ANP) selected for successfully probing each target. Researchers have discovered that geolocation accuracy primarily depends on the minimum delay between probers and the target. Inspired by that, we propose TinyG, a prober selection algorithm designed to reduce the ANP needed to find probers within a sufficiently small delay from the target. TinyG divides the probing process into multiple rounds and leverages previous measurement results to guide the selection of probers for subsequent rounds. Experimental results show that when the associated minimum delay is within 2 ms, various active geolocation methods can provide credible geolocation results. TinyG outperforms other algorithms in reducing the ANP needed to obtain credible results. Compared to using more than 1,300 probers, TinyG can achieve an ANP of 6.7 with only a 6% coverage loss of credible results.
Hui Wang 0011, Jilong Wang 0001, Peiran Wang
CNSM2
2023 Enhancing Federated Learning by One-Shot Transferring of Intermediate Features from Clients
abstract
Federated learning (FL) is an emerging paradigm using a parameter server (PS) to coordinate multiple decentralized clients for training a common model without exposing their raw data. Despite its amazing capability in preserving data privacy, FL confronts two significant challenges that have not been sufficiently addressed by existing works, which are: 1) heterogeneous data distributed on clients given that the PS cannot alter data locations. 2) limited computation resources of FL clients who may conduct model training with mobile devices. To tackle these challenges, we propose a novel federated one-shot transferring of intermediate features (FedOTF) algorithm. More specifically, FedOTF consists of two stages: feature extraction and model reconstruction. To overcome challenge 1, clients in stage 1 only collaboratively train a small model in a federated manner, with the objective to extract features. In stage 2, intermediate features (generated by the small model trained by stage 1) are transferred from clients to the PS so that a large model can be trained to overcome challenge 2. In addition, the design of FedOTF is robust, which can flexibly diminish the number of communication rounds of stage 1 when network capacity is limited, and reduce the amount of exposed features when privacy is concerned. To verify the superiority of FedOTF, we conduct comprehensive experiments with real datasets. The experiment results demonstrate that FedOTF can significantly improve the model utility of FL because a better model can be finally obtained by the PS without incurring heavy computational load on clients. Besides, we conduct robustness evaluation of FedOTF which can achieve stable performance when varying network capacity and privacy requirement.
Youxingzhu Deng, Yipeng Zhou, Gang Liu 0028, Hui Wang 0011, Shui Yu 0001
DSAA4
2023 Top AS Router Geolocation in Databases: Performance and Techniques
abstract
Autonomous systems (ASes) at the top of the global transit hierarchy play core roles in the Internet. Thus, accurately geolocating their routers is crucial for drawing credible conclusions on topics such as network resilience and traffic censorship. IP geolocation databases (DBs) are frequently used for this purpose. However, the accuracy of DBs on top AS routers has not been fully evaluated due to the lack of a comprehensive ground truth dataset. In this study, we address this gap by constructing a geolocation ground truth dataset that contains more than 12,000 router interfaces in the top 10 ASes, utilizing delay measurements from carefully selected Looking Glass vantage points. We evaluate the coverage and accuracy of 6 DBs, including 4 widely used ones and 2 new ones. Our evaluation shows that most DBs exhibit poor accuracy when geolocating top AS routers. We conduct an in-depth analysis to uncover the primary techniques behind DBs. Our investigation reveals the reasons behind the poor performance of certain DBs. Moreover, we discover that the best-performing DB heavily relies on hostnames, and all DBs perform poorly in geolocating routers without hostnames. Our dataset, which is the largest ground truth dataset of top AS routers to the best of our knowledge, will be publicly accessible to the Internet research community.
Hui Wang 0011, Jilong Wang 0001, Peiran Wang
GLOBECOM2
2023 Optimal Collaborative Uploading in Crowdsensing with Graph Learning
abstract
It is pivotal and challenging for crowdsensing systems to guarantee the reliable uploading of sensory data from source devices (workers) to a centralized platform, in order to process sensing tasks accurately and fast. On one hand, with limited communication resources, uploading a massive amount of sensory data is not cost-effective. On the other hand, the disruption of uploading is inevitable because of stochastic network environments and worker dropout, resulting in extra wasting of resources. To address that, we focus on a collaborative uploading scenario and propose to reduce the uploading latency of sensory data by adaptive data allocation while retaining data integrity at the destination. A key technical challenge is to identify proper collaborative paths such that corresponding data allocation and uploading are reliable enough. As such, we formulate a joint optimization problem with the minimization goal of uploading latency by considering both path selection and data allocation. To mine helpful information from unstructured topology-aware data, we propose a new diffusion graph convolution module by forming information aggregation based on the diffusion process that characterizes the stochastic correlation of devices. After transforming the original problem into a primal-dual problem, an algorithm is then developed by adapting Advantage Actor-Critic (A2C) framework embedded with the diffusion graph convolution module. With extensive experiments, it is validated that the newly developed algorithm improves collaborative uploading by reducing uploading latency and also stabilizing the queue state of intermediate devices, compared to existing heuristic and learning-based methods.
Yao Zhang 0005, Tom H. Luan, Hui Wang 0011, Liang Wang 0017, Zhiwen Yu 0001, Bin Guo 0001
ICC3
2023 Janus: A Unified Distributed Training Framework for Sparse Mixture-of-Experts Models
abstract
Scaling models to large sizes to improve performance has led a trend in deep learning, and sparsely activated Mixture-of-Expert (MoE) is a promising architecture to scale models. However, training MoE models in existing systems is expensive, mainly due to the All-to-All communication between layers.
Juncai Liu, Hui Wang 0011
SIGCOMM2
2023 HMPT: a human-machine cooperative program translation method
abstract
Abstract Program translation aims to translate one kind of programming language to another, e.g., from Python to Java. Due to the inefficiency of translation rules construction with pure human effort (software engineer) and the low quality of machine translation results with pure machine effort, it is suggested to implement program translation in a human–machine cooperative way. However, existing human–machine program translation methods fail to utilize the human’s ability effectively, which require human to post-edit the results (i.e., statically modified directly on the model generated code). To solve this problem, we propose HMPT (Human-Machine Program Translation), a novel method that achieves program translation based on human–machine cooperation. It can (1) reduce the human effort by introducing a prefix-based interactive protocol that feeds the human’s edit into the model as the prefix and regenerates better output code, and (2) reduce the interactive response time resulted by excessive program length in the regeneration process from two aspects: avoiding duplicate prefix generation with cache attention information, as well as reducing invalid suffix generation by splicing the suffix of the results. The experiments are conducted on two real datasets. Results show compared to the baselines, our method reduces the human effort up to 73.5% at the token level and reduces the response time up to 76.1%.
Xin Zhang 0157, Zhiwen Yu 0001, Jiaqi Liu 0002, Hui Wang 0011, Liang Wang 0017, Bin Guo 0001
Autom. Softw. Eng.4
2023 Coverage Optimization for Directional Sensor Networks: A Novel Sensor Redeployment Scheme
abstract
The ever-growing Internet of Things (IoT) provides a powerful means for complex and changeable environmental monitoring. Directional sensor networks (DSNs), as a typical architecture of IoT, can efficiently facilitate various digital and intelligent IoT applications. In the DSNs, due to the asymmetry in coverage focus and diversity in detection angle of the directional IoT sensors, how to enhance the coverage performance with the limited sensors becomes a new challenge. To this end, we develop a novel sensor redeployment scheme based on the minimum exposure path (MEP) to optimize the coverage performance of the DSNs. Specifically, we first propose a minimum exposure path searching algorithm based on the particle swarm optimization (MEP-PSO) algorithm with the target of obtaining the MEP in the DSNs. With this algorithm, the traditional MEP problem can be analyzed and simplified by conducting the grid discretization and building the weighted undirected graph. Then, an MEP-based coverage optimization (MEP-CO) algorithm is proposed to determine the optimal deployment locations and the dispatch sensors so that the IoT sensors can be dynamically redeployed to achieve the coverage optimization. After that, we derive the formula for the coverage upper bound (CUB) and develop a CUB algorithm to provide a benchmark for evaluating the effectiveness of different coverage optimization algorithms. Simulation results demonstrate that the proposed coverage optimization scheme can significantly promote the minimum exposure value (MEV) and coverage ratio of the monitoring area compared with the existing algorithms.
Xuelian Cai, Luqiao Wang, Yilong Hui, Wenwei Yue, Hui Wang 0011, Yao Zhang 0005, Nan Cheng 0001, Changle Li
IEEE Internet Things J.6
2023 Computation Offloading Method Using Stochastic Games for Software-Defined-Network-Based Multiagent Mobile Edge Computing
abstract
In the scenario of Industry 4.0, mobile smart devices (SDs) on production lines have to process massive amounts of data. These computing tasks sometimes far exceed the computing capability of SDs and require lots of energy and time to process. How to effectively reduce energy consumption and latency is necessary to be solved. To this end, we first propose a software-defined network (SDN)-based mobile edge computing (MEC) system. In the MEC system, SDs can offload computation tasks to edge servers to decrease the processing latency and avoid the waste of energy. At the same time, taking advantage of SDN’s programmability, scalability, and isolation of the control plane and the data plane, an SDN controller can manage edge devices within the MEC system. Second, based on a stochastic game, we study the computation offloading and resource allocation problems in the MEC system and establish a stochastic game-based computation offloading model. Furthermore, we prove that the multiuser stochastic game in this system can achieve Nash Equilibrium. We further consider each SD as an independent agent and design a stochastic game-based resource allocation algorithm with prioritized experience replays (SGRA-PERs) to minimize energy consumption and processing latency with Multiagent Reinforcement Learning. Experiment results demonstrate that the proposed SGRA-PER is superior to MADDPG,$Q$-Mix, and MAPPO algorithms, which can significantly reduce the processing delay and energy consumption with dynamic resource allocation. Moreover, SGRA-PER can still keep a higher performance under the increase of SDs, which can be applied in a large-scale MEC system.
Guowen Wu, Hui Wang 0011, Hong Zhang 0046, Shui Yu 0001, Shigen Shen
IEEE Internet Things J.2
2023 Communication compression techniques in distributed deep learning: A survey
abstract
Nowadays, the training data and neural network models are getting increasingly large. The training time of deep learning will become unbearably long on a single machine. To reduce the computation and storage burdens, distributed deep learning has been put forward to collaboratively train a large neural network model with multiple computing nodes in parallel. The unbalanced development of computation and communication capabilities has led to training time being dominated by communication time, making the communication overhead a major challenge toward efficient distributed deep learning. Communication compression is an effective method to alleviate communication overhead, and it has evolved from simple random sparsification or quantization to versatile strategies or data structures. In this survey, existing communication compression techniques are reviewed and classified to provide a bird’s eye view. The main properties of each class of compression methods are analyzed, and their applications or theoretical convergence are described if necessary. This survey is potentially helpful for researchers and engineers to understand the up-to-date achievements on the communication compression techniques that accelerate the training of large deep learning models.
Zeqin Wang, Yuedong Xu 0001, Yipeng Zhou, Hui Wang 0011
J. Syst. Archit.5
2023 Optimizing the Numbers of Queries and Replies in Convex Federated Learning With Differential Privacy
abstract
Federated learning (FL) empowers distributed clients to collaboratively train a shared machine learning model through exchanging parameter information. Despite the fact that FL can protect clients’ raw data, malicious users can still crack original data with disclosed parameters. To amend this flaw, differential privacy (DP) is incorporated into FL clients to disturb original parameters, which however can significantly impair the accuracy of the trained model. In this work, we study an imperative question which has been vastly overlooked by existing works: what are the optimal numbers of queries and replies in FL with DP so that the final model accuracy is maximized. In FL, the parameter server (PS) needs to query participating clients for multiple global iterations to complete training. Each client responds a query from the PS by conducting a local iteration. We consider FL that will uniformly and randomly select participating clients to conduct local iterations with the FedSGD algorithm. Our work investigates how many times the PS should query clients and how many times each client should reply the PS by incorporating two most extensively used DP mechanisms (i.e., the Laplace mechanism and Gaussian mechanisms). Through conducting convergence rate analysis, we can determine the optimal numbers of queries and replies in FL with DP so that the final model accuracy can be maximized. Finally, extensive experiments are conducted with publicly available datasets: MNIST and FEMNIST, to verify our analysis and the results demonstrate that properly setting the numbers of queries and replies can significantly improve the final model accuracy in FL with DP.
Yipeng Zhou, Xuezheng Liu, Di Wu 0001, Hui Wang 0011, Shui Yu 0001
IEEE Trans. Dependable Secur. Comput.5
2023 Semi-Decentralized Federated Edge Learning With Data and Device Heterogeneity
abstract
Federated edge learning (FEEL) emerges as a privacy-preserving paradigm to effectively train deep learning models from the distributed data in 6G networks. Nevertheless, the limited coverage of a single edge server results in an insufficient number of participating client nodes, which may impair the learning performance. In this paper, we investigate a novel FEEL framework, namelysemi-decentralized federated edge learning(SD-FEEL), where multiple edge servers collectively coordinate a large number of client nodes. By exploiting the low-latency communication among edge servers for efficient model sharing, SD-FEEL incorporates more training data, while enjoying lower latency compared with conventional federated learning. We detail the training algorithm for SD-FEEL with three steps, including local model update, intra-cluster, and inter-cluster model aggregations. The convergence of this algorithm is proved on non-independent and identically distributed data, which reveals the effects of key parameters and provides design guidelines. Meanwhile, the heterogeneity of edge devices may cause the straggler effect and deteriorate the convergence speed of SD-FEEL. To resolve this issue, we propose an asynchronous training algorithm with a staleness-aware aggregation scheme, of which, the convergence is also analyzed. The simulations demonstrate the effectiveness and efficiency of the proposed algorithms for SD-FEEL and corroborate our analysis.
Yuchang Sun 0001, Jiawei Shao, Yuyi Mao, Hui Wang 0011, Jun Zhang 0004
IEEE Trans. Netw. Serv. Manag.4
2023 Optimizing Parameter Mixing Under Constrained Communications in Parallel Federated Learning
abstract
In vanilla Federated Learning (FL) systems, a centralized parameter server (PS) is responsible for collecting, aggregating and distributing model parameters with decentralized clients. However, the communication link of a single PS can be easily overloaded by concurrent communications with a massive number of clients. To overcome this drawback, multiple PSes can be deployed to form a parallel FL (PFL) system, in which each PS only communicates with a subset of clients and its neighbor PSes. On one hand, each PS conducts iterations with clients in its subset. On the other hand, PSes communicate with each other periodically to mix their parameters so that they can finally reach a consensus. In this paper, we propose a novel parallel federated learning algorithm called Fed-PMA, which optimizes such parallel FL under constrained communications by conducting parallel parameter mixing and averaging with theoretic guarantees. We formally analyze the convergence rate of Fed-PMA with convex loss, and further derive the optimal number of times each PS should mix with its neighbor PSes so as to maximize the final model accuracy within a fixed span of training time. Theoretical study manifests that PSes should mix their parameters more frequently if the connection between PSes is sparse or the time cost of mixing is low. Inspired by our analysis, we propose the Fed-APMA algorithm that can adaptively determine the near-optimal number of mixing times with non-convex loss under dynamic communication conditions. Extensive experiments with realistic datasets are carried out to demonstrate that both Fed-PMA and its adaptive version Fed-APMA significantly outperform the state-of-the-art baselines.
Xuezheng Liu, Zirui Yan, Yipeng Zhou, Di Wu 0001, Xu Chen 0004, Hui Wang 0011
IEEE/ACM Trans. Netw.6
2023 Offloading Elastic Transfers to Opportunistic Vehicular Networks Based on Imperfect Trajectory Prediction
abstract
Due to the high cost of cellular networks, vehicle users would like to offload elastic traffic through vehicular networks as much as possible. This demand prompts researchers to consider how to make the vehicular network system achieve better performance for requests coming online, such as maximizing throughput. The traffic in vehicular networks is transferred through opportunistic contacts between vehicles and infrastructures. When making scheduling decisions, the scheduler must be aware of vehicles’ future trajectories. Vehicles’ future trajectories are usually predicted by trajectory prediction algorithms when users are unwilling to report their future trips. Unfortunately, no trajectory prediction algorithm can be completely accurate, and these inaccurate prediction results will degrade the throughput achieved by scheduling algorithms. In this paper, we focus on reducing the negative impact of inaccurate predictions. Specifically, we measure two data-driven trajectory prediction algorithms that have been widely used for trajectory predictions and understand the characteristics of the accuracy of predicted contacts. Based on the enlightenment from the measurement, we design a system, i.e., i-Offload, to offload elastic traffic under imperfect trajectory predictions. The experimental results show that our system has good throughput and high scheduling efficiency even under imperfect trajectory predictions. Compared with existing scheduling algorithms, our method improves the throughput by about one time.
Chao Xu 0015, Hui Wang 0011, Jilong Wang 0001, Yipeng Zhou, Yuedong Xu 0001, Di Wu 0001, Changqing An
IEEE/ACM Trans. Netw.2
2022 Distributed Routing Controller for Large-scale Live Video Streams in Real-Time Networks
abstract
Overlay routing control for live video delivery has become an important yet challenging task for Real-Time Networks (RTNs), but existing approaches designed for traditional Content Delivery Networks (CDNs) fall short of meeting the challenge. In this paper, we develop a distributed overlay routing controller for RTNs to deliver massive high-quality live video streams in low latencies. We first formulate a joint optimization that offers rich control flexibility and can yield low-latency and cost-effective routing solutions. To obtain the appealing potential value of the optimal solution in the context of large-scale live videos, we develop a distributed control framework that can find near-optimal solutions at fine-grained timescales. Evaluations on real-world live video traces show that our distributed controller derives high-quality (in terms of both performance and cost) overlay routing solutions while reducing the decision latency by 38%-89% compared to the state-of-the-art centralized controller.
Hui Wang 0011, Chao Xu 0015, Jilong Wang 0001
GLOBECOM2
2022 PipeCompress: Accelerating Pipelined Communication for Distributed Deep Learning
abstract
Distributed learning is widely used to accelerate the training of deep learning models, but it is known that communication efficiency limits the scalability of distributed learning systems. Current gradient compression techniques provide promising methods to reduce communication time, but the extra time incurred by compression is not negligible. After compression techniques are applied, the communication time is significantly reduced because the data size needed to communicate becomes much smaller, but compressing gradients is time-consuming and it becomes a new bottleneck. In this paper, we design and implement PipeCompress, a system to decouple compression and backpropagation operations into two processes and pipeline the two processes to hide compression time. We also propose a specialized inter-process communication mechanism based on the characteristics of DNN distributed training to improve the efficiency of passing messages between the two processes, which makes sure that the decoupling does not bring much extra inter-process communication time cost. As far as we know, this is the first work that notices the overhead of compression and pipelines backpropagation and compression operations to hide compression time in distributed learning. Experiments show that PipeCompress can significantly hide compression time, reduce iteration time, and accelerate the training process on various DNN models.
Juncai Liu, Hui Wang 0011, Chenghao Rong, Jilong Wang 0001
ICC2
2022 CoupHM: Task Scheduling Using Gradient Based Optimization for Human-Machine Computing Systems
abstract
We witnessed great advancement in Artificial Intelligence (AI) powered technologies in recent years, and yet, when applied to certain high-stake contexts, such as medical diagnosis, automatic driving and criminal justice, they are not qualified. This matter can be greatly settled by Human-Machine Computing (HMC), which is an effective computing paradigm that couples the expertise and demonstration abilities of humans with the high-performance computing power of machines. This work studies an optimal task scheduling problem for HMC systems, where various tasks are decomposed and dispatched to humans and AI-enabled machines to provide significantly better benefits compared to either type of computing resources in isolation. However, designing such optimal task scheduling is challenging because of the stochastic hybrid features of machines, as well as various human professional abilities. Considering the Quality of Service (QoS) and the heterogeneity of human-machine computing resources, we propose CoupHM, a feasible task scheduler using gradient based optimization for HMC systems. In particular, we firstly present the underlying architecture of HMC system and details of the task-driven workload model. On that basis, we then formulate the objective optimization problem to be solved and describe the composition of the CoupHM scheduler. Finally, the performance of our solution is evaluated by the simulation experiments, and the results indicate that the proposed scheduler has preferable performance both in balancing resources and guaranteeing QoS, which can serve as guidelines for future research on HMC systems.
Hui Wang 0011, Zhuoli Ren, Zhiwen Yu 0001, Yao Zhang 0005, Jiaqi Liu 0002, Helei Cui
ICPADS1
2022 An Efficient HPU Resource Virtualization Framework for Human-Machine Computing Systems
abstract
Driven by state-of-the-art AI technologies, human-AI collaboration has become an important area in computer supported teamwork research. Principles for Human-Machine Computing (HMC) have been discussed to accomplish complex goals by outsourcing some computational steps to humans and collaboratively achieving more accurate results. In HMC systems, however, the human participant brings great challenges to efficient provisioning of resources. Virtualization provides ideas for increasing the agility, flexibility and scalability of resources and has been applied in traditional computer systems, such as cloud computing. Unfortunately, existing hardware virtualization scheme is not ready to address utilization, and performance limitations associated with Human Processing Unit (HPU) resources. To tackle this problem, in this paper, we propose an efficient HPU resource virtualization framework for HMC systems. In particular, we firstly describe the modeling details of the HPU resource. And on this basis, we present the Time Division Multiplexing (TDM)-based virtualization scheme which aims to establish the mapping between each real HPU (rHPU) and its virtual HPUs (vHPUs). Secondly, we apply our minds to address the vHPU reconfiguration problem by managing the vHPU waiting queue, and propose DvR-PSO algorithm. Finally, the performance of our proposed HPU virtualization framework is evaluated through simulation experiments, and the results show that our solution can make remarkable effectiveness, which can serve as guidelines for future research on HMC systems.
Hui Wang 0011, Zhiwen Yu 0001, Zhuoli Ren, Yao Zhang 0005, Bin Guo 0001
Internetware1
2022 Multi-hop Precision Time Protocol: an Internet Applicable Time Synchronization Scheme
abstract
Precise time synchronization is essential for 5G systems, data centers, industrial automation systems, military fields, and more. Although the precision of IEEE 1588 PTP can achieve sub-hundred-nanosecond accuracy, it works only when being deployed hop-by-hop within a LAN with limited range. Hop-by-hop deployment leads to high deployment costs and makes it inapplicable over the Internet. In this paper, we propose the multi-hop precision time protocol (M-PTP), a high-precision and low-cost time synchronization protocol, which does not require hop-by-hop deployment, and no special functions need to be added to the network relay devices such as routers and switches. M-PTP leverages two key ideas. First, to mitigate the "asymmetry in forward delay and reverse delay" problem, SVM-based delay estimation is used to calculate the distribution of positive and negative random delays, then L-estimator is leveraged to estimate the time offset. Second, based on time offset, M-PTP exploits loop effect optimization among nodes. We implemented the protocol and tested its performance on variance hops under different traffic conditions and CPU loads. The experimental results show that M-PTP can achieve a precision of 11.61ns at 5 hops, which is approximately 3 times the precision of HUYGENS and approximately 30 times the precision of PTP.
Kunling He, Changqing An, Hui Wang 0011, Tianshu Li, Linmei Zu
NOMS3
2022 Predicting Unseen Links Using Learning-based Matrix Completion
abstract
Researchers have noticed the AS-level Internet topology that can be observed from the current measurement infrastructure is far from complete, which means researchers have to deploy more measurement vantage points (VPs) and conduct measurements for more source/destination pairs to fully understand the whole Internet. Unfortunately, it is known that blindly deploying more points and conducting more measurements to achieve the goal is inefficient, if not infeasible. In this paper, we try to improve the efficiency by predicting where unseen AS links might be located from the observed AS paths to guide the measurements towards a more complete AS-level topology. We formulate the prediction of unseen links as a matrix completion problem. However, the traditional matrix completion methods have limited learning capacities and cannot deal with the complex constraints on the underlying topology. We develop a learning-based matrix completion method specifically for the unseen AS link prediction problem. The method exploits a neural network and utilizes side-information which is carefully chosen from AS attributes based on our understanding on Internet peering practices, therefore our method is able to learn more expressive latent vectors and achieves outstanding prediction performance in our scenario. Experiments performed on a real-world dataset show the prediction results can achieve a high AUC (Area Under the Receiver Operating Characteristic Curve) of 0.834.
Shuying Zhuang, Hui Wang 0011, Jilong Wang 0001, Changqing An, Yuedong Xu 0001, Tianhao Wu 0010
NOMS2
2022 RouteInfer: Inferring Interdomain Paths by Capturing ISP Routing Behavior Diversity and Generality
Tianhao Wu 0010, Hui Wang 0011, Jilong Wang 0001, Shuying Zhuang
PAM2
2022 HM-MDS: A Human-machine Collaboration based Online Medical Diagnosis System
abstract
Online medical diagnosis refers to diagnosing diseases and providing treatment suggestions on the websites. It develops rapidly and has become a new choice for patients to seek medical treatment. Although manual online medical diagnosis is reliable, it has problems such as low efficiency, heavy burden on doctors, and long waiting time for patients. Relying on machines for automatic disease diagnosis is highly efficient, which, however, has low accuracy and reliability. In general, online medical diagnosis usually has two stages: inquiry and diagnosis. Inquiry stage refers to asking about the patient’s physiological, where the questions are usually streamlined, and thus can be handled by the machine. Diagnosis stage is to diagnose the disease and provide medical recommendations, which has strict requirements for accuracy and safety, and thus should be handled by the human. Inspired by this, in the paper we propose a human-machine collaboration based online medical diagnosis system, i.e., HM-MDS. In inquiry stage, the system employs the machine. It uses the BERT+CRF to identify symptoms in the patient’s dialogue and uses a DQN-based method to ask about symptoms. In diagnosis stage, the system employs both the machine and the human. The machine generates a pre-diagnosis result by calculating disease probability. Then the human doctor gives the final diagnosis result by checking the pre-diagnosis result and revising it if necessary. Obviously, HM-MDS can effectively save human doctor’s time as well as patient’s time, while ensure the accuracy of the diagnosis result. We conduct experiments on a real-world dataset. The results show our approach improves the online medical diagnosis’s reliability as well as patient satisfaction, and ensures diagnosis accuracy. The time cost for a reliable medical diagnosis is reduced to 36% compared with pure manual work.
Yixuan Chen 0011, Jiaqi Liu 0002, Zhiwen Yu 0001, Hui Wang 0011, Liang Wang 0017, Bin Guo 0001
SMC4
2022 Exploring the Layered Structure of Containers for Design of Video Analytics Application Migration
abstract
The existing solutions to the migration of container-based applications are not suitable for live video analytics applications because these solutions can result in excessive migration time. Intuitively, it is possible to exploit the layered structure of containers to improve the migration performance, but we still need a good understanding of the characteristics of the containers related to video analytics applications to justify the intuitive idea and design a high-performance solution. In this paper, we pull 3735 representative images from Docker Hub. We analyze these images and get three main findings: (1) the images of video analytics applications have more layers and larger sizes than that of general applications; (2) we can cache images in the destination servers and reuse the same layers cached in the destination server to reduce the migration time when an image is migrated from its source server to its destination server; (3) the size of the remaining data to be transferred during migrations is still large and a high-performance migration scheme is still necessary. Based on the above findings, we propose a pipelined migration scheme to optimize migration performance. Evaluations show that pipelined migration performs significantly better than other migration schemes.
Chenghao Rong, Hui Wang 0011, Juncai Liu, Jilong Wang 0001
WCNC2
2022 Semi-Decentralized Federated Edge Learning for Fast Convergence on Non-IID Data
abstract
Federated edge learning (FEEL) has emerged as an effective approach to reduce the large communication latency in Cloud-based machine learning solutions, while preserving data privacy. Unfortunately, the learning performance of FEEL may be compromised due to limited training data in a single edge cluster. In this paper, we investigate a novel framework of FEEL, namely semi-decentralized federated edge learning (SD-FEEL). By allowing model aggregation across different edge clusters, SD-FEEL enjoys the benefit of FEEL in reducing the training latency, while improving the learning performance by accessing richer training data from multiple edge clusters. A training algorithm for SD-FEEL with three main procedures in each round is presented, including local model updates, intra-cluster and inter-cluster model aggregations, which is proved to converge on non-independent and identically distributed (non-IID) data. We also characterize the interplay between the network topology of the edge servers and the communication overhead of inter-cluster model aggregation on the training performance. Experiment results corroborate our analysis and demonstrate the effectiveness of SD-FFEL in achieving faster convergence than traditional federated learning architectures. Besides, guidelines on choosing critical hyper-parameters of the training algorithm are also provided.
Yuchang Sun 0001, Jiawei Shao, Yuyi Mao, Hui Wang 0011, Jun Zhang 0004
WCNC4
2022 Faster Activity and Data Detection in Massive Random Access: A Multiarmed Bandit Approach
abstract
This article investigates the grant-free random access mechanism for massive Internet of Things (IoT) devices. By embedding the data symbols in the signature sequences, joint device activity detection and data decoding can be achieved, which, however, significantly increases the computational complexity. Coordinate descent algorithms that enjoy a low per-iteration complexity have been employed to solve this detection problem, but previous works typically employ a random coordinate selection policy which leads to slow convergence. In this article, we develop multiarmed bandit (MAB) approaches for more efficient detection via coordinate descent, which achieves a delicate tradeoff betweenexplorationandexploitationin coordinate selection. Specifically, we first propose a bandit-based strategy, i.e., Bernoulli sampling, to speed up the convergence rate of coordinate descent, by learning which coordinates will result in more aggressive descent of thenonconvex objective function. To further improve the convergence rate, an inner MAB problem is established to learn the exploration policy of Bernoulli sampling. Both convergence rate analysis and simulation results are provided to show that the proposed bandit-based algorithms enjoy faster convergence rates with a lower time complexity compared with the state-of-the-art algorithm. Furthermore, our proposed algorithms are generally applicable to different scenarios, e.g., massive random access with low-precision analog-to-digital converters (ADCs).
Jialin Dong, Jun Zhang 0004, Yuanming Shi, Hui Wang 0011
IEEE Internet Things J.4
2022 Gain Without Pain: Offsetting DP-Injected Noises Stealthily in Cross-Device Federated Learning
abstract
Federated learning (FL) is an emerging paradigm through which decentralized devices can collaboratively train a common model. However, a serious concern is the leakage of privacy from exchanged gradient information between clients and the parameter server (PS) in FL. To protect gradient information, clients can adopt differential privacy (DP) to add additional noises and distort original gradients before they are uploaded to the PS. Nevertheless, the model accuracy will be significantly impaired by DP noises, making DP impracticable in real systems. In this work, we propose a novel noise information secretly sharing (NISS) algorithm to alleviate the disturbance of DP noises by sharing negated noises among clients. We theoretically prove that: 1) if clients are trustworthy, DP noises can be perfectly offset on the PS and 2) clients can easily distort negated DP noises to protect themselves in case that other clients are not totally trustworthy, though the cost lowers model accuracy. NISS is particularly applicable for FL across multiple Internet of Things (IoT) systems, in which all IoT devices need to collaboratively train a model. To verify the effectiveness and the superiority of the NISS algorithm, we conduct experiments with the MNIST and CIFAR-10 data sets. The experimental results verify our analysis and demonstrate that NISS can improve model accuracy by 19% on average and obtain better privacy protection if clients are trustworthy.
Wenzhuo Yang, Yipeng Zhou, Miao Hu 0001, Di Wu 0001, James Xi Zheng, Hui Wang 0011, Song Guo 0001, Chao Li 0067
IEEE Internet Things J.6
2022 USST: A two-phase privacy-preserving framework for personalized recommendation with semi-distributed training
Yipeng Zhou, Jun Liu 0001, Hui Wang 0011, Jilong Wang 0001, Guanfeng Liu 0001, Di Wu 0001, Chao Li 0067, Shui Yu 0001
Inf. Sci.3
2022 Towards Hit-Interruption Tradeoff in Vehicular Edge Caching: Algorithm and Analysis
abstract
Recent advancements in edge computing and edge caching provide a feasible solution to support a plethora of new applications such as on-demand videos, AR/VR, road surveillance. However, to apply edge caching in vehicular scenarios is still difficult due to the unkonwn request pattern of vehicular users and intermittent service links between vehicles and edge servers (e.g., Road Side Units, RSUs). In this paper, we aim to investigate the vehicular edge caching problem in practical vehicular scenarios by considering higher hit ratio, while avoiding interruption of caching services. Specifically, to obtain a higher hit ratio, we firstly propose an on-demand adaptive cache algorithm. The algorithm can adjust the eviction time of cached contents by tracking the dynamics of requests and content popularity. We then develop an analysis framework to model the interruption performance of caching services from RSUs. Through diffraction approximation theory, the service process can be modeled as a joint process of the movement and stopping of vehicles to deduce the interruption ratio. To apply the on-demand adaptive cache algorithm in practical scenarios, the final caching decisions should be corrected by incorporating the interruption performance. Therefore, a$\alpha $-fair utility-oriented vehicular edge caching scheme is developed, which can achieve the tradeoff of hit ratio and interruption ratio. Performance evaluation shows the advantages of our proposed vehicular caching scheme in hit ratio, accuracy of analysis model, utility, respectively.
Yao Zhang 0005, Changle Li, Tom H. Luan, Chau Yuen, Yuchuan Fu, Hui Wang 0011, Weigang Wu
IEEE Trans. Intell. Transp. Syst.6
2022 Scheduling Massive Camera Streams to Optimize Large-Scale Live Video Analytics
abstract
In smart cities, more and more government departments will make use of live analytics of videos from surveillance cameras in their tasks, such as vehicle traffic monitoring and criminal detection. Obviously, it is costly for each individual department to deploy its own infrastructure,i.e., cameras and analytics system. In this paper, we consider a scenario in which a city deploys an infrastructure and departments submit requests to access and analyze videos for their own purposes. The live analytics of massive streams is computation-intensive and the tasks might be latency-critical, which makes scheduling massive streams to optimize all tasks an essential and challenging work. We exploit an end-edge-cloud architecture and propose an adaptive system to schedule the massive camera streams and tasks, which considers all factors affecting the computation and networking resource consumption,e.g., sharing of model computation, video quality, model partition, and task placement. Particularly, the resource consumption ofFaster R-CNN + ResNet101under each partition scheme is profiled for the first time and we notice the partition must be used together with lossless compression techniques to be beneficial. Furthermore, sometimes tasks might be required to migrate because the scheduling decision made by the system changes to adapt to the changing resource supply and demand. In order to avoid the performance degradation during migration, we propose a non-destructive migration scheme and implement it in the system. Simulations demonstrate our system achieves a total utility close to the maximum and our analytics system performs better than state-of-the-art solutions.
Chenghao Rong, Hui Wang 0011, Juncai Liu, Jilong Wang 0001, Sharon X. Huang
IEEE/ACM Trans. Netw.2
2021 Discovering obscure looking glass sites on the web to facilitate internet measurement research
abstract
Despite researchers have noticed that Looking Glass (LG) vantage points (VPs) are valuable for Internet measurement researches, they can only exploit VPs from well-known LG sites published on several LG portal pages. There should be a lot of LG sites that are not published in these portal pages, namely obscure LG sites, which are not easy to be found and exploited by researchers. In this paper, we design an efficient focused crawler to discover as many LG sites as possible which can avoid unnecessary resource consumption on analyzing irrelevant pages. Our designed focused crawler takes a similarity-guided search that exploits the well-developed search engines and comprehensively mines the common features shared by known LG sites to discover more LG pages. Moreover, the focused crawler takes a two-step PU learning classifier based on carefully selected LG features to efficiently discard irrelevant URLs, thus avoiding a lot of unnecessary resource consumption. As far as we know, we are the first to develop a method to discover obscure LG sites on the web. Experimental results show the effectiveness of our focused crawler. To facilitate practical applications, we further develop an automation tool, which can successfully retrieve 910 obscure automatable LG VPs from relevant pages obtained through our focused crawler. The 910 LG VPs significantly increase the geographic and network coverage of available VPs and we show their potential values in improving the completeness of AS-level Internet topology by a simple case study. Our method and the final VP list are beneficial to the measurement community.
Shuying Zhuang, Hui Wang 0011, Jilong Wang 0001, Zujiang Pan, Tianhao Wu 0010
CoNEXT2
2021 FedPA: An adaptively partial model aggregation strategy in Federated Learning
Juncai Liu, Hui Wang 0011, Chenghao Rong, Yuedong Xu 0001, Jilong Wang 0001
Comput. Networks2
2021 Collaboratively Replicating Encoded Content on RSUs to Enhance Video Services for Vehicles
abstract
With the development of smart cities, Internet services will be pervasively accessible for moving vehicles. It is envisioned that the video content demand of vehicles will explode in the near future. However, the strategy to efficiently distribute video content in large-scale vehicular networks is still absent due to challenges arising from the huge video population, heavy bandwidth consumption, heterogeneous user devices, and vehicles’ mobility. In this work, we propose to collaboratively replicate video content on Roadside Units (RSUs) to enhance video distribution services based on the fact that the contact period between moving vehicles and a single RSU is not long enough to complete video downloading. In our design, a video file is split into multiple chunks. Each RSU replicates a small number of original chunks and chunks encoded by network coding. Replicating encoded chunks can reduce redundancy of chunks on different RSUs so that RSUs can complement each other better, whereas original chunks can be transrated to chunks with lower bitrates flexibly to fit in users’ devices. Therefore, we replicate both original and encoded chunks on RSUs to take advantages of both sides. Stochastic models are employed to analyze chunk download processes and a convex optimization problem is formulated to determine the optimal partition of space allocated to each kind of chunks. Furthermore, we extend our strategy to support video streaming services and empirically prove that the influence caused by limitations of network coding is moderate. In the end, we conduct extensive simulations which not only validate the accuracy of our models but also demonstrate that our strategy can effectively boost video distribution services.
Yipeng Zhou, Guoqiao Ye, Di Wu 0001, Hui Wang 0011, Min Chen 0003
IEEE Trans. Mob. Comput.5
2020 MEP-PSO Algorithm-Based Coverage Optimization in Directional Sensor Networks
abstract
As a sub-class of internet of things (IoTs), wireless sensor networks (WSNs) are becoming ubiquitous in recent years, which makes the efficient coverage of sensors challenging. Traditionally, WSNs are composed of omni-directional sensors, which, however, are still limited to unadjustable sensing angle and superfluous energy consumption. Fortunately, these limitations can be overcome by deploying directional sensors in WSNs, thus forming directional sensor networks, namely DSNs. Therefore, it is necessary to propose efficient coverage optimization methods for DSNs to solve the minimum exposure path (MEP) problem that refers to a path along which the intruder can go through WSNs with lowest detection probability. In this paper, a novel MEP-PSO algorithm-based coverage optimization mechanism is proposed to improve the coverage quality in DSNs. With our coverage optimization mechanism, the traditional MEP problem is analyzed by means of discrete geometric theories while the path searching performance is improved based on the particle swarm optimization (PSO) algorithm. Specifically, the deployment scenario is firstly discretized into multiple square grids with uniform sizes. The weighted undirected graph is thus constructed in which the path segment exposure of MEP can be analyzed by discrete geometric theory. Based on the analysis, the feasibility of PSO is evaluated and enhanced in terms of MEP searching. Using our algorithm, the coverage performance of DSNs can be improved significantly by dynamically adjusting the positions of directional sensors. Finally, we conduct extensive experiments to validate the effectiveness of our work.
Luqiao Wang, Changle Li, Hui Wang 0011, Yao Zhang 0005
GLOBECOM3
2020 A Scheme on Pedestrian Detection using Multi-Sensor Data Fusion for Smart Roads
abstract
Transforming our roads into smart roads is an indispensable step towards future self-driving systems, and therefore has drawn increasing attention from both academia and industry. To this end, this paper develops a novel cost-effective IoT-based target detection system utilizing the multi-sensor data fusion technology with a particular focus on pedestrian detection, as an important component of smart road system. Particularly, the developed intelligent pedestrian detection module (${i}$PDM) consists of three major sensors, i.e., Doppler microwave radar sensor, passive infrared (PIR), and geomagnetic sensor. A multi-sensor data fusion algorithm is developed to fuse the sensor data and achieves reliable target detection. After that, ${i}$PDM sends the relevant warning signal wirelessly to nearby base station and vehicles. Experiments are conducted on real traffic environment to evaluate the performance of ${i}$PDM. The results validate the high reliability of ${i}$PDM with an average 91.7% detection accuracy. Moreover, to our best knowledge, ${i}$PDM is the first IoT-based implementation for pedestrian detection of smart roads. It is necessary to highlight that ${i}$PDM is a low-cost, low-power, wide-coverage pedestrian detection system where the cost of a single ${i}$PDM is only US $ 30, which makes it suitable to large-scale deployment.
Hui Wang 0011, Changle Li, Yao Zhang 0005, Yilong Hui, Guoqiang Mao
VTC Spring1
2020 Understanding the latency to visit websites in China: An infrastructure perspective
Shuying Zhuang, Hui Wang 0011, Pei Zhang 0003, Jilong Wang 0001
Comput. Networks2
2020 Prediction Based Vehicular Caching: Where and What to Cache?
Yao Zhang 0005, Changle Li, Tom H. Luan, Yuchuan Fu, Hui Wang 0011
Mob. Networks Appl.5
2020 Exploiting user reviews for automatic movie tagging
Canrui Wu, Chen Wang 0008, Yipeng Zhou, Di Wu 0001, Min Chen 0003, Hui Wang 0011, Harry Qin
Multim. Tools Appl.6
2020 Squeezing the Gap: An Empirical Study on DHCP Performance in a Large-Scale Wireless Network
abstract
Dynamic Host Configuration Protocol (DHCP) is widely used to dynamically assign IP addresses to users. However, due to little knowledge on the behavior and performance of DHCP, it is challenging to configure lease time and divide IP addresses for address pools properly in large-scale wireless networks. In this paper, we conduct the largest known measurement on the behavior and performance of DHCP in the wireless network of T University (TWLAN). We find the performance of DHCP is far from satisfactory: (1) The non-authenticated devices lead to a waste of 25% of addresses at the rush hour. (2) Address pool utilization varies greatly under the current address division strategy. (3) A device does not generate traffic for 67% of the lease time on average. Meanwhile, we observe devices of different locations and operating systems show diverse online patterns. A unified lease time setting could result in an inefficient usage of addresses. To address the problems, taking account of authentication information and online patterns, we propose a new leasing strategy. The results show it outperforms three state-of-the-art baselines and reduces the number of assigned addresses by 24% and the average total lease time by 17% without significantly increasing the DHCP server load. Besides, we further propose an adaptive address division strategy to balance the address utilization of pools, which can be deployed in parallel with the new leasing strategy and reduce the risk of address exhaustion.
Haibo Wang 0004, Hui Wang 0011, Jilong Wang 0001, Weizhen Dang, Jing'an Xue, Jinzhe Shan
IEEE/ACM Trans. Netw.2
2019 WRT: Constructing Users' Web Request Trees from HTTP Header Logs
abstract
As web traffic has already dominated Internet, massive web logs are being generated ceaselessly. It is essential and meaningful for operators to mine valuable information and knowledge from the log data. However, the state-less feature of HTTP and increasing dynamics and complexities of web services bring a challenge to web mining in web logs. To solve the problem, in this paper, we introduce Web Request Tree (WRT) to reconstruct users' web request behaviors from web logs with HTTP header information, which can be applied in many fields. Compared to previous related work, our method pays special attentions to modern technologies to handle cases caused by these technologies such as PJAX. We evaluate the feasibility of our method with a measurement study on referrer policies of Alexa top websites and results show that our method can achieve a high accuracy for most websites. We also conduct experiments on a real-world dataset collected from official websites of a top university in China. We find that a lot of abnormal requests exist in log data by analyzing WRT and WRT is rich in information that is valuable for operators to analyze user behaviors, detect anomalies, optimize performance and so on.
Shengchao Liu, Jilong Wang 0001, Hui Wang 0011, Haibo Wang 0004
ICC3
2019 BDAC: A Behavior-aware Dynamic Adaptive Configuration on DHCP in Wireless LANs
abstract
DHCP is widely used to dynamically allocate IP addresses to the devices on local area networks, but the explosive increases of WiFi devices and their frequent mobility pose great challenges on DHCP performance in wireless LANs. In this paper, by analyzing large scale real network traces, we observe that the dynamic WiFi user behavior (e.g., online time pattern and spatio-temporal mobility pattern) leads to the poor DHCP performance. The IP pools in some VLANs have been exhausted in rush hours although the total IP utilization in WLAN is only 24%. Therefore, we have to configure IP lease times and IP pools dynamically and make sure that they are adaptive to the WiFi user behavior. In order to achieve this goal, we characterize and model the user behavior across online time pattern and spatiotemporal mobility pattern. Then we propose BDAC, a behaviour-aware dynamic adaptive configuration, which is combined of two strategies: adaptive IP lease time configuration and dynamic IP pool configuration. The former is to set adaptive lease times across user roles and area types based on online time pattern to reclaim IP addresses in time and reduce the peak IP usage, while the latter dynamically migrates the IP addresses across VLANs based on spatio-temporal mobility correlation to save the IP addresses. Using the real network traces of a different week, we conduct experiments to evaluate the performance of BDAC. Results show that BDAC can save up to 60% of IP addresses and the actual IP utilization rises from 24% to 59%. Furthermore, BDAC maintains high IP utilization when the number of VLANs in a WLAN increases.
Congcong Miao, Jilong Wang 0001, Tianying Ji, Hui Wang 0011, Chao Xu 0015, Fengyuan Ren
ICNP4
2019 Evaluating performance and inefficient routing of an anycast CDN
abstract
Anycast has been increasingly deployed for content delivery networks to map clients to their nearby replicas, which relies on the underlying routing. However, the simplicity of operation comes at cost of less precise client-mapping control. Although many works have measured anycast DNS, anycast CDNs, with different service goals and engineering, are still not fully understood. In this paper, we design novel methods and combine large-scale traceroute and HTTP measurement to evaluate the overall client-proximity and inefficient routing of the largest anycast CDN, Cloudflare. We find that 90% paths traverse only 2-4 ASes, which highlights its direct networks providers. By further identifying and characterizing direct providers at finer granularity of facilities, we quantitatively shows that Cloudflare unevenly uses few large transit providers to delivery the majority of contents. Inspired by the observations, we propose an anycast routing pathology and diagnosis methodology. Investigation reveals that few huge providers have outsized impact in that they are not only related to many inter-domain inflations, but also have path inflation inside their own networks, thus deserving priority focus when troubleshooting.
Jing'an Xue, Weizhen Dang, Haibo Wang 0004, Jilong Wang 0001, Hui Wang 0011
IWQoS5
2019 A survey on resource scheduling for data transfers in inter-datacenter WANs
Hui Wang 0011, Jilong Wang 0001, Changqing An, Qianli Zhang
Comput. Networks1
2019 Towards predictable performance via two-layer bandwidth allocation in cloud datacenter
Hui Yu 0004, Jiahai Yang 0001, Hui Wang 0011, Hui Zhang 0052
J. Parallel Distributed Comput.3
2019 An Adaptive Online Scheme for Scheduling and Resource Enforcement in Storm
abstract
As more and more applications need to analyze unbounded data streams in a real-time manner, data stream processing platforms, such as Storm, have drawn the attention of many researchers, especially the scheduling problem. However, there are still many challenges unnoticed or unsolved. In this paper, we propose and implement an adaptive online scheme to solve three important challenges of scheduling. First, how to make a scaling decision in a real-time manner to handle the fluctuant load without congestion? Second, how to minimize the number of affected workers during rescheduling while satisfying the resource demand of each instance? We also point out that the stateful instances should not be placed on the same worker with stateless instances. Third, currently, the application performance cannot be guaranteed because of resource contention even if the computation platform implements an optimal scheduling algorithm. In this paper, we realize resource isolation using Cgroup, and then the performance interference caused by resource contention is mitigated. We implement our scheduling scheme and plug it into Storm, and our experiments demonstrate in some respects our scheme achieves better performance than the state-of-the-art solutions.
Shengchao Liu, Jianping Weng, Hui Wang 0011, Changqing An, Yipeng Zhou, Jilong Wang 0001
IEEE/ACM Trans. Netw.3
2019 On The Robustness of Price-Anticipating Kelly Mechanism
abstract
The price-anticipating Kelly mechanism (PAKM) is one of the most extensively used strategies to allocate divisible resources for strategic users in communication networks and computing systems. The users are deemed as selfish and also benign, each of which maximizes his individual utility of the allocated resources minus his payment to the network operator. However, in many applications a user can use his payment to reduce the utilities of his opponents, thus playing a misbehaving role. It remains mysterious to what extent the misbehaving user can damage or influence the performance of benign users and the network operator. In this work, we formulate a non-cooperative game consisting of a finite amount of benign users and one misbehaving user. The maliciousness of this misbehaving user is captured by his willingness to pay to trade for unit degradation in the utilities of benign users. The network operator allocates resources to all the users via the price-anticipating Kelly mechanism. We present six important performance metrics with regard to the total utility and the total net utility of benign users, and the revenue of network operator under three different scenarios: with and without the misbehaving user, and the maximum. We quantify the robustness of PAKM against the misbehaving actions by deriving the upper and lower bounds of these metrics. With new approaches, all the theoretical bounds are applicable to an arbitrary population of benign users. Our study reveals two important insights: 1) the performance bounds are very sensitive to the misbehaving user's willingness to pay at certain ranges and 2) the network operator acquires more revenues in the presence of the misbehaving user which might disincentivize his countermeasures against the misbehaving actions.
Yuedong Xu 0001, Zhujun Xiao, Tianyu Ni, Hui Wang 0011, Xin Wang 0002, Eitan Altman
IEEE/ACM Trans. Netw.4
2018 A Multi-dimension Measurement Study of a Large Scale Campus WiFi Network
abstract
The growing trend of wireless devices and WiFi networks poses significant management challenges to network administrators. Characterizing WiFi user behavior and understanding WiFi network usage pattern are helpful to identify the management challenges so that network administrators could manage WiFi networks more efficiently. In this work, we collect comprehensive datasets, i.e., DHCP dataset, AAA dataset, SNMP dataset of ACs in a large campus WiFi network. We provide a detailed measurement study from multiple dimensions, i.e., server plane, temporal plane, spatial plane and traffic plane. We observe that the WiFi network under study is far from optimal. First, the phenomenon of IP waste is severe due to the isolation between DHCP server and AAA server. Second, current deployment of network infrastructure resources is based on network administrators' experience and it results in that the WiFi performance varies a lot across different areas. Furthermore, we also study the user behavior with different types of devices and in different kinds of buildings. Our observations indicate that the WiFi network could be improved and managed more efficiently from multiple dimensions. We believe that this measurement study is helpful for network administrators and researchers to understand more about large scale WiFi networks.
Congcong Miao, Jilong Wang 0001, Hui Wang 0011, Jun Zhang 0004, Shengchao Liu
LCN3
2018 A study on geographic properties of internet routing
Hui Wang 0011, Changqing An
Comput. Networks1
2018 Root Cause Analysis of Anomalies of Multitier Services in Public Clouds
Jianping Weng, Hui Wang 0011, Jiahai Yang 0001, Yang Yang 0004
IEEE/ACM Trans. Netw.2
2017 Root cause analysis of anomalies of multitier services in public clouds
abstract
Anomalies of multitier services running in cloud platform can be caused by components of the same tenant or performance interference from other tenants. If the performance of a multitier service degrades, we need to find out the root causes precisely to recover the service as soon as possible. In this paper, we argue that cloud providers are in a better position than tenants to solve this problem, and the solution should be non-intrusive to tenants' services or applications. Based on these two considerations, we propose a solution for cloud providers to help tenants to localize root causes of any anomaly. We design a non-intrusive method to capture the dependency relationships of components, which improves the feasibility of root cause localization system. Our solution can find out root causes no matter they are in the same tenant as the anomaly or from other tenants. Our proposed two-step localization algorithm exploits measurement data of both application layer and underlay infrastructure and a random walk procedure to improve its accuracy. Our realworld experiments of a three-tier web application running in a small-scale cloud platform show a 38.9% improvement in mean average precision compared to current methods.
Jianping Weng, Hui Wang 0011, Jiahai Yang 0001, Yang Yang 0004
IWQoS2
2016 Towards online anomaly detection by combining multiple detection methods and Storm
abstract
In this paper, we illustrate the significance and advantage of combining the results of multiple detection methods. We implement these methods as bolts in a Apache Storm cluster which is a famous real-time computation framework. We simulate two kinds of anomalies — one involving large number of small network flows and the other involving small number of large network flows. The experiments show that combining multiple methods outperforms any single detection method from the point of view of statistics. Besides, we observe that all the results are outputted in real time without delay, which means that our detection platform is indeed an effective online system.
Ziyu Wang 0007, Jiahai Yang 0001, Hui Zhang 0052, Shize Zhang, Hui Wang 0011
NOMS6
2016 SpongeNet: Towards bandwidth guarantees of cloud datacenter with two-phase VM placement
abstract
In today's production-grade cloud datacenter, cloud service providers do not offer any bandwidth guarantees between VMs, which results in unpredictable performance of tenants' applications. To address this issue, we present SpongeNet, a solution that provides bandwidth guarantees for tenants with a novel network abstraction model and a two-phase VM placement algorithm. Prior solutions have significant limitations: 1) the existing coarse-grained network abstraction models cannot fully express tenants' network requirements and waste a lot of bandwidth resources in demand level; 2) the prior VM placement algorithms, take neither the two scheduling phases nor the tenants' requirements into consideration. As an extension of the existing studies, the proposed network abstraction model in this paper, called Fine-grained Virtual Cluster or FGVC, provides a more precise and flexible way for tenants to specify network requirements and realizes bandwidth saving. SpongeNet also proposes a novel two-phase VM placement algorithm that provides the optimal combinations of ordering policies and dispatching policies in consideration of different goals. Extensive simulations based on real application traces and 3-level tree topology show that SpongeNet provides 48% bandwidth saving than the state-of-art solutions (e.g., the Oktopus system), while significantly improving the throughput rates by 18% and response times by 92%.
Hui Yu 0004, Jiahai Yang 0001, Hui Wang 0011, Zi Liang
NOMS4
2015 MuLTI: Multiple location tags inference for users in social networks
abstract
Social networks, with tremendous popularity all over the world, have become the most important platform for many services in the past years. Location, as part of users' basic information, is always the key to many recommendation services in social networks. Most of the previous research works focus on inferring on the users' home locations. However, it is not enough as many people in social networks have multiple location tags, including home location, work location and on. In this paper, we propose a multiple location tags inference algorithm, i.e. MuLTI to build complete location profiles for users in social networks. We formulate the correlations between the users' location tags and their friendships, tweets, and then infer the users' locations in each of their friendships and tweets. It reflects the activity level of users to be in different locations. Apart from the activity level, we also consider the time span of users to be in different locations, so as to infer the users' long-term location tags better, as we find that users may also be active in their temporal locations. Experiments show that MuLTI improves the precision by about 15%, and the recall by about 25% compared with the state-of-the-art algorithms.
Zejia Chen, Jiahai Yang 0001, Hui Wang 0011
ISCC3
2015 Solving multicast problem in cloud networks using overlay routing
Hui Wang 0011, Jeffrey Cai, Jerry Lu, Kevin Yin, Jiahai Yang 0001
Comput. Commun.1
2015 Supporting Serendipitous Social Interaction Using Human Mobility Prediction
abstract
Leveraging the regularities of people's trajectories, mobility prediction can help forecast social interaction opportunities. In this paper, in order to facilitate real-world social interaction, we aim to predict “serendipitous” social interactions, which are defined as unplanned encounters and interaction opportunities and regarded as emerging social interactions. We collected GPS trajectory data from people' daily life on campus and use it as empirical mobility traces to generate decision trees and model trees to predict next venues, arrival times, and user encounter. Mobility regularities are mainly considered in these prediction models, and mobility contexts (e.g., time, location, and speed) act as decision nodes in the classification trees. Experimental results using collected GPS data showed that our system achieves 90% accuracy for predicting a user's next venue using a decision tree algorithm, with minute-level (around 5 min) prediction error for arrival time using the model tree algorithm. Two prototype applications were developed to support serendipitous social interaction on campus, and the feedback from a user study with 25 users demonstrated the usability of these two applications.
Zhiwen Yu 0001, Hui Wang 0011, Bin Guo 0001, Tao Gu 0001, Tao Mei 0001
IEEE Trans. Hum. Mach. Syst.2
2014 A cascading framework for uncovering spammers in social networks
abstract
With tremendous popularity, OSNs have become the most important platform for marketing and advertising during the past years. Meanwhile, spamming has already become a very serious problem in OSNs, drawing the attention of both academic and industry communities. In this paper, we investigate the problem of spammer detection from the perspective of user behaviors, including relation creation, user activeness, user interaction and tweet content. We quantitatively explore their correlations with spammer detection and find that tweet content is the most important factor for spammer detection, followed by relation creation. Based on these behavior factors, we propose a novel cascading framework CWB-SPAM for spammer detection in OSNs. Experiments on dataset crawled from Sina Microblog show that the proposed algorithm outperforms over all classical algorithms we investigated in terms of F-score1• Experiments also demonstrate that as a probabilistic classification model, the proposed CWB-SPAM has a good ranking quality. It enables the OSN operators to make tradeoff between precision and recall easily so that the proposed algorithm can be used in different scenarios. Besides, we also note that the proposed framework can be used in other probabilistic binary classification models and thus applied in more scenarios.
Zejia Chen, Jiahai Yang 0001, Hui Wang 0011
Networking3
2014 A value based framework for provider selection of regional ISPs
abstract
To access to the Internet, regional Internet Service Providers (ISPs) have to buy transit service from global ISPs. Provider selection strategies are related closely to ISPs' economic interests. With the growing number of potential transit provider and the flattening topology of the Internet, it's getting harder for ISPs to select upstream provider empirically as before. In this paper, we propose a concept of bargaining power as an important decision-making criterion of ISPs during their provider selection process, and design a value based framework to help ISPs' provider selection based on it. The bargaining power of each involved ISP is computed by applying the Shapley Value based transit value distribution mechanism to each involved traffic flow, taking into consideration the cooperative possibility among ISPs and the market roles these ISPs play, i.e., potential providers or potential competitors. It reflects not only the cost and link level transit performance, geographical constraints, but also includes the influence of interconnection impacts, demand/supply relationships by analyzing the traffic content and commercial relationships among ISPs. We then design a quantitative provider selection framework and instantiate our framework using the operation data of a real-world network, CERNET, a national ISP in China. In addition, we evaluate our provider selection results for CERNET and the experimental results show the effectiveness and practicability of our solution in this paper.
Hui Wang 0011, Jiahai Yang 0001
NOMS2
2012 AMIR: Another Multipath Interdomain Routing
abstract
Multipath routing is an important and promising technique to increase the Internet's reliability and to give users greater control over the service they receive. Currently the interdomain routing protocol limits each router to using a single route for a destination network, which does not satisfy the diverse requirements of end users. In this paper, in order to support the effective and efficient multipath routing, we propose a multipath interdomain routing system(AMIR), which not only provides more novel paths but also realizes a new AS-level routing scheme. In the control plane, the topology information is collected from neighboring ASes around the primary path, and based on this topology, multipath of the special node pairs is calculated by our multipath discovery algorithm. In the data plane, we use the interdomain source routing to forward the packets. Experiments with Internet topology and routing data demonstrate that AMIR is practical and feasible, and offers tremendous flexibility and diversity for path selection with reasonable overhead.
Donghong Qin, Jiahai Yang 0001, Zhuolin Liu, Hui Wang 0011, Bin Zhang 0016
AINA4
2012 Flattening and preferential attachment in the internet evolution
abstract
Understanding of the Internet evolution is important for many research topics, such as network planning, optimal routing design, etc. In this paper, we try to analyze CAIDA AS-level topology dataset from 2004 to 2010 to validate two conjectures on the Internet evolution, i.e., the Internet flattening trend and the preferential attachment rule. Our analysis shows that the evolvement of the Internet core is different from the edge of Internet. We classify the Internet into several layers using different layering methods, i.e., Rich Club coefficient based method, k-core decomposition method and SARK hierarchy model, and then study the changes of the features of these layers. Under all of these laying methods, we find that the boundaries between neighboring layers in the Internet core are more and more blurred; ASes in the core distribute more evenly and different layers are closer to each other in size, while the Internet edge still has a distinct hierarchical characteristic. It is more evident in Asia and Europe than North America. The other difference between Internet core and Internet edge is that link births/deaths in the Internet core follow the “Preferential Attachment/de-attachment” rule, while link births/deaths in the Internet edge follow a super linear preferential attachment/de-attachement rule. On the other hand, in both Internet core and Internet edge, link births caused by AS births present stronger preference than link rewiring.
Hui Wang 0011, Jiahai Yang 0001
APNOMS2
2012 Separating identifier from locator with extended DNS
abstract
Although most researchers have agreed that the locator/identifier separation is beneficial for the Internet, there is no consensus on how to define the “identifier”. In this paper, we propose a scheme in which identifiers are distributed by authorities to endpoints, and the authorities are responsible for maintaining real-time locators of endpoints with identifiers they distributed. This scheme can be helpful to accounting, security and other network management tasks. We also present in details how to implement this scheme with extended DNS and a new infrastructure, i.e., ID Mapping System (IDMS).
Hui Wang 0011, Mingwei Xu 0001, Jiahai Yang 0001
ICC1
2011 MIB design and application for source address validation improvement protocol
abstract
In this paper, we present SAVI-MIB, a management information base (MIB) designed to support configuration and monitoring of SAVI protocol which can provide fine granularity source address validation. Objects are defined to meet the detailed management requirement of local networks and accommodate different scenarios. SAVI-MIB is implemented in switches and deployed in some campus networks. Objects of SAVI-MIB are retrieved and used to help find configuration errors in SAVI deployment and profile behavior of end hosts. SAVI-MIB can also be used in parameter optimization, auto configuration, anomaly detection, etc.
Changqing An, Hui Wang 0011, Jiahai Yang 0001
ISCC2
2011 A study of traffic, user behavior and pricing policies in a large campus network
Hui Wang 0011, Changqing An, Jiahai Yang 0001
Comput. Commun.1
2011 A study on key strategies in P2P file sharing systems and ISPs' P2P traffic management
Hui Wang 0011, Chungang Wang, Jiahai Yang 0001, Changqing An
Peer-to-Peer Netw. Appl.1
2010 A measure of growth of user community in OSNs
abstract
Online Social Networks (OSNs) become more and more popular in recent years, with almost several hundred million users involved. A lot of research efforts have been done on features like analysis of the structure of user network, statistics of static properties and dynamic characteristics in OSNs, while little is known about the way how the communities of users grow bigger. This paper proposes a new methodology in OSN research to study the long-run evolution of user community and give guidance for managing network traffic and web caching policy, and a few discussions are also presented.
Jiahai Yang 0001, Hui Wang 0011, Hui Zhang 0052
IWQoS3
2009 Understanding Web Hosting Utility of Chinese ISPs
Guanqun Zhang, Hui Wang 0011, Jiahai Yang 0001
APNOMS2
2008 Understanding IPv6 Usage: Communities and Behaviors
Shaojun Huang, Changqing An, Hui Wang 0011, Jiahai Yang 0001
APNOMS3
2008 A game-theoretic analysis of the implications of overlay network traffic on ISP peering
Hui Wang 0011, Dah-Ming Chiu, John C. S. Lui
Comput. Networks1
2007 Inter-AS Inbound Traffic Engineering via ASPP
abstract
AS Path Prepending (ASPP) is a popular method for the inter-AS inbound traffic engineering, which is known to be more difficult than the outbound traffic engineering. Although the ASPP approach has been extensively practised by many ASes, it is surprising that there still lacks a systematic study of this approach and the basic understanding of its effectiveness. In this paper, we introduce the concept, applicability and potential instability problem of the ASPP approach. Some guidelines are given as the first step to study the method to avoid instability problem. Finally, we study the dynamic prepending behavior of ISPs and show a real-world pathologic case of prepending instability based on our measurement study of RouteViews data.
Hui Wang 0011, Dah-Ming Chiu, John C. S. Lui, Rocky K. C. Chang
IEEE Trans. Netw. Serv. Manag.1
2006 Modeling the Peering and Routing Tussle between ISPs and P2P Applications
abstract
The connectivity between millions of nodes on the Internet is provided by the interconnection of many ISPs' networks. These ISPs, in their decisions to peer with each other, define a set of transit relationships. These transit relationships are the primary factors that dictate how traffic flows through the Internet. BGP-based inter-domain routing that implements these transit relationships can be considered economically efficient. The advent of peer-to-peer (P2P) applications and overlay networks, however, changes the rules by providing traffic routing favoring the applications' needs. This can lead to reduced economic efficiency and upset the ISPs' business model. In this paper, we propose simple models to represent P2P traffic demand, peering and routing in a market place of two competing ISPs to illustrate this tussle of the Internet. Based on these models, we also propose and investigate alternative peering and provisioning strategies available to the ISPs and analyze their effectiveness
Hui Wang 0011, Dah-Ming Chiu, John C. S. Lui
IWQoS1
2005 PORT: A Price-Oriented Reliable Transport Protocol for Wireless Sensor Networks
abstract
In wireless sensor networks, to obtain reliability and minimize energy consumption, a dynamic rate-control and congestion-avoidance transport scheme is very important. We notice that reporting packets may contribute to the sink's fidelity of its knowledge on the phenomenon of interest to different extents. Thus, reliability cannot simply be measured by the sink's total incoming packet rate as considered in current schemes. Also, communication costs between sources and the sink may be different and may change dynamically. Based on these considerations, we propose PORT (price-oriented reliable transport protocol) to facilitate the sink to achieve reliability. Under the constraint that the sink must obtain enough fidelity for reliability purpose, PORT minimizes energy consumption with two schemes. One is based on the sink's application-based optimization approach that feeds back the optimal reporting rates. The other is a locally optimal routing scheme according to the feedback of downstream communication conditions. PORT can adapt well to the communication conditions for energy saving while maintaining the necessary level of reliability. Simulation results in an application case study demonstrate the effectiveness of PORT
Yangfan Zhou 0002, Michael R. Lyu, Jiangchuan Liu, Hui Wang 0011
ISSRE4