VLDB 2026 Research / reviewers in the wild / expert
Zihui Ge
dblp:76/5778
· DBLP profile ↗
62ranked-venue papers
4as first author
4since 2021 · last 2023
0000-0003-0114-7584ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 48 · 3 first-author · 4 since 2021Systems, architecture and hardware · 6Security and privacy · 6Software engineering, systems software and programming languages · 3Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Chroma: Learning and Using Network Contexts to Reinforce Performance Improving ConfigurationsabstractManaging network configuration and improving service experience effectively is essential for cellular service providers (CSPs). This is challenging because of cellular networks' large scale and complexity, the wide variety of configuration parameters, and the performance impact tradeoffs resulting across multiple metrics and geographical locations. This paper focuses on learning and using network contexts to recommend performance-improving configurations. While learning contexts, one must carefully account for the configuration parameter dependency, performance impact confusion that can arise due to co-occurring unrelated changes, and uneven change deployment distribution across locations. We present a new solution Chroma that addresses the above challenges. Using real-world data collected from a large operational LTE and 5G cellular service provider, we thoroughly evaluate and demonstrate the efficacy of Chroma. We successfully trial Chroma on an operational cellular network and highlight its benefits in practical settings. Changhan Ge, Zihui Ge, Xuan Liu 0002, Ajay Mahimkar, Yusef Shaqalle, Yu Xiang 0003, Shomik Pathak |
MobiCom | 2 |
| 2022 | Aurora: conformity-based configuration recommendation to improve LTE/5G serviceabstractCellular service operators frequently tune the network configuration to optimize coverage, support seamless handovers, minimize channel interference, and improve the service performance experience to the end-users. Tuning such a complicated network is highly challenging because of the many configuration parameters, evolving complexity of cellular networks, and diverse requirements across voice, video, and data applications. Any misconfigurations or even poor settings can significantly negatively impact service quality. In this paper, we propose a new approach Aurora that derives best practices knowledge from exploration of the massive existing configuration in the network and uses conformity-based recommendation with performance-based filtering to improve cellular service. We implemented and evaluated Aurora using data from a very large LTE and 5G cellular service provider. Our operational experience over the last one year highlights the benefits of Aurora and exposes exciting research opportunities and challenges in configuration tuning and performance management. Ajay Mahimkar, Zihui Ge, Xuan Liu 0002, Yusef Shaqalle, Yu Xiang 0003, Jennifer Yates, Shomik Pathak, Rick Reichel |
IMC | 2 |
| 2021 | Minimizing Effort and Risk with Network Change Deployment PlanningabstractNetworks undergo continuous changes to introduce new services and improve existing ones. Network change deployment involves carefully deciding when each change activity will be executed and who will be executing the change. This is a complex process because each service group has to plan its activities following a set of operational and technological constraints. Besides, multiple groups may be working on the same or dependent nodes at the same time, and they must coordinate their deployment plans. If they do not co-ordinate, conflicting change execution could result in unexpected impacts. Traditionally, change deployment has been a tedious and time-consuming task. To address this, we propose an innovative solution Zapper that aims for minimal human effort to coordinate the changes, minimal risk to service quality, and efficient plans to rapidly deploy the changes. Zapper maps change scheduling constraints into mathematical equations and then uses optimization algorithms to generate conflict-free change plans that satisfy all constraints across service groups. We have deployed Zapper at a large service provider and it is being used regularly by the network operations teams for more than two years to schedule over 4.5 million change activities. Carlos Eduardo de Andrade, Ajay Mahimkar, Rakesh K. Sinha, Weiyi Zhang 0001, André Augusto Ciré, Giritharan Rana, Zihui Ge, Sarat C. Puthenpura, Jennifer Yates, Robert Riding |
Networking | 7 |
| 2021 | Auric: using data-driven recommendation to automatically generate cellular configurationabstractCellular service providers add carriers in the network in order to support the increasing demand in voice and data traffic and provide good quality of service to the users. Addition of new carriers requires the network operators to accurately configure their parameters for the desired behaviors. This is a challenging problem because of the large number of parameters related to various functions like user mobility, interference management and load balancing. Furthermore, the same parameters can have varying values across different locations to manage user and traffic behaviors as planned and respond appropriately to different signal propagation patterns and interference. Manual configuration is time-consuming, tedious and error-prone, which could result in poor quality of service. In this paper, we propose a new data-driven recommendation approach Auric to automatically and accurately generate configuration parameters for new carriers added in cellular networks. Our approach incorporates new algorithms based on collaborative filtering and geographical proximity to automatically determine similarity across existing carriers. We conduct a thorough evaluation using real-world LTE network data and observe a high accuracy (96%) across a large number of carriers and configuration parameters. We also share experiences from our deployment and use of Auric in production environments. Ajay Mahimkar, Ashiwan Sivakumar, Zihui Ge, Shomik Pathak, Karunasish Biswas |
SIGCOMM | 3 |
| 2019 | Egret: simplifying traffic management for physical and virtual network functionsabstractTraffic migration is a common procedure performed by operators during planned maintenance and unexpected incidents to prevent/reduce service disruptions. However, current practices of traffic migration often couple operators' intentions (e.g. device upgrades) with network setups (e.g. load-balancers), resulting in poor re-usability and substantial operational complexities. Our study of 205 Methods of Procedure (MOPs) from a major U.S. carrier suggests that generalizing traffic migration with a unified model is feasible. Such generalization along with SDN's automation capability is key to scalable and flexible management of traffic, especially for virtualized network functions with unprecedented scale, heterogeneity, and fast iteration. In this paper, we propose Egret, a generic traffic migration system that simplifies traffic management for physical and virtual network functions. Egret (1) hides intricate implementation details from operators with generic intention-based interfaces, and (2) modularizes common traffic migration procedures to enable plug-and-play by developers and vendors. Leveraging a novel mask-based abstraction of traffic migration jobs, Egret can further simplify reverse traffic migration and enable job interleaving. Yikai Lin, Ajay Mahimkar, Bo Han 0001, Zihui Ge, Vijay Gopalakrishnan, Z. Morley Mao |
CoNEXT | 4 |
| 2019 | Rigorous, Effortless and Timely Assessment of Cellular Network ChangesabstractCellular service providers continuously deploy changes in their network in the form of new software releases, service feature introductions, configuration changes, equipment re-homes, firmware upgrades, and topology modifications. It is important to carefully assess the impact of these changes on service performance to validate expected behaviors and take mitigation actions in a timely fashion in case of any unexpected degradation. The diverse nature of the network changes, complex interactions across different layers of the cellular network, and the rapid evolution of the network make it challenging to accurately conduct the assessment. In this paper, we present the design and implementation of our system that enables rigorous, effortless and timely assessment of performance around network changes. We share our lessons learned from the deployment in an operational cellular network over the last eight years. Ajay Mahimkar, Zihui Ge, Sanjeev Ahuja, Shomik Pathak, Nauman Shafi |
DSN | 2 |
| 2018 | Predictive Analysis in Network Function Virtualization
Zhijing Li 0001, Zihui Ge, Ajay Mahimkar, Jia Wang 0001, Ben Y. Zhao, Haitao Zheng 0001, Joanne Emmons, Laura Ogden |
Internet Measurement Conference | 2 |
| 2017 | AutoFocus: Automatically scoping the impact of anomalous service eventsabstractNetworks, and the services they enable, are increasingly diverse and highly utilized. From DSL and fiber-to-the-home access networks, to cellular mobile networks, to contentdelivery networks; all require extensive monitoring in order to meet the increase of user expectations of the availability and quality of those services provided to them. The complexity of these networks and services require better management on the part of providers as the data resulting from service monitoring experiences an increase in dimensionality, making it difficult to fully interpret anomalies in the data. For example, anomaly detection generally says “I found an anomaly with mobile phone A in market Z”. But it is more useful to know what other phones and what other markets are also experiencing the same anomaly. Ren Quinn, Zihui Ge, Jacobus E. van der Merwe |
CNSM | 2 |
| 2017 | Reflection: Automated test location selection for cellular network upgradesabstractCellular networks are constantly evolving due to frequent changes in radio access and end user equipment technologies, dynamic applications and associated trafflc mixes. Network upgrades should be performed with extreme caution since millions of users heavily depend on the cellular networks for a wide range of day to day tasks, including emergency and alert notifications. Before upgrading the entire network, it is important to conduct field evaluation of upgrades. Field evaluations are typically cumbersome and can be time consuming; however if done correctly they can help alleviate a lot of the deployment issues in terms of service quality degradation. The choice and number of field test locations have significant impacts on the time-to-market as well as confidence in how well various network upgrades will work out in the rest of the network. In this paper, we propose a novel approach — Reflection to automatically determine where to conduct the upgrade field tests in order to accurately identify important features that affect the upgrade. We demonstrate the effectiveness of Reflection using extensive evaluation based on real traces collected from a major US cellular network as well as synthetic traces. Mubashir Adnan Qureshi, Ajay Mahimkar, Lili Qiu, Zihui Ge, Sarat C. Puthenpura, Nabeel Mir, Sanjeev Ahuja |
ICNP | 4 |
| 2017 | Coordinating rolling software upgrades for cellular networksabstractCellular service providers continuously upgrade their network software on base stations to introduce new service features, fix software bugs, enhance quality of experience to users, or patch security vulnerabilities. A software upgrade typically requires the network element to be taken out of service, which can potentially degrade the service to users. Thus, the new software is deployed across the network using a rolling upgrade model such that the service impact during the roll-out is minimized. A sequential roll-out guarantees minimal impact but increases the deployment time thereby incurring a significant human cost and time in monitoring the upgrade. A network-wide concurrent roll-out guarantees minimal deployment time but can result in a significant service impact. The goal is to strike a balance between deployment time and service impact during the upgrade. In this paper, we first present our findings from analyzing upgrades in operational networks and discussions with network operators and exposing the challenges in rolling software upgrades. We propose a new framework Concord to effectively coordinate software upgrades across the network that balances the deployment time and service impact. We evaluate Concord using real-world data collected from a large operational cellular network and demonstrate the benefits and tradeoffs. We also present a prototype deployment of Concord using a small-scale LTE testbed deployed indoors in a corporate building. Mubashir Adnan Qureshi, Ajay Mahimkar, Lili Qiu, Zihui Ge, Max Zhang, Ioannis Broustis |
ICNP | 4 |
| 2017 | Monitoring quality-of-experience for operational cellular networks using machine-to-machine trafficabstractIt is crucial for cellular data network operators to understand the service quality perceived by its customers. The state-of-art systems deployed in cellular networks mostly report service quality aggregated on cell site level, which is typically an aggregation of tens or hundreds of customers depending on the locations of the cell sites. In this paper, we propose to enhance the measurement of customer-perceived service quality by leveraging M2M devices as sensors in the field, which provide an unprecedented opportunity for cellular network operators to measure what end-users experience with better accuracy and coverage. Our approach is to identify a set of M2M devices which are stationary and communicate continuously over the cellular network over an indefinite period of time. We use these M2M devices to estimate the customer-perceived service quality during cell site outages. We implement our methodology as a system called M2MScan and evaluate M2MScan with both synthetic outages and real outages from a large-scale operational cellular network. To the best of our knowledge, this is the first work that employs M2M devices to measure the service quality perceived by customers in operational cellular networks at a large scale. Faraz Ahmed, Jeffrey Erman, Zihui Ge, Alex X. Liu, Jia Wang 0001 |
INFOCOM | 3 |
| 2017 | AESOP: Automatic Policy Learning for Predicting and Mitigating Network Service ImpairmentsabstractEfficient management and control of modern and next-gen networks is of paramount importance as networks have to maintain highly reliable service quality whilst supporting rapid growth in traffic demand and new application services. Rapid mitigation of network service degradations is a key factor in delivering high service quality. Automation is vital to achieving rapid mitigation of issues, particularly at the network edge where the scale and diversity is the greatest. This automation involves the rapid detection, localization and (where possible) repair of service-impacting faults and performance impairments. However, the most significant challenge here is knowing what events to detect, how to correlate events to localize an issue and what mitigation actions should be performed in response to the identified issues. These are defined as policies to systems such as ECOMP. Supratim Deb, Zihui Ge, Sastry Isukapalli, Sarat C. Puthenpura, Shobha Venkataraman, Jennifer Yates |
KDD | 2 |
| 2017 | Firewall Fingerprinting and Denial of Firewalling AttacksabstractFirewalls are critical security devices handling all traffic in and out of a network. Firewalls, like other software and hardware network devices, have vulnerabilities, which can be exploited by motivated attackers. However, just like any other networking and computing devices, firewalls often have vulnerabilities that can be exploited by attackers. In this paper, first, we investigate some possible firewall fingerprinting methods and surprisingly found that these methods can achieve quite high accuracy. Second, we study what we call denial of firewalling (DoF) attacks, where attackers use carefully crafted traffic to effectively overload a firewall. To the best of our knowledge, this paper represents the first study of firewall fingerprinting and DoF attacks. Alex X. Liu, Amir R. Khakpour, Joshua W. Hulst, Zihui Ge, Dan Pei, Jia Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2017 | Detecting and Localizing End-to-End Performance Degradation for Cellular Data Services Based on TCP Loss Ratio and Round Trip TimeabstractProviding high end-to-end (E2E) performance experienced by users is critical for cellular service providers to best serve their customers. This paper focuses on the detection and localization of E2E performance degradation (such as slow webpage page loading and unsmooth video playing) at cellular service providers. Detecting and localizing E2E performance degradation is crucial for cellular service providers, content providers, device manufactures, and application developers to jointly troubleshoot root causes. To the best of our knowledge, the detection and localization of E2E performance degradation at cellular service providers has not been previously studied. In this paper, we propose a holistic approach to detecting and localizing E2E performance degradation at cellular service providers across the four dimensions of user locations, content providers, device types, and application types. Our approach consists of three steps: modeling, detection, and localization. First, we use training data to build models that can capture the normal performance of every E2E instance, which means the flows corresponding to a specific location, content provider, device type, and application type. Second, we use our models to detect performance degradation for each E2E instance on an hourly basis. Third, after each E2E instance has been labeled as non-degrading or degrading, we use association rule mining techniques to localize the source of performance degradation. Our system detected performance degradation instances over a period of one week. In 80% of the detected degraded instances, content providers, device types, and application types were the only factors of performance degradation. Faraz Ahmed, Jeffrey Erman, Zihui Ge, Alex X. Liu, Jia Wang 0001 |
IEEE/ACM Trans. Netw. | 3 |
| 2016 | Quantifying the service performance impact of self-organizing network actionsabstractAs smartphone users increasingly rely on cellular networks to access voice, video, and web applications, guaranteeing good performance and high availability is more important than ever. Historically, managing cellular network configuration has been a manual, error-prone process; recently, automated solutions such as SON (Self-Organizing Networks) controllers are being deployed for dynamic tuning of network configuration to improve end-user service performance under dynamic network and traffic conditions. SON automates many aspects of cellular network configuration, but it is nonetheless susceptible to software bugs and expected traffic changes that could result in sub-optimal performance. In this paper, we propose a capability (Veracity) to analyze and quantify the performance effects of SON actions. Assessing the effects of SON control is difficult because of the dynamic nature of SON and the dependency of end-user performance on factors such as radio channel quality, mobility and traffic load. Veracity addresses these using model-driven impact detection and quantification. Our evaluation using data collected from an operational cellular network demonstrates that Veracity is accurate. Veracity is now being used by the service providers' field operation teams for the assessment of SON effectiveness in arenas and stadiums. Swati Roy, David L. Applegate, Zihui Ge, Ajay Mahimkar, Shomik Pathak, Sarat C. Puthenpura |
CNSM | 3 |
| 2016 | Detecting and localizing end-to-end performance degradation for cellular data servicesabstractProviding high end-to-end (E2E) performance is critical for cellular service providers to best serve their customers. Detecting and localizing E2E performance degradation is crucial for cellular service providers, content providers, device manufactures, and application developers to jointly troubleshoot root causes. To the best of our knowledge, detection and localization of E2E performance degradation at cellular service providers has not been previously studied. In this paper, we propose a holistic approach to detecting and localizing E2E performance degradation at cellular service providers across the four dimensions of user locations, content providers, device types, and application types. First, we use training data to build models that can capture the normal performance of every E2E-instance, which means flows corresponding to a specific location, content provider, device type, and application type. Second, we use our models to detect performance degradation for each E2E-instance on an hourly basis. Third, after each E2E-instance has been labeled as non-degrading or degrading, we use association rule mining techniques to localize the source of performance degradation. Our system detected performance degradation instances over a period of one week. In 80% of the detected degraded instances, content providers, device types, and application types were the only factors of performance degradation. Faraz Ahmed, Jeffrey Erman, Zihui Ge, Alex X. Liu, Jia Wang 0001 |
INFOCOM | 3 |
| 2016 | Libra: Impact assessment of cellular load balancingabstractLoad on cellular towers is one of the key metrics that cellular service providers monitor as part of their operational and management tasks. Increased load on the towers can lead to congestion, which in turn can severely degrade quality of service perceived by users. Hence, it is of great interest to cellular service providers to minimize the maximum load at cell towers, and thereby minimize chances of congestion in the event of a sudden increase in load due to user demand changes. This goal can be achieved by proactive load balancing among neighboring cell towers, i.e., proactively identify opportunities to balance the load through re-binding of users from heavily loaded cell towers to lightly loaded neighboring towers. In this paper, we propose a new tool Libra to effectively assess the impact of load balancing related parameter changes. Libra provides an objective measure of the degree of load imbalance across multiple network locations and identifies if the measure improves or degrades after parameter changes. Our evaluation of Libra using real-world data collected from a large cellular provider demonstrates its effectiveness in accurately capturing the degree of imbalance at multiple cell towers. Kanthi Nagaraj, Ajay Mahimkar, Zihui Ge, Aman Shaikh, Jia Wang 0001, Kevin Mohr, Mark Stockert |
INFOCOM | 3 |
| 2016 | Automated Test Location Selection For Cellular Network UpgradesabstractCellular networks are constantly evolving due to frequent changes in radio access and end user equipment technologies, applications, and traffic. Network upgrades should be performed with extreme caution since millions of users heavily depend on the cellular networks. Before upgrading the entire network, it is important to conduct field evaluation of upgrades.The choice and number of field test locations have significant impact on the time-to-market and confidence in how well various network upgrades will work out in the rest of the network. We propose a novel approach -- Reflection to automatically determine where to conduct the upgrade field tests to accurately identify important features that affect the upgrade and predict for the performance of untested locations. We demonstrate its effectiveness using real traces collected from a major US cellular network as well as synthetic traces. Mubashir Adnan Qureshi, Ajay Mahimkar, Lili Qiu, Zihui Ge, Sarat C. Puthenpura, Nabeel Mir, Sanjeev Ahuja |
SIGMETRICS | 4 |
| 2015 | Magus: minimizing cellular service disruption during network upgradesabstractPlanned upgrades in cellular networks occur every day, may often need to be performed on weekdays, and can potentially degrade service for customers. In this paper, we explore the problem of tuning network configurations in order to mitigate any potential impact due to a planned upgrade which takes the base station off-air. The objective is to recover the loss in service performance or coverage which would have occurred without any modifications. To our knowledge, impact mitigation for planned base station downtimes has not been explored before in the literature. The primary contribution of this work is a proactive approach based on a predictive model that uses operational data of user density distributions and path loss (rather than idealized analytical models of these) to quickly estimate the best power and tilt configuration of neighboring base stations that enables high recovery. A secondary contribution is an approach to minimize synchronized handovers. These ideas, embodied in a capability called Magus, enables us to recover up to 76% of the potential performance loss due to planned upgrades in some cases for a large US mobile network, and this recovery varies as a function of base station density. Moreover, Magus is able to reduce synchronized handovers by a factor of 8. Ioannis Broustis, Zihui Ge, Ramesh Govindan, Ajay Mahimkar, N. K. Shankaranarayanan, Jia Wang 0001 |
CoNEXT | 3 |
| 2015 | ABSENCE: Usage-based Failure Detection in Mobile NetworksabstractWe present our proposed ABSENCE system which detects service disruptions in mobile networks using aggregated customer usage data. ABSENCE monitors aggregated customer usage to detect when aggregated usage is lower than expected in a given geographic region (e.g., zip code), across a given customer device type, or for a given service. Such a drop in expected usage is interpreted as a sign of a potential service disruption being experienced in that region/device type/service. ABSENCE effectively deals with users' mobility and scales to detect failures in various mobile services (e.g., voice, data, SMS, MMS, etc). We perform a systematic evaluation of our proposed approach by introducing synthetic failures in measurements obtained from a US operator. We also compare our results with ground truth (real service disruptions) obtained from the mobile operator. Binh Nguyen 0003, Zihui Ge, Jacobus E. van der Merwe, Jennifer Yates |
MobiCom | 2 |
| 2015 | Detecting and Localizing End-to-End Performance Degradation for Cellular Data ServicesabstractNowadays mobile device (e.g., smartphone) users not only have a high expectation on the availability of the cellular data service, but also increasingly depend on the high end-to-end (E2E) performance of their applications. Since the E2E performance of individual application sessions may vary greatly, depending on factors such as the cellular network condition, the content provider, the type/model of the mobile devices, and the application software, detecting and localizing service performance degradations in a timely manner at large scale is of great value to cellular service providers. In this paper, we build a holistic measurement system that tracks session-level E2E performance metrics along with the service attributes for these factors. Using data collected from a major cellular service provider, we first model the expected E2E service performance with a regression based approach, detect performance degradation conditions based on the time series of fine-grained measurement data, and finally localize the service degradation using association-rule-mining techniques. Our deployment experience reveals that in 80% of the detected problem instances, performance degradation can be attributed to non-network-location specific factors, such as a common content provider, or a set of applications running on certain models of devices. Faraz Ahmed, Jeffrey Erman, Zihui Ge, Alex X. Liu, Jia Wang 0001 |
SIGMETRICS | 3 |
| 2014 | Crossroads: A Practical Data Sketching Solution for Mining Intersection of StreamsabstractThe explosive increase in cellular network traffic, users, and applications, as well as the corresponding shifts in user expectations, has created heavy needs and demands on cellular data providers. In this paper we address one such need: mining the logs of cellular voice and data traffic to rapidly detect network performance anomalies and other events of interest. The core challenge in solving this problem is the issue that it is impossible to predict beforehand where in the traffic the event may appear, requiring us to be able to query arbitrary subsets of the network traffic (e.g., longer than usual round-trip times for users in a specific urban area to connect to FunContent.com using a particular model of phone). Since it is infeasible to store all combinations of such data, especially when it is collected in real-time, we need to be able to summarize the traffic data using succinct sketch data structures to answer these queries. Zhenglin Yu, Zihui Ge, Ashwin Lall, Jia Wang 0001, Jun (Jim) Xu |
Internet Measurement Conference | 2 |
| 2014 | NetSearch: Googling large-scale network management dataabstractIn order to ensure the service quality, modern Internet Service Providers (ISPs) invest tremendously on their network monitoring and measurement infrastructure. Vast amount of network data, including device logs, alarms, and active/passive performance measurement across different network protocols and layers, are collected and stored for analysis. As network measurement grows in scale and sophistication, it becomes increasingly challenging to effectively “search” for the relevant information that best support the needs of network operations. In this paper, we look into techniques that have been widely applied in the information retrieval and search engine domain and explore their applicability in network management domain. We observe that unlike the textural information on the Internet, network data are typically annotated with time and location information, which can be further augmented using information based on network topology, protocol and service dependency. We design NetSearch, a system that pre-processes various network data sources on data ingestion, constructs index that matches both the network spatial hierarchy model and the inherent timing/textual information contained in the data, and efficiently retrieves the relevant information that network operators search for. Through case study, we demonstrate that NetSearch is an important capability for many critical network management functions such as complex impact analysis. Tongqing Qiu, Zihui Ge, Dan Pei, Jia Wang 0001, Jun (Jim) Xu |
Networking | 2 |
| 2013 | Robust assessment of changes in cellular networksabstractCellular network service providers often have to conduct small scale testing in the operational network before a change (e.g., a new feature) is fully rolled out across the entire network. This is referred to as the First Field Application (FFA). However, assessing the effectiveness of FFA changes is challenging because of overlapping external factors: seasonality (foliage, leaves budding), weather (rain, snow, hurricanes, storms), traffic pattern changes due to big events (e.g., games at stadiums, students returning to school after holidays), and network events such as outages or other maintenance activities in different regions. In this paper, we first highlight the technical challenges in assessing the service performance impact of changes in operational cellular networks. We then propose Litmus, a new approach based on a spatial dependency model for robust assessment of changes. We evaluate the effectiveness of Litmus using real-world data from operational cellular networks (GSM, UMTS and LTE). Our operational experiences demonstrate accurate inferences of the service performance impact of changes in the field. Ajay Mahimkar, Zihui Ge, Jennifer Yates, Chris Hristov, Vincent Cordaro, Shane Smith, Mark Stockert |
CoNEXT | 2 |
| 2013 | Proactive call drop avoidance in UMTS networksabstractThe rapid advancement of smartphones has instigated tremendous data applications for cell phones. Supporting simultaneous voice and data services in a cellular network is not only desirable but also becoming indispensable. However, if the voice and data are serviced through the same antenna (like the 3G UMTS network), a voice call with data sessions requires better radio connection than a voice-only call. In this paper, we systematically study the coordination between the voice and data transmissions in UMTS networks. From analyzing a large carrier's UMTS network recording data, we first identify the most relevant network measurements/features indicating a potential call drop, then propose a drop-call predictor based on AdaBoost. Moreover, we develop an intelligent call management strategy to voluntarily block data sessions when the voice is predicted to be dropped. Our analysis utilizing real service provider's data sets shows that our proposed scheme can not only predict drop calls with a very high accuracy but also achieve the highest user satisfaction compared to the other existing call management strategies. Jie Yang 0003, Dahai Xu, Guangzhi Li, Yu Jin 0001, Zihui Ge, Mario Kosseifi, Robert D. Doverspike, Yingying Chen 0001, Lei Ying 0001 |
INFOCOM | 6 |
| 2013 | Modeling Cellular User Mobility Using a Leap Graph
Nick G. Duffield, Zihui Ge, Seungjoon Lee, Jeffrey Pang |
PAM | 3 |
| 2012 | Firewall fingerprintingabstractFirewalls are critical security devices handling all traffic in and out of a network. Firewalls, like other software and hardware network devices, have vulnerabilities, which can be exploited by motivated attackers. However, because firewalls are usually placed in the network such that they are transparent to the end users, it is very hard to identify them and use their corresponding vulnerabilities to attack them. In this paper, we study firewall fingerprinting, in which one can use firewall decisions on TCP packets with unusual flags and machine learning techniques for inferring firewall implementation. Amir R. Khakpour, Joshua W. Hulst, Zihui Ge, Alex X. Liu, Dan Pei, Jia Wang 0001 |
INFOCOM | 3 |
| 2012 | Threshold compression for 3G scalable monitoringabstractWe study the problem of scalable monitoring of operational 3G wireless networks. Threshold-based performance monitoring in large 3G networks is very challenging for two main factors: large network scale and dynamics in both time and spatial domains. A fine-grained threshold setting (e.g., perlocation hourly) incurs prohibitively high management complexity, while a single static threshold fails to capture the network dynamics, thus resulting in unacceptably poor alarm quality (up to 70% false/miss alarm rates). In this paper, we propose a scalable monitoring solution, called threshold-compression that can characterize the location- and time-specific threshold trend of each individual network element (NE) with minimal threshold setting. The main insight is to identify groups of NEs with similar threshold behaviors across location and time dimensions, forming spatial-temporal clusters to reduce the number of thresholds while maintaining acceptable alarm accuracy in a large-scale 3G network. Our evaluations based on the operational experience on a commercial 3G network have demonstrated the effectiveness of the proposed solution. We are able to reduce the threshold setting up to 90% with less than 10% false/miss alarms. Suk-Bok Lee, Dan Pei, Mohammad Hajiaghayi, Ioannis Pefkianakis, Songwu Lu, Zihui Ge, Jennifer Yates, Mario Kosseifi |
INFOCOM | 7 |
| 2012 | Argus: End-to-end service anomaly detection and localization from an ISP's point of viewabstractRecent trends in the networked services industry (e.g., CDN, VPN, VoIP, IPTV) see Internet Service Providers (ISPs) leveraging their existing network connectivity to provide an end-to-end solution. Consequently, new opportunities are available to monitor and improve the end-to-end service quality by leveraging the information from inside the network. We propose a new approach to detect and localize end-to-end service quality issues in such ISP-managed networked services by utilizing traffic data passively monitored at the ISP side, the ISP network topology, routing tables and geographic information. This paper presents the design of a generic service quality monitoring system “Argus”. Argus has been successfully deployed in a tier-1 ISP to monitor millions of users of its CDN service and assist operators to detect and localize end-to-end service quality issues. This operational experience demonstrates that Argus is effective in accurate, quick detection and localization of important service quality issues. Ashley Flavel, Zihui Ge, Alexandre Gerber, Daniel Massey, Christos Papadopoulos, Hiren Shah, Jennifer Yates |
INFOCOM | 3 |
| 2012 | ALERT-ID: Analyze Logs of the Network Element in Real Time for Intrusion Detection
Jie Chu 0003, Zihui Ge, Richard Huber, Ping Ji 0002, Jennifer Yates, Yung-Chao Yu |
RAID | 2 |
| 2012 | G-RCA: a generic root cause analysis platform for service quality management in large IP networksabstractAn increasingly diverse set of applications, such as Internet games, streaming videos, e-commerce, online banking, and even mission-critical emergency call services, all relies on IP networks. In such an environment, best-effort service is no longer acceptable. This requires a transformation in network management from detecting and replacing individual faulty network elements to managing the end-to-end service quality as a whole. In this paper, we describe the design and development of a Generic Root Cause Analysis platform (G-RCA) for service quality management (SQM) in large IP networks. G-RCA contains a comprehensive service dependency model that incorporates topological and cross-layer relationships, protocol interactions, and control plane dependencies. G-RCA abstracts the root cause analysis process into signature identification for symptom and diagnostic events, temporal and spatial event correlation, and reasoning and inference logic. G-RCA provides a flexible rule specification language that allows operators to quickly customize G-RCA and provide different root cause analysis tools as new problems need to be investigated. G-RCA is also integrated with data trending, manual data exploration, and statistical correlation mining capabilities. G-RCA has proven to be a highly effective SQM platform in several different applications, and we present results regarding BGP flaps, PIM flaps in Multicast VPN service, and end-to-end throughput degradation in content delivery network (CDN) service. Lee Breslau, Zihui Ge, Daniel Massey, Dan Pei, Jennifer Yates |
IEEE/ACM Trans. Netw. | 3 |
| 2011 | Rapid detection of maintenance induced changes in service performanceabstractService quality in operational IP networks can be impacted due to planned or unplanned maintenance. During any maintenance activity, the responsibility of the operations team is to complete the work order and perform a check-up to ensure there are no unexpected service disruptions. Once the maintenance is complete, it is crucial to continuously monitor the network and look for any performance impacts. What operations lack today are effective tools to rapidly detect maintenance induced performance changes. The large scale and heterogeneity of network elements and performance metrics makes the problem extremely challenging. Ajay Mahimkar, Zihui Ge, Jia Wang 0001, Jennifer Yates, Yin Zhang 0001, Joanne Emmons, Brian Huntley, Mark Stockert |
CoNEXT | 2 |
| 2011 | Q-score: proactive service quality assessment in a large IPTV systemabstractIn large-scale IPTV systems, it is essential to maintain high service quality while providing a wider variety of service features than typical traditional TV. Thus service quality assessment systems are of paramount importance as they monitor the user-perceived service quality and alert when issues occurs. For IPTV systems, however, there is no simple metric to represent user-perceived service quality and Quality of Experience (QoE). Moreover, there is only limited user feedback, often in the form of noisy and delayed customer calls. Therefore, we aim to approximate the QoE through a selected set of performance indicators in a proactive (i.e., detect issues before customers reports to call centers) and scalable fashion. Han Hee Song, Zihui Ge, Ajay Mahimkar, Jia Wang 0001, Jennifer Yates, Yin Zhang 0001, Andrea Basso 0001, Min Chen 0010 |
Internet Measurement Conference | 2 |
| 2011 | Scalable monitoring via threshold compression in a large operational 3G networkabstractThreshold-based performance monitoring in large 3G networks is very challenging for two main factors: large network scale and dynamics in both time and spatial domains. There exists a fundamental tradeoff between the size of threshold settings and the alarm quality. In this paper, we propose a scalable monitoring solution, called threshold-compression that characterizes the tradeoff via intelligent threshold aggregation. The main insight behind our solution is to identify groups of network elements with similar threshold behaviors across location and time dimensions, thus forming spatial-temporal clusters and generating the associated compressed thresholds within the optimization framework. Our evaluations on a commercial 3G network have demonstrated the effectiveness of our threshold-compression solution, e.g., threshold setting reduction up to 90% within 10% false/miss alarms. Suk-Bok Lee, Dan Pei, Mohammad Hajiaghayi, Ioannis Pefkianakis, Songwu Lu, Zihui Ge, Jennifer Yates, Mario Kosseifi |
SIGMETRICS | 7 |
| 2010 | G-RCA: a generic root cause analysis platform for service quality management in large IP networksabstractAs IP networks have become the mainstay of an increasingly diverse set of applications ranging from Internet games and streaming videos, to e-commerce and online-banking, and even to mission-critical 911, best effort service is no longer acceptable. This requires a transformation in network management from detecting and replacing individual faulty network elements to managing the service quality as a whole. Lee Breslau, Zihui Ge, Daniel Massey, Dan Pei, Jennifer Yates |
CoNEXT | 3 |
| 2010 | Listen to me if you can: tracking user experience of mobile network on social mediaabstractSocial media sites such as Twitter continue to grow at a fast pace. People of all generations use social media to exchange messages and share experiences of their life in a timely fashion. Most of these sites make their data available. An intriguing question is can we exploit this real-time and massive data-flow to improve business in a measurable way. In this paper, we are particularly interested in tweets (Twitter messages) that are relevant to mobile network performance. We compare tweets with a more traditional source of user experience, i.e., customer care tickets, and correlate both of them with a list of major network incidents. From our study, we have the following observations. First, Twitter users and users who call customer service tend to report different types of performance issues. Second, we observe that tweets typically appear more rapidly in response to network problems than customer tickets. They also appear to respond to a wider range of network issues. Third, significant spikes in the number of tweets appear to indicate short term performance impairments which are not reported in our current list of major network incidents. These observations together indicate that Twitter is an attractive, complementary source for monitoring service performance and its impact on user experience. Tongqing Qiu, Junlan Feng, Zihui Ge, Jia Wang 0001, Jun (Jim) Xu, Jennifer Yates |
Internet Measurement Conference | 3 |
| 2010 | What happened in my network: mining network events from router syslogsabstractRouter syslogs are messages that a router logs to describe a wide range of events observed by it. They are considered one of the most valuable data sources for monitoring network health and for trou- bleshooting network faults and performance anomalies. However, router syslog messages are essentially free-form text with only a minimal structure, and their formats vary among different vendors and router OSes. Furthermore, since router syslogs are aimed for tracking and debugging router software/hardware problems, they are often too low-level from network service management perspectives. Due to their sheer volume (e.g., millions per day in a large ISP network), detailed router syslog messages are typically examined only when required by an on-going troubleshooting investigation or when given a narrow time range and a specific router under suspicion. Automated systems based on router syslogs on the other hand tend to focus on a subset of the mission critical messages (e.g., relating to network fault) to avoid dealing with the full diversity and complexity of syslog messages. In this project, we design a Sys-logDigest system that can automatically transform and compress such low-level minimally-structured syslog messages into meaningful and prioritized high-level network events, using powerful data mining techniques tailored to our problem domain. These events are three orders of magnitude fewer in number and have much better usability than raw syslog messages. We demonstrate that they provide critical input to network troubleshooting, and net- work health monitoring and visualization. Tongqing Qiu, Zihui Ge, Dan Pei, Jia Wang 0001, Jun (Jim) Xu |
Internet Measurement Conference | 2 |
| 2010 | Crowdsourcing service-level network event monitoringabstractThe user experience for networked applications is becoming a key benchmark for customers and network providers. Perceived user experience is largely determined by the frequency, duration and severity of network events that impact a service. While today's networks implement sophisticated infrastructure that issues alarms for most failures, there remains a class of silent outages (e.g., caused by configuration errors) that are not detected. Further, existing alarms provide little information to help operators understand the impact of network events on services. Attempts to address this through infrastructure that monitors end-to-end performance for customers have been hampered by the cost of deployment and by the volume of data generated by these solutions. David R. Choffnes, Fabián E. Bustamante, Zihui Ge |
SIGCOMM | 3 |
| 2010 | Detecting the performance impact of upgrades in large operational networksabstractNetworks continue to change to support new applications, improve reliability and performance and reduce the operational cost. The changes are made to the network in the form of upgrades such as software or hardware upgrades, new network or service features and network configuration changes. It is crucial to monitor the network when upgrades are made because they can have a significant impact on network performance and if not monitored may lead to unexpected consequences in operational networks. This can be achieved manually for a small number of devices, but does not scale to large networks with hundreds or thousands of routers and extremely large number of different upgrades made on a regular basis. Ajay Mahimkar, Han Hee Song, Zihui Ge, Aman Shaikh, Jia Wang 0001, Jennifer Yates, Yin Zhang 0001, Joanne Emmons |
SIGCOMM | 3 |
| 2010 | Network prefix-level traffic profiling: Characterizing, modeling, and evaluation
Hongbo Jiang 0001, Zihui Ge, Shudong Jin, Jia Wang 0001 |
Comput. Networks | 2 |
| 2009 | Modeling user activities in a large IPTV systemabstractInternet Protocol Television (IPTV) has emerged as a new delivery method for TV. In contrast with native broadcast in traditional cable and satellite TV system, video streams in IPTV are encoded in IP packets and distributed using IP unicast and multicast. This new architecture has been strategically embraced by ISPs across the globe, recognizing the opportunity for new services and its potential toward a more interactive style of TV watching experience in the future. Since user activities such as channel switches in IPTV impose workload beyond local TV or set-top box (different from broadcast TV systems), it becomes essential to characterize and model the aggregate user activities in an IPTV network to support various system design and performance evaluation functions such as network capacity planning. In this work, we perform an in-depth study on several intrinsic characteristics of IPTV user activities by analyzing the real data collected from an operational nation-wide IPTV system. We further generalize the findings and develop a series of models for capturing both the probability distribution and time-dynamics of user activities. We then combine theses models to design an IPTV user activity workload generation tool called SIMUL WATCH, which takes a small number of input parameters and generates synthetic workload traces that mimic a set of real users watching IPTV. We validate all the models and the prototype of SIMUL WATCH using the real traces. In particular, we show that SIMUL WATCH can estimate the unicast and multicast traffic accurately, proving itself as a useful tool in driving the performance study in IPTV systems. Tongqing Qiu, Zihui Ge, Seungjoon Lee, Jia Wang 0001, Jun (Jim) Xu, Qi Zhao 0006 |
Internet Measurement Conference | 2 |
| 2009 | Towards automated performance diagnosis in a large IPTV networkabstractIPTV is increasingly being deployed and offered as a commercial service to residential broadband customers. Compared with traditional ISP networks, an IPTV distribution network (i) typically adopts a hierarchical instead of mesh-like structure, (ii) imposes more stringent requirements on both reliability and performance, (iii) has different distribution protocols (which make heavy use of IP multicast) and traffic patterns, and (iv) faces more serious scalability challenges in managing millions of network elements. These unique characteristics impose tremendous challenges in the effective management of IPTV network and service. In this paper, we focus on characterizing and troubleshooting performance issues in one of the largest IPTV networks in North America. We collect a large amount of measurement data from a wide range of sources, including device usage and error logs, user activity logs, video quality alarms, and customer trouble tickets. We develop a novel diagnosis tool called Giza that is specifically tailored to the enormous scale and hierarchical structure of the IPTV network. Giza applies multi-resolution data analysis to quickly detect and localize regions in the IPTV distribution hierarchy that are experiencing serious performance problems. Giza then uses several statistical data mining techniques to troubleshoot the identified problems and diagnose their root causes. Validation against operational experiences demonstrates the effectiveness of Giza in detecting important performance issues and identifying interesting dependencies. The methodology and algorithms in Giza promise to be of great use in IPTV network operations. Ajay Mahimkar, Zihui Ge, Aman Shaikh, Jia Wang 0001, Jennifer Yates, Yin Zhang 0001, Qi Zhao 0006 |
SIGCOMM | 2 |
| 2008 | Troubleshooting chronic conditions in large IP networksabstractChronic network conditions are caused by performance impairing events that occur intermittently over an extended period of time. Such conditions can cause repeated performance degradation to customers, and sometimes can even turn into serious hard failures. It is therefore critical to troubleshoot and repair chronic network conditions in a timely fashion in order to ensure high reliability and performance in large IP networks. Today, troubleshooting chronic conditions is often performed manually, making it a tedious, time-consuming and error-prone process. Ajay Mahimkar, Jennifer Yates, Yin Zhang 0001, Aman Shaikh, Jia Wang 0001, Zihui Ge, Cheng Tien Ee |
CoNEXT | 6 |
| 2007 | Joint Traffic Blocking and Routing Under Network Failures and MaintenancesabstractUnder device failures and maintenance activities, network resources reduce and congestion may arise inside networks. In this paper, we study a dual approach that combines traffic blocking (rate-limiting) at the edge of a network and traffic rerouting inside the network. We formulate a joint ingress blocking and routing optimization problem and develop mechanisms to introduce blocking differentiations among users with different service priorities and with different level of impact to network congestions. Our evaluation result shows that by blocking only a small fraction of traffic, one can greatly reduce network congestion under severe failures and maintenance activities. Our solution efficiently identifies the optimal blocking among heterogeneous users and achieves much better performance in comparison with proportional traffic blocking. The proposed algorithms can be easily adopted by network service providers in their traffic engineering practices. Chao Liang 0003, Zihui Ge, Yong Liu 0013 |
GLOBECOM | 2 |
| 2007 | OPTWALL: A Hierarchical Traffic-Aware Firewall
Subrata Acharya, Bryan N. Mills, Mehmud Abliz, Taieb Znati, Jia Wang 0001, Zihui Ge, Albert G. Greenberg |
NDSS | 6 |
| 2007 | Design and analysis of a demand adaptive and locality aware streaming media server cluster
Zihui Ge, Ping Ji 0002, Prashant J. Shenoy |
Multim. Syst. | 1 |
| 2007 | A comparison of hard-state and soft-state signaling protocols
Ping Ji 0002, Zihui Ge, James F. Kurose, Don Towsley |
IEEE/ACM Trans. Netw. | 2 |
| 2006 | Traffic-Aware Firewall Optimization StrategiesabstractThe overall performance of a firewall is crucial in enforcing and administrating security, especially when the network is under attack. The continuous growth of the Internet, coupled with the increasing sophistication of the attacks, is placing stringent demands on firewall performance. In this paper, we describe a traffic-aware optimization framework to improve the operational cost of firewalls. Based on this framework, we design a set of tools that inspect and analyze both multidimensional firewall rules and traffic logs and construct the optimal equivalent firewall rules based on the observed traffic characteristics. To the best of our knowledge, this work is the first to use traffic characteristics in firewall optimization. Furthermore, we develop a novel adaptation mechanism that dynamically detects anomalous traffic behavior and adaptively alters the firewall rules to avoid serious performance degradation due to the traffic anomaly. To evaluate the performance of our approaches, we collected a large set of firewall rules and traffic logs at tens of enterprise networks managed by a Tier-1 service provider. Our evaluation results find these approaches very effective. In particular, we achieve more than 10 fold performance improvement by using the proposed traffic-aware firewall optimization. Subrata Acharya, Jia Wang 0001, Zihui Ge, Taieb Znati, Albert G. Greenberg |
ICC | 3 |
| 2006 | Path Protection Routing with SRLG Constraints to Support IPTV in WDM Mesh NetworksabstractThe distribution of broadcast TV across large provider networks has become a highly topical subject as satellite distribution capacity exhausts and competitive pressures increase. In a typical IPTV architecture, broadcast TV is distributed from two sources (for redundancy) to multiple destinations. The aim of this paper is to examine how IPTV can be reliably and cost effectively supported in wavelength division multiplexed (WDM) networks. WDM networks have evolved to mesh topologies and recently to support multicast, which is particularly valuable in reducing the network cost in broadcast TV applications. Our goal is to find two trees with a minimal total cost such that we have two physically (or Shared Risk Link Group (SRLG)) diverse paths to each of the destinations - one from each of the sources. Any two links that belong to a common SRLG are subject to a single point of failure, be it a channel or wavelength failure, a fiber cut, or a complete conduit cut. We first show that our path protection routing problem is NP-complete. We then propose an Integer Programming (IP) formulation for this problem. Using real network topology data, we show that the real networks are amenable to the IP problem formulation and yield optimal solutions. Meeyoung Cha, W. Art Chaovalitwongse, Zihui Ge, Jennifer Yates, Sue B. Moon |
INFOCOM | 3 |
| 2006 | Dynamic cache reconfiguration strategies for cluster-based streaming proxy
Yang Guo 0001, Zihui Ge, Bhuvan Urgaonkar, Prashant J. Shenoy, Don Towsley |
Comput. Commun. | 2 |
| 2005 | Finding Critical Traffic MatricesabstractA traffic matrix represents the amount of traffic between origin and destination in a network. It has tremendous potential utility for many IP network engineering applications, such as network survivability analysis, traffic engineering, and capacity planning. Recent advances in traffic matrix estimation have enabled ISPs to measure traffic matrices continuously. Yet a major challenge remains towards achieving the full potential of traffic matrices. In practical networking applications, it is often inconvenient (if not infeasible) to deal with hundreds or thousands of measured traffic matrices. So it is highly desirable to be able to extract a small number of "critical" traffic matrices. Unfortunately, we are not aware of any good existing solutions to this problem (other than a few ad hoc heuristics). This seriously limits the applicability of traffic matrices. To bridge the gap between the measurement and the actual application of traffic matrices, we study the critical traffic matrices selection (CritMat) problem in this paper. We developed a mathematical problem formalization after identifying the key requirements and properties of CritMat in the context of network design and analysis. Our complexity analysis showed that CritMat is NP-hard. We then developed several clustering-based approximation algorithms to CritMat. We evaluated these algorithms using a large collection of real traffic matrices collected in AT&T's North American backbone network. Our results demonstrated that these algorithms are very effective and that a small number (e.g., 12) of critical traffic matrices suffice to yield satisfactory performance. Yin Zhang 0001, Zihui Ge |
DSN | 2 |
| 2005 | Optimizing Event Distribution in Publish/Subscribe Systems in the Presence of Policy-Constraints and Composite EventsabstractIn the publish/subscribe paradigm, information is disseminated from publishers to subscribers that are interested in receiving the information. In practice, information dissemination is often restricted by policy constraints due to concerns such as security or confidentiality agreement. Meanwhile, to avoid overwhelming subscribers by the vast amount of primitive information, primitive pieces of information can be combined at so-called brokers in the network, a process called composition. Information composition provides subscribers the desirable ability to express interests in an efficiently selective way. In this paper, we formulate the min-cost event distribution problem in pub/sub systems with policy constraints and information composition. Our goal is to minimize the total cost of event transmission while satisfying policy constraints and enabling information composition. This optimization problem is shown to be NP-complete. Our simulation study shows that our heuristics work efficiently, especially in a policy-constrained system. We also find that by increasing the number of broker nodes in a pub/sub system, we are able to reduce the total cost of event delivery. Weifeng Chen 0001, James F. Kurose, Don Towsley, Zihui Ge |
ICNP | 4 |
| 2005 | Optimal Routing with Multiple Traffic Matrices Tradeoff between Average andWorst Case PerformanceabstractIn this paper, we consider the problem of finding an "efficient" and "robust" set of routes in the face of changing/uncertain traffic. The changes/uncertainty in exogenous traffic is characterized by multiple traffic matrices. Our goal is to find a set of routes that result in good average case performance over the set of traffic matrices, while avoiding bad worst case performance for any single traffic matrix. With multiple traffic matrices, previous work aims solely to optimize the average case performance Chun Zhang, et al., (2005), or the worst case performance David Applegate, et al., (2003). For a given set of traffic matrices, different sets of routes offer a different tradeoff between the average case and the worst case performance. In this paper, we quantify the performance of a routing configuration at both network level and link level. We propose a simple metric-a weighted sum of the average case and the worst case performance-to control the tradeoff between these two considerations. Despite of its simple form, this metric is very effective. We prove that optimizing routing using this metric has desirable properties, such as the average case performance being a decreasing, convex and differentiable function to the worst case performance. By extending previous work Chun Zhang, et al., (2005) Bernard Fortz, et al., (2002), we derive methods to find the optimal routes with respect to the proposed metric for two classes of intra-domain routing protocols: MPLS and OSPF/IS-IS. We evaluate our approach with data collected from an operational tier-I ISP. For MPLS, we find that there exists significant tradeoff (e.g., 15%-23% difference) between optimizing solely on the average case performance and solely on the worst case performance. Our approach can identify solutions that can dramatically improve the worst case performance (13%-15%) while only slightly sacrificing the average case performance (2.2%-3%), in comparison to that by optimizing solely on the average case performance. For OSPF/IS-IS, we still find a significant difference between the two optimization objectives, however, a fine-grained tradeoff is difficult to achieve due to the limited control that OSPF/IS-IS provide. Chun Zhang 0002, James F. Kurose, Don Towsley, Zihui Ge, Yong Liu 0013 |
ICNP | 4 |
| 2005 | Network Anomography
Yin Zhang 0001, Zihui Ge, Albert G. Greenberg, Matthew Roughan |
Internet Measurement Conference | 2 |
| 2004 | Index-server optimization for P2P file sharing in mobile ad hoc networksabstractIn this paper, we compare two basic approaches towards providing peer-to-peer file-sharing (or more generally, information search) in mobile ad-hoc networks (MANET). The flooding approach broadcasts a query (e.g., to locate a node holding a given file) to all network nodes. The index-server approach adds additional servers (known as index servers) that cache directory information about which nodes have which files. With index servers, a node wishing to locate a file first queries its local index server, which then queries other index servers, as needed. The use of index servers presents the possibility of locating a file index quickly in an index server cache, but requires additional overhead to maintain cache consistency. We compare the performance of the flooding approach to two index-server caching approaches: consistent caching and local caching. We quantify the reduction in search overhead using the index-server scheme rather than flooding in MANET, and study how the optimal number of index servers varies according to network size, query rate, and index generation rate. We compare the flooding scheme and the consistent caching and local caching schemes, for two types of queries: history queries and latest queries. Numerical results show how one can choose between the alternatives of consistent caching and local caching depending on network size, index generation rate and query rate. Chikara Ohta, Zihui Ge, Yang Guo 0001, James F. Kurose |
GLOBECOM | 2 |
| 2004 | On Dynamic Subset Difference Revocation Scheme
Weifeng Chen 0001, Zihui Ge, Chun Zhang 0002, James F. Kurose, Don Towsley |
NETWORKING | 2 |
| 2004 | Modeling frame-level errors in GSM wireless channels
Ping Ji 0002, Benyuan Liu, Don Towsley, Zihui Ge, James F. Kurose |
Perform. Evaluation | 4 |
| 2003 | Matchmaker: Signaling for Dynamic Publish/Subscribe ApplicationsabstractThe publish/subscribe (pub/sub) paradigm provides content-oriented data dissemination in which communication channels are established between content publishers and content subscribers based on a matching of subscribers interest in the published content provided - a process we refer to as "matchmaking". Once an interest match has been made, content forwarding state can be installed at intermediate nodes (e.g., active routers, application-level relay nodes) on the path between a content provider and an interested subscriber. In dynamic pub/sub applications, where published content and subscriber interest change frequently the signaling overhead needed to perform matchmaking can be a significant overhead. We first formalize the matchmaking process as an optimization problem, with the goal of minimizing the amount of matchmaking signaling messages. We consider this problem for both shared and per-source multicast data (content) distribution topologies. We characterize the fundamental complexity of the problem, and then describe several efficient solution approaches. The insights gained through our analysis are then embodied in a novel active matchmaker signaling protocol (AMSP). AMSP dynamically adapts to applications' changing publication and subscription requests through a link-marking approach. We simulate AMSP and two existing broadcast-based approaches for conducting matchmaking, and find that AMSP significantly reduces signaling overhead. Zihui Ge, Ping Ji 0002, James F. Kurose, Don Towsley |
ICNP | 1 |
| 2003 | Modeling Peer-Peer File Sharing SystemsabstractPeer-peer networking has recently emerged as a new paradigm for building distributed networked applications. We develop simple mathematical models to explore and illustrate fundamental performance issues of peer-peer file sharing systems. The modeling framework introduced and the corresponding solution methods are flexible enough to accommodate different characteristics of such systems. Through the specification of model parameters, we apply our framework to three different peer-peer architectures: centralized indexing, distributed indexing with flooded queries, and distributed indexing with hashing directed queries. Using our model, we investigate the effects of system scaling, freeloaders, file popularity and availability on system performance. In particular, we observe that a system with distributed indexing and flooded queries cannot exploit the full capacity of peer-peer systems. We further show that peer-peer file sharing systems can tolerate a significant number of freeloaders without suffering much performance degradation. In many cases, freeloaders can benefit from the available spare capacity of peer-peer systems and increase overall system throughput. Our work shows that simple models coupled with efficient solution methods can be used to understand and answer questions related to the performance of peer-peer file sharing systems. Zihui Ge, Daniel R. Figueiredo 0001, Sharad Jaiswal, James F. Kurose, Don Towsley |
INFOCOM | 1 |
| 2003 | A comparison of hard-state and soft-state signaling protocolsabstractOne of the key infrastructure components in all telecommunication networks, ranging from the telephone network, to VC-oriented data networks, to the Internet, is its signaling system. Two broad approaches towards signaling can be identified: so-called hard-state and soft-state approaches. Despite the fundamental importance of signaling, our understanding of these approaches - their pros and cons and the circumstances in which they might best be employed - is mostly anecdotal (and occasionally religious). In this paper, we compare and contrast a variety of signaling approaches ranging from a "pure" soft state, to soft-state approaches augmented with explicit state removal and/or reliable signaling, to a "pure" hard state approach. We develop an analytic model that allows us to quantify state inconsistency in single- and multiple-hop signaling scenarios, and the "cost" (both in terms of signaling overhead, and application-specific costs resulting from state inconsistency) associated with a given signaling approach and its parameters (e.g., state refresh and removal timers). Among the class of soft-state approaches, we find that a soft-state approach coupled with explicit removal substantially improves the degree of state consistency while introducing little additional signaling message overhead. The addition of reliable explicit setup/update/removal allows the soft-state approach to achieve comparable (and sometimes better) consistency than that of the hard-state approach. Ping Ji 0002, Zihui Ge, James F. Kurose, Don Towsley |
SIGCOMM | 2 |
| 2002 | A Demand Adaptive and Locality Aware (DALA) streaming media server cluster architectureabstractThe wide availability of broadband networking technologies such as cable modems and DSL coupled with the growing popularity of the Internet has led to a dramatic increase in the availability and the use of online streaming media. With the "last mile" network bandwidth no longer a constraint, the bottleneck for video streaming has been pushed closer to the server. Streaming high quality audio and video to a myriad of clients imposes significant resource demands on the server. In this work, we propose a demand adaptive and locality aware (DALA) clustered media server architecture that can dynamically allocate resources to adapt to changing demand and also maximize the number of clients serviced by the server cluster. Moreover, our design exploits temporal locality among requests by dispatching newly arriving requests to servers that are already servicing prior requests for those objects, thereby extracting the benefits of locality. We explore the efficacy of the DALA clustered architecture using simulations. Our simulation results show that DALA is highly adaptive, exhibits significant performance gains when compared to static schemes, and has a low system overhead. Our results demonstrate that DALA is a simple, yet effective approach for designing clustered media servers. Zihui Ge, Ping Ji 0002, Prashant J. Shenoy |
NOSSDAV | 1 |
| 2001 | Channelization Problem in Large Scale Data DisseminationabstractIn many large scale data dissemination systems, a large number of information flows must be delivered to a large number of information receivers. However, because of differences in interests among receivers, not all receivers are interested in all of the information flows. Multicasting provides the opportunity to deliver a subset of the information flows to a subset of the receivers. With a limited number of multicast groups available, the channelization problem is to find an optimal mapping of information flows to a fixed number of multicast groups, and a subscription mapping of receivers to multicast groups so as to minimize a function of the total bandwidth consumed and the amount of unwanted information received by receivers. We formally define two versions of the channelization problem and subscription problem (a subcomponent of the channelization problem). We analyze the complexity of each version of the channelization problem and show that they are both NP-complete. We also find that the subscription problem is NP-complete when one flow can be assigned to multiple multicast groups. We also study and compare different approximation algorithms to solve the channelization problem, finding that one particular heuristic, flow-based-merge, finds good solutions over a range of problem configurations. Micah Adler, Zihui Ge, James F. Kurose, Don Towsley, Steve Zabele |
ICNP | 2 |