Zhihan Fang

dblp:162/3674 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-1882-6252ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Cellular Infrastructure Sharing for Network Robustness: A Citywide Empirical Study
abstract
Individual cellular networks have been very robust to random cell tower failures due to redundant cell tower deployments. However, a large-scale clustered failure with multiple cell towers can lead to the loss of services of a cellular network. Recently, off-the-shelf smartphones can support multiple network standards, so cellular network infrastructure sharing is a promising direction to improve the service robustness under potential large-scale clustered tower failures. The existing work on cellular network robustness is usually limited to large-scale studies of individual networks or small-scale studies of multiple networks. In this work, we conduct the first investigation, to our knowledge, on cross-network infrastructure sharing benefits for enhancing robustness with afull cellular penetration rate. Our work is based on all cellular networks in Shenzhen, China, covering over 10 million cellular users. Specifically, we design a new metric to quantify cellular network robustness with or without cross-network sharing under both random and clustered cell tower failures. We further study the impact of different factors on robustness, including the number of networks, spatiotemporal dynamics, contextual factors, and a case study at two key transportation hubs. We provide a set of lessons learned based on our study, along with discussions of the results.
Zhihan Fang, Guang Yang 0028, Wenjun Lyu, Zhiqing Hong, Shuxin Zhong, Weijian Zuo, Yuelei Xie, Yu Yang 0010, Guang Wang 0001, Yunhuai Liu, Desheng Zhang 0002
IEEE Trans. Mob. Comput.1
2024 Improving Network Robustness via Cellular Infrastructure Sharing: An Empirical Study of Infrastructure Failure with All Cellular Operators in a City
abstract
Individual cellular networks have been very robust to random cell tower failure due to redundant cell tower deployments. However, a large-scale clustered failure (e.g., due to fiber cut or cyber attacks) with multiple cell towers can lead to the loss of services of a cellular network. Recently, off-the-shelf smartphones can support multiple network standards, so cellular network infrastructure sharing is a promising direction to improve the service robustness under potential large-scale clustered cell tower failure. The existing work on cellular network robustness is usually limited to large-scale studies of individual networks or small-scale studies of multiple networks. In this work, we conduct the first investigation, to our knowledge, into the benefits of cross-network infrastructure sharing for enhancing robustness at a full cellular penetration rate. We design a new metric to quantify cellular network robustness with or without cross-network sharing under both random and clustered cell tower failures. We further study the impact of spatial dynamics on cellular network robustness.
Zhihan Fang, Guang Yang 0028, Wenjun Lyu, Zhiqing Hong, Shuxin Zhong, Weijian Zuo, Yu Yang 0010, Guang Wang 0001, Desheng Zhang 0002
SIGSPATIAL/GIS1
2022 VeMo: Enable Transparent Vehicular Mobility Modeling at Individual Levels With Full Penetration
abstract
Understanding and predicting real-time vehicle mobility patterns on highways are essential to address traffic congestion and respond to the emergency. However, almost all existing works (e.g., based on cellphones, onboard devices, or traffic cameras) suffer from high costs, low penetration rates, or only aggregate results. To address these drawbacks, we utilize Electric Toll Collection systems (ETC) as a large-scale sensor network and design a system called VeMo to transparently model and predict vehicle mobility at the individual level with a full penetration rate. Our novelty is how we address uncertainty issues (i.e., unknown routes and speeds) due to sparse implicit ETC data based on a key data-driven insight, i.e., individual driving behaviors are strongly correlated with crowds of drivers under certain spatiotemporal contexts and can be predicted by combining both personal habits and context information. We evaluate VeMo with (i) a large-scale ETC system with tracking devices at 773 highway entrances and exits capturing more than 2 million vehicles every day; (ii) a fleet consisting of 114 thousand vehicles with GPS data as ground truth. Compared with state-of-the-art benchmark mobility models, the experimental results show that VeMo outperforms them by 10 percent on average.
Yu Yang 0010, Xiaoyang Xie, Zhihan Fang, Fan Zhang 0019, Yang Wang 0015, Desheng Zhang 0002
IEEE Trans. Mob. Comput.3
2021 MoCha: Large-Scale Driving Pattern Characterization for Usage-based Insurance
abstract
Given widely adopted vehicle tracking technologies, usage-based insurance has been a rising market over the past few years. With potential discounts from insurance companies, customers voluntarily install sensing devices in their vehicles for insurance companies, which are utilized to analyze their historical driving patterns to derive the risks of future driving. However, it is challenging to characterize and predict driving patterns, especially for new users with limited data. To address this issue, we propose and evaluate a system called MoCha to accurately characterize driving patterns for usage-based insurance. The key question we aim to explore with MoCha is whether we can fully explore long-term driving patterns of new users with only limited historical data of themselves by leveraging abundant data of other users and contextual information. To answer this question, we design (i) a multi-level driving pattern modeling component to capture the spatial-temporal dependency on both individual and group level, and (ii) a multi-task learning method to utilize underlying relations of driving metrics and predict multiple driving metrics simultaneously. We implement and evaluate MoCha with real-world on-board diagnostics data from a large insurance company with more than 340,000 vehicles. Further, we validate the usefulness of MoCha by predicting driving risks based on real-world claim data in a Chinese city, Shenzhen.
Zhihan Fang, Guang Yang 0028, Dian Zhang 0001, Xiaoyang Xie, Guang Wang 0001, Yu Yang 0010, Fan Zhang 0019, Desheng Zhang 0002
KDD1
2021 Pricing-aware Real-time Charging Scheduling and Charging Station Expansion for Large-scale Electric Buses
abstract
We are witnessing a rapid growth of electrified vehicles due to the ever-increasing concerns on urban air quality and energy security. Compared to other types of electric vehicles, electric buses have not yet been prevailingly adopted worldwide due to their high owning and operating costs, long charging time, and the uneven spatial distribution of charging facilities. Moreover, the highly dynamic environment factors such as unpredictable traffic congestion, different passenger demands, and even the changing weather can significantly affect electric bus charging efficiency and potentially hinder the further promotion of large-scale electric bus fleets. To address these issues, in this article, we first analyze a real-world dataset including massive data from 16,359 electric buses, 1,400 bus lines, and 5,562 bus stops. Then, we investigate the electric bus network to understand its operating and charging patterns, and further verify the necessity and feasibility of a real-time charging scheduling. With such understanding, we design busCharging , a pricing-aware real-time charging scheduling system based on Markov Decision Process to reduce the overall charging and operating costs for city-scale electric bus fleets, taking the time-variant electricity pricing into account. To show the effectiveness of busCharging , we implement it with the real-world data from Shenzhen, which includes GPS data of electric buses, the metadata of all bus lines and bus stops, combined with data of 376 charging stations for electric buses. The evaluation results show that busCharging dramatically reduces the charging cost by 23.7% and 12.8% of electricity usage simultaneously. Finally, we design a scheduling-based charging station expansion strategy to verify our busCharging is also effective during the charging station expansion process.
Guang Wang 0001, Zhihan Fang, Xiaoyang Xie, Shuai Wang 0008, Huijun Sun, Fan Zhang 0019, Yunhuai Liu, Desheng Zhang 0002
ACM Trans. Intell. Syst. Technol.2
2021 A Measurement Framework for Explicit and Implicit Urban Traffic Sensing
abstract
Urban traffic sensing has been investigated extensively by different real-time sensing approaches due to important applications such as navigation and emergency services. Basically, the existing traffic sensing approaches can be classified into two categories by sensing natures, i.e., explicit and implicit sensing. In this article, we design a measurement framework called EXIMIUS for a large-scale data-driven study to investigate the strengths and weaknesses of two sensing approaches by using two particular systems for traffic sensing as concrete examples. In our investigation, we utilize TB-level data from two systems: (i) GPS data from five thousand vehicles, (ii) signaling data from three million cellphone users, from the Chinese city Hefei. Our study adopts a widely used concept called crowdedness level to rigorously explore the impacts of contexts on traffic conditions including population density, region functions, road categories, rush hours, holidays, weather, and so on, based on various context data. We quantify the strengths and weaknesses of these two sensing approaches in different scenarios and then we explore the possibility of unifying two sensing approaches for better performance by using a truth discovery-based data fusion scheme. Our results provide a few valuable insights for urban sensing based on explicit and implicit data from transportation and telecommunication domains.
Zhou Qin 0001, Zhihan Fang, Yunhuai Liu, Desheng Zhang 0002
ACM Trans. Sens. Networks2
2020 CellRep: Usage Representativeness Modeling and Correction Based on Multiple City-Scale Cellular Networks
abstract
Understanding representativeness in cellular web logs at city scale is essential for web applications. Most of the existing work on cellular web analyses or applications is built upon data from a single network in a city, which may not be representative of the overall usage patterns since multiple cellular networks coexist in most cities in the world. In this paper, we conduct the first comprehensive investigation of multiple cellular networks in a city with a 100% user penetration rate. We study web usage pattern (e.g., internet access services) correlation and difference between diverse cellular networks in terms of spatial and temporal dimensions to quantify the representativeness of web usage from a single network in usage patterns of all users in the same city. Moreover, relying on three external datasets, we study the correlation between the representativeness and contextual factors (e.g., Point-of-Interest, population, and mobility) to explain the potential causalities for the representativeness difference. We found that contextual diversity is a key reason for representativeness difference, and representativeness has a significant impact on the performance of real-world applications. Based on the analysis results, we further design a correction model to address the bias of single cellphone networks and improve representativeness by 45.8%.
Zhihan Fang, Guang Wang 0001, Shuai Wang 0008, Chaoji Zuo, Fan Zhang 0019, Desheng Zhang 0002
WWW1
2019 VeMo: Enabling Transparent Vehicular Mobility Modeling at Individual Levels with Full Penetration
abstract
Understanding and predicting real-time vehicle mobility patterns on highways are essential to address traffic congestion and respond to the emergency. However, almost all existing works (e.g., based on cellphones, onboard devices, or traffic cameras) suffer from high costs, low penetration rates, or only aggregate results. To address these drawbacks, we utilize Electric Toll Collection systems (ETC) as a large-scale sensor network and design a system called VeMo to transparently model and predict vehicle mobility at the individual level with a full penetration rate. Our novelty is how we address uncertainty issues (i.e., unknown routes and speeds) due to sparse implicit ETC data based on a key data-driven insight, i.e., individual driving behaviors are strongly correlated with crowds of drivers under certain spatiotemporal contexts and can be predicted by combining both personal habits and context information. More importantly, we evaluate VeMo with (i) a large-scale ETC system with tracking devices at 773 highway entrances and exits capturing more than 2 million vehicles every day; (ii) a fleet consisting of 114 thousand vehicles with GPS data as ground truth. We compared VeMo with state-of-the-art benchmark mobility models, and the experimental results show that VeMo outperforms them by average 10% in terms of accuracy.
Yu Yang 0010, Xiaoyang Xie, Zhihan Fang, Fan Zhang 0019, Yang Wang 0015, Desheng Zhang 0002
MobiCom3
2019 Fine-grained travel time sensing in heterogeneous mobile networks: poster abstract
abstract
Sensing and estimating real-time passenger travel time in urban mobile networks provides essential information for passengers to select different transportation modalities, improving their travel experience. Most existing work on travel time sensing has been focused on individual transportation modalities and its riding time based on the static time tables. However, passengers often consider different modes of transportation, e.g. taxis, subways, buses or personal vehicles, and a significant portion of the travel time is spent in the uncertain waiting or walking, which cannot be accurately obtained by static time tables. In this paper, we design a real-time data-driven framework FineTravel for fine-grained travel time sensing based on multi-source real-time data from multi-modal mobile networks including taxi, bus, subway, and private vehicle. The key challenge we address in FineTravel is to estimate implicit components (including walking, waiting and riding time) of the total travel time without direct measurement. In contrast to the existing work based on single-modal transportation networks mostly with offline data, the novelty of the FineTravel is based on its real-time multi-source data-driven (including both vehicle GPS and smartcard data) modeling for a comprehensive travel time sensing.
Zhihan Fang, Fan Zhang 0019, Desheng Zhang 0002
SenSys1
2018 EXIMIUS: A Measurement Framework for Explicit and Implicit Urban Traffic Sensing
abstract
Urban traffic sensing has been investigated extensively by different real-time sensing approaches due to important applications such as navigation and emergency services. Basically, the existing traffic sensing approaches can be classified into two categories, i.e., explicit and implicit sensing. In this paper, we design a measurement framework called EXIMIUS for a large-scale data-driven study to investigate the strengths and weaknesses of these two sensing approaches by using two particular systems for traffic sensing as concrete examples, i.e., a vehicular system as a crowdsourcing-based explicit sensing and a cellular system as an infrastructure-based implicit sensing. In our investigation, we utilize TB-level data from two systems: (i) vehicle GPS data from 3 thousand private cars and 2 thousand commercial vehicles, (ii) cellular signaling data from 3 million cellphone users, from the Chinese city Hefei. Our study adopts a widely-used concept called crowdedness level to rigorously explore the impacts of various spatiotemporal contexts on real-time traffic conditions including population density, region functions, road categories, rush hours, etc. based on a wide range of context data. We quantify the strengths and weaknesses of these two sensing approaches in different scenarios then we explore the possibility of unifying these two sensing approaches for better performance. Our results provide a few valuable insights for urban sensing based on explicit and implicit data from transportation and telecommunication domains.
Zhou Qin 0001, Zhihan Fang, Yunhuai Liu, Desheng Zhang 0002
SenSys2
2017 Prioritizing manual test cases in rapid release environments
abstract
Summary Test case prioritization is an important testing activity, in practice, specially for large scale systems. The goal is to rank the existing test cases in a way that they detect faults as soon as possible, so that any partial execution of the test suite detects the maximum number of defects for the given budget. Test prioritization becomes even more important when the test execution is time consuming, for example, manual system tests versus automated unit tests. Most existing test case prioritization techniques are based on code coverage, which requires access to source code. However, manual testing is mainly performed in a black‐box manner (manual testers do not have access to the source code). Therefore, in this paper, the existing test case prioritization techniques (e.g. diversity‐based and history‐based techniques) are examined and modified to be applicable on manual black‐box system testing. An empirical study on four older releases of desktop Firefox showed that none of the techniques were strongly dominating the others in all releases. However, when nine more recent releases of desktop Firefox, where the development has been moved from a traditional to a more agile and rapid release environment, were studied, a very significant difference between the history‐based approach and its alternatives was observed. The higher effectiveness of the history‐based approach compared with alternatives also held on 28 additional rapid releases of other Firefox projects – mobile Firefox and tablet Firefox. The conclusion of the paper is that test cases in rapid release environments can be very effectively prioritized for execution, based on their historical failure knowledge. In particular, it is the recency of historical knowledge that explains its effectiveness in rapid release environments rather than other changes in the process. Copyright © 2016 John Wiley & Sons, Ltd.
Hadi Hemmati, Zhihan Fang, Mika Mäntylä, Bram Adams
Softw. Test. Verification Reliab.2
2015 Prioritizing Manual Test Cases in Traditional and Rapid Release Environments
abstract
Test case prioritization is one of the most practically useful activities in testing, specially for large scale systems. The goal is ranking the existing test cases in a way that they detect faults as soon as possible, so that any partial execution of the test suite detects maximum number of defects for the given budget. Test prioritization becomes even more important when the test execution is time consuming, e.g., manual system tests vs. automated unit tests. Most existing test case prioritization techniques are based on code coverage, which requires access to source code. However, manual testing is mainly done in a black- box manner (manual testers do not have access to the source code). Therefore, in this paper, we first examine the existing test case prioritization techniques and modify them to be applicable on manual black-box system testing. We specifically study a coverage- based, a diversity-based, and a risk driven approach for test case prioritization. Our empirical study on four older releases of Mozilla Firefox shows that none of the techniques are strongly dominating the others in all releases. However, when we study nine more recent releases of Firefox, where the development has been moved from a traditional to a more agile and rapid release environment, we see a very signifiant difference (on average 65% effectiveness improvement) between the risk-driven approach and its alternatives. Our conclusion, based on one case study of 13 releases of an industrial system, is that test suites in rapid release environments, potentially, can be very effectively prioritized for execution, based on their historical riskiness; whereas the same conclusions do not hold in the traditional software development environments.
Hadi Hemmati, Zhihan Fang, Mika Mäntylä
ICST2