VLDB 2026 Research / reviewers in the wild / expert
Juanjuan Zhao 0001
dblp:34/10239-1
· DBLP profile ↗
46ranked-venue papers
6as first author
31since 2021 · last 2025
0000-0003-1002-9272ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 13 · 10 since 2021Artificial intelligence and machine learning · 12 · 7 since 2021Computer networks · 9 · 4 since 2021Systems, architecture and hardware · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Incremental Tucker Decomposition for Scalable and Efficient Multilingual Speech RecognitionabstractMultilingual Automatic Speech Recognition (ASR) systems face challenges in efficiently integrating new languages while maintaining scalability and performance. This paper introduces an Incremental Tucker-based Low-Rank Adaptation (LoRA) framework designed to address these limitations. By leveraging Tucker decomposition to compress multilingual LoRA parameters and introducing an efficient incremental learning mechanism, the framework enables seamless addition of new languages without re-computing the full decomposition. Experiments on the Mozilla Common Voice dataset demonstrate that the proposed approach reduces parameter requirements by up to 80% compared to traditional LoRA-based methods, while achieving competitive Word Error Rates (WER) across diverse languages. This framework provides a scalable and resource-efficient solution for modern multilingual ASR systems, offering significant advantages in memory efficiency and adaptability to dynamic language expansion. Guanghui Song, Ye Hong, Juanjuan Zhao 0001, Kejiang Ye |
IJCNN | 4 |
| 2025 | On the Adversarial Robustness of Visual-Language Chat ModelsabstractWith the rapid development of large language models (LLMs), there has been a strong interest in integrating other modalities such as image comprehension capabilities. While they have shown impressive performance in various multimodal tasks, the robustness of Visual Language Models (VLMs) has not been thoroughly investigated. We mainly focus on the robustness of VLMs on visual adversarial examples. In this work, we explore the capability of adversarial examples targeting VLMs. We highlight that the multimodal nature of VLMs presents a unique attack surface to manipulate the outputs of the LLMs, and the continuous nature of visual inputs further enhances the effectiveness of adversarial attacks against language generative models. Furthermore, we demonstrate three application scenarios for adversarial examples targeting VLMs: image description, jailbreaking, and information hiding. We conduct experiments on several leading open-source VLMs and demonstrate the successful application of adversarial examples in all the proposed scenarios. We hope that our findings would enable the development of multimodal models more robust to adversarial attacks. Our code is available at https://github.com/lafeat/m3-break. Tianrui Qin, Xuan Wang 0029, Juanjuan Zhao 0001, Kejiang Ye, Cheng-Zhong Xu 0001 |
ICMR | 3 |
| 2025 | SDF-Guided Multi-modal Big Data Road Extraction
Juanjuan Zhao 0001, Kejiang Ye |
PAKDD (7) | 2 |
| 2025 | Offline Map Matching Based on Localization Error Distribution Modeling
Ruilin Xu 0009, Kaijie Li, Kejiang Ye, Fan Zhang 0019, Juanjuan Zhao 0001 |
PAKDD (6) | 7 |
| 2025 | An Online Map Matching Algorithm for Path-Free Trajectories by Integrating Path-Constrained TrajectoriesabstractOnline map matching (MM) aligns real-time GPS trajectories with digital road networks, playing a vital role in vehicle navigation, route planning, and traffic analysis. Hidden Markov Models (HMMs) are widely used for their interpretability and ability to handle low GPS sampling rates. However, in urban scenarios characterized by complex road networks, significant GPS localization error, and dynamic traffic conditions, existing HMM-based methods face challenges such as large road search spaces due to uniform GPS localization error distributions (GLED) and inaccurate route accessibility estimates stemming from inadequate consideration of real-time traffic conditions. This paper proposes an improved HMM-based MM method, recognizing that urban vehicle trajectories can be categorized into two types: path-free (e.g., taxis, private cars) and path-constrained (e.g., buses). Analyzing path-constrained trajectories helps estimate fine-grained GLED and real-time traffic states of path-free vehicles more precisely. The novelty of our approach lies in two aspects: i) Using a hierarchical spectral clustering algorithm based on GPS localization errors of path-constrained bus trajectories, a city is divided into fine-grained sub-regions with consistent GLED. This enables the HMM an adaptive road search scopes, improving online MM efficiency. ii) Gradient boosting trees, known for their interpretability, estimate free-flow speeds by integrating path-constrained trajectories with the factors like road attributes and time, optimizing HMM state transition probabilities for path-free trajectory MM. Experiments on real-world data demonstrate that the optimized HMM methods, leveraging different trajectory types, significantly enhance MM efficiency and accuracy compared to baseline models.The codebase of our methods and datasets are available at https://github.com/jacklee018/onlineMM-IPCT. Kaijie Li, Juanjuan Zhao 0001, Kejiang Ye, Cheng-Zhong Xu 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Urban Transport Mode Split Prediction: A Hybrid Deep Learning Framework Considering Spatiotemporal DependencyabstractTransport Mode Split (TMS) represents the distribution of trips among transport modes between city regions. Accurate TMS prediction is crucial for urban planning and traffic management. Traditional methods, such as curve models and discrete choice models, often fail to capture user travel mode pReferences due to dataset limitations and inadequate spatiotemporal modeling. While deep learning models have been applied in related domains such as ride-sharing and traffic prediction, they typically focus on a single transport mode and do not model the competitive relationships between modes. We propose a novel Deep Learning-based Framework for Transport Mode Split Prediction (DMSP), integrating restricted-route (bus/subway) and free-route (taxi) data. It introduces two key innovations: 1) efficient extraction of TMS samples from sparse raw data, enriched with spatiotemporal context, and 2) a prediction model extending discrete choice theory via maximum utility principles. Mode utilities are estimated using CNNs for travel mode-specific effects and GNN for spatiotemporal dependencies. Experiments on six months of Shenzhen data show that DMSP reduces MAPE by over 4% compared to existing methods. Juanjuan Zhao 0001, Furong Zheng, Fan Zhang 0019, Kejiang Ye, Cheng-Zhong Xu 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | RCDP: A Privacy-Preserving Approach for Synthesizing Realistic Commuting DataabstractPublishing commuting trajectory data, including information of home and workplace locations, commuting distances and working hours, provides valuable insights for urban transportation planning. However, the data also contains sensitive personal information, raising privacy concerns even after the removal of unique identifiers. While traditional privacy-preserving methods, such as k-anonymity and differential privacy, have been widely applied, they mainly focus on single-trip travel patterns (e.g., point sequences or paths) and fail to capture the unique characteristics of commuting behavior. In this paper, we propose RCDP, a novel differential privacy-based model for synthesizing realistic commuting data using a prefix tree structure. RCDP introduces two key innovations: (1) it models round-trip commuting patterns through adaptive spatio-temporal generalization and a prefix tree, ensuring that the synthesized data retains the key commuting characteristics; (2) it employs a hierarchical privacy budget allocation mechanism that dynamically adjusts the budget across tree levels, along with a distribution-based node insertion method to maintain tree consistency, effectively balancing privacy and utility. Validation using public transport smart card data in Shenzhen, China demonstrates that RCDP outperforms existing k-anonymity and differential privacy approaches in preserving essential commuting features while ensuring strong privacy protection. Juanjuan Zhao 0001, Kejiang Ye |
IEEE Big Data | 2 |
| 2024 | Fine-Grained Geo-Obfuscation to Protect Workers' Location Privacy in Time-Sensitive Spatial Crowdsourcing
Chenxi Qiu, Yuede Ji, Anna Cinzia Squicciarini, Ram Dantu, Juanjuan Zhao 0001, Cheng-Zhong Xu 0001 |
EDBT | 6 |
| 2024 | GCompletor: A Graph-Based Deep Learning Method for Traffic State Imputation on Urban Road Networks
Kaijie Li, Juanjuan Zhao 0001, Li Yan 0004, Ye Li 0002, Kejiang Ye |
ICPR (6) | 2 |
| 2024 | MPRG: A Method for Parallel Road Generation Based on Trajectories of Multiple Types of Vehicles
Bingru Han, Juanjuan Zhao 0001, Kejiang Ye, Fan Zhang 0019 |
PAKDD (5) | 2 |
| 2024 | Enhanced HMM Map Matching Model Based on Multiple Type Trajectories
Juanjuan Zhao 0001, Fan Zhang 0019, Kejiang Ye |
PAKDD (5) | 2 |
| 2024 | GSPM: An Early Detection Approach to Sudden Abnormal Large Outflow in a Metro System
Juanjuan Zhao 0001, Fan Zhang 0019, Kejiang Ye |
PAKDD (5) | 2 |
| 2024 | FMSYS: Fine-Grained Passenger Flow Monitoring in a Large-Scale Metro System Based on AFC Smart Card Data
Juanjuan Zhao 0001, Fan Zhang 0019, Kejiang Ye |
PAKDD (5) | 2 |
| 2024 | STRmt: A state transition based model for real-time crowd counting in a metro systemabstractSummary Real‐time estimation of crowd counting in underground metro systems, constrained by limited space, is crucial for managing heightened pedestrian volumes and responding promptly to emergencies. To address this challenge, we propose a passenger state transition‐based model, called STRmt, designed for the seamless and continuous monitoring of real‐time crowd movement within service areas of stations and trains, leveraging auto fare collection systems (AFC) as a comprehensive sensor network. Our innovation lies in modeling the dynamic movement of passengers within a metro system over time as a state transition process aligned with the train schedule. To achieve this, we introduce a spatio‐temporal deep learning framework, denoted as STnet, designed to dynamically predict these state transitions. The performance of our method is rigorously assessed through extensive experiments conducted spanning 2 years in Shenzhen, China, utilizing AFC data, train schedule data, and weather data. The results demonstrate that the proposed method surpasses baseline methods, achieving an estimation precision of 0.92. Juanjuan Zhao 0001, Jun Zhang 0014, Fan Zhang 0019, Kejiang Ye, Cheng-Zhong Xu 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2024 | A Heterogeneous Graph Convolution Based Method for Short-Term OD Flow Completion and Prediction in a Metro SystemabstractShort-term OD flow (i.e. the number of passenger traveling between stations) prediction is crucial to traffic management in metro systems. The delayed effect in latest complete OD flow collection and complex spatiotemporal correlations of OD flows in high dimension make it challengeable to predict short-term OD flow. Existing methods need to be improved due to not fully utilizing the real-time passenger mobility data and not sufficiently modeling the implicit correlation of the mobility patterns between stations. In this paper, we propose a Completion based Adaptive Heterogeneous Graph Convolution Spatiotemporal Predictor. The novelty is mainly reflected in two aspects. The first is to model real-time mobility evolution by establishing the implicit correlation between observed OD flows and the prediction target OD flows in high dimension based on a key data-driven insight: the destination distributions of the passengers departing from a station are correlated with other stations sharing similar attributes (e.g. geographical location, region function). The second is to complete the latest incomplete OD flows by estimating the destination distribution of unfinished trips through considering the real-time mobility evolution and the time cost between stations, which is the base of time series prediction and can improve the model’s dynamic adaptability. Extensive experiments on two real world metro datasets demonstrate the superiority of our model over other competitors with the biggest model performance improvement being nearly 4%. In addition, the data complete framework we propose can be integrated into other models to improve their performance up to 2.1%. Jiexia Ye, Juanjuan Zhao 0001, Furong Zheng, Cheng-Zhong Xu 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | AdvDiffuser: Natural Adversarial Example Synthesis with Diffusion ModelsabstractPrevious work on adversarial examples typically involves a fixed norm perturbation budget, which fails to capture the way humans perceive perturbations. Recent work has shifted towards natural unrestricted adversarial examples (UAEs) that breaks ℓpperturbation bounds but nonetheless remain semantically plausible. Current methods use GAN or VAE to generate UAEs by perturbing latent codes. However, this leads to loss of high-level information, resulting in low-quality and unnatural UAEs. In light of this, we propose AdvDiffuser, a new method for synthesizing natural UAEs using diffusion models. It can generate UAEs from scratch or conditionally based on reference images. To generate natural UAEs, we perturb predicted images to steer their latent code towards the adversarial sample space of a particular classifier. We also propose adversarial inpainting based on class activation mapping to retain the salient regions of the image while perturbing less important areas. On CIFAR-10, CelebA and ImageNet, we demonstrate that it can defeat the most robust models on the RobustBench leaderboard with near 100% success rates. Furthermore, The synthesized UAEs are not only more natural but also stronger compared to the current state-of-the-art attacks. Specifically, compared with GA-attack, the UAEs generated with AdvDiffuser exhibit 6× smaller LPIPS perturbations, 2 ~ 3× smaller FID scores and 0.28 higher in SSIM metrics, making them perceptually stealthier. Finally, adversarial training with AdvDiffuser further improves the model robustness against attacks with unseen threat models.1 Juanjuan Zhao 0001, Kejiang Ye, Cheng-Zhong Xu 0001 |
ICCV | 3 |
| 2023 | Destruction-Restoration Suppresses Data Protection Perturbations against Diffusion ModelsabstractDiffusion models have become popular in computer vision applications due to their ability to generate high-quality images quickly, with some models achieving a high degree of realism. However, the use of these models for novel image manipulation and generation poses significant risks of being used for harmful purposes, such as the spread of misinformation or copyright infringement. Recent research has proposed techniques that involve adding imperceptible perturbations to images to mitigate privacy and copyright concerns. However, trivial image processing methods can largely destroy these protective perturbations. In this paper, we propose a new method called DiffCleaner, a destruction-restoration pipeline, that combines the advantages of image corruption and restoration methods. It can not only remove protective perturbations but also further preserve high quality generation, as well as fidelity of the original features in generated images. DiffCleaner produces specialized images with visualization and quantitative metrics closely matching unprotected results. We pitch a range of circumventing methods vs. two state-of-the-art data protection methods against diffusion models, and show that both are fragile and offer inadequate data protection. By providing a benchmark comprising a range of protective perturbations removal methods, we hope this paper will inspire and facilitate new research towards developing more resilient methods for data protection. Tianrui Qin, Juanjuan Zhao 0001, Kejiang Ye |
ICTAI | 3 |
| 2023 | Practical model with strong interpretability and predictability: An explanatory model for individuals' destination prediction considering personal and crowd travel behaviorabstractAbstract Real‐time individuals' destination prediction is of great significance for real‐time user tracking, service recommendation and other related applications. Traditional technology mainly used statistical methods based on the travel patterns mined from personal history travel data. However, it is not clear how to predict the destinations of individuals with only limited personal historical data. In this paper, taking the public transportation metro systems as example, we design a practical method called practical model with strong interpretability and predictability to predict each passenger's destination. Our main novelties are two aspects: (1) We propose to predict individuals' destination by combining personal and crowd behavior under certain context. (2) An explanatory model combining discrete choice model and neural network model is proposed to predict individuals' stochastic trip's destination, which can be applied to other transportation analysis scenarios about individuals' choice behavior such as travel mode choice or route choice. We validate our method based on extensive experiments, using smart card data collected by automatic fare collection system and weather data in Shenzhen, China. The experimental results demonstrate that our approach can achieve better performance than other baselines in terms of prediction accuracy. Juanjuan Zhao 0001, Jiexia Ye, Minxian Xu, Cheng-Zhong Xu 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2023 | Completion and augmentation-based spatiotemporal deep learning approach for short-term metro origin-destination matrix prediction under limited observable data
Jiexia Ye, Juanjuan Zhao 0001, Furong Zheng, Cheng-Zhong Xu 0001 |
Neural Comput. Appl. | 2 |
| 2023 | Metro OD Matrix Prediction Based on Multi-View Passenger Flow Evolution Trend ModelingabstractShort-term Origin-Destination(OD) matrix prediction in metro systems aims to predict the number of passenger demands from one station to another during a short time period. That is crucial for dynamic traffic operations, e.g., route recommendation, metro scheduling. However, existing methods need further improvement due to that they fail to take full use of the real-time traffic information and model the complex spatiotemporal correlation of traffic flows. In this paper, aMulti-ViewPassengerFlow (MVPF) evolution trend based OD matrix prediction method is proposed. It consists of two components focusing on individual station and cross-station learning. Specifically, the individual station level part uses Gate Recurrent Unit and Extended Graph Attention Networks combined model to learn the high-level spatiotemporal-dependent representation of each station as the roles of origin and destination respectively, by considering multiple views of real-time traffic information (i.e., Inflow, destination allocation of Inflow, Outflow, origin allocations of Outflow). The cross-station part aims to learn passenger mobility pattern from each origin to destination through defining a transition matrix under spatiotemporal context. Compared with state-of-the-art solutions, MVPF increases the OD prediction performance metric of WMAPE by 2.5% on average. The experimental results demonstrate the superiority of MVPF against other competitors. The source code is available athttps://github.com/zfrInSIAT/MVPF-code. Furong Zheng, Juanjuan Zhao 0001, Jiexia Ye, Kejiang Ye, Cheng-Zhong Xu 0001 |
IEEE Trans. Big Data | 2 |
| 2023 | MobiCharger: Optimal Scheduling for Cooperative EV-to-EV Dynamic Wireless ChargingabstractWith the advancement of dynamic wireless charging for Electric Vehicles (EVs), Mobile Energy Disseminator (MED), which can charge an EV in motion, becomes available. However, existing wireless charging scheduling methods for wireless sensors, which are the most related works to MED deployment, are not directly applicable for city-scale EV-to-EV dynamic wireless charging. We presentMobiCharger:aMobile wirelessChargerguidance system that determines the number of serving MEDs, and their optimal routes. We studied a metropolitan-scale vehicle mobility dataset, and found: most vehicles have routines, and the number of driving EVs changes over time, which means MED deployment should adaptively change as well. We combine EVs' current trajectories and routines to estimate EV density and the cruising graph for MED coverage. Then, we develop an offline MED deployment method that utilizes multi-objective optimization to determine the number of serving MEDs and the driving route of each MED, and an online method that utilizes Reinforcement Learning to adjust the MED deployment when the real-time vehicle traffic changes. Our trace-driven experiments show that compared with previous methods,MobiChargerincreases the medium State-of-Charge of all EVs by 50% during all time slots, and the number of charges of EVs by almost 100%. Li Yan 0004, Haiying Shen, Liuwang Kang, Juanjuan Zhao 0001, Zhe Zhang 0048, Cheng-Zhong Xu 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2022 | TrafficAdaptor: an adaptive obfuscation strategy for vehicle location privacy against traffic flow aware attacksabstractOne of the most popular location privacy-preserving mechanisms applied in location-based services (LBS) is location obfuscation, where mobile users are allowed to report obfuscated locations instead of their real locations to services. Many existing obfuscation approaches consider mobile users that can move freely over a region. However, this is inadequate for protecting the location privacy of vehicles, as their mobility is restricted by external factors, such as road networks and traffic flows. This auxiliary information about external factors helps an attacker to shrink the search range of vehicles' locations, increasing the risk of location exposure. Chenxi Qiu, Li Yan 0004, Anna Cinzia Squicciarini, Juanjuan Zhao 0001, Cheng-Zhong Xu 0001, Primal Pappachan |
SIGSPATIAL/GIS | 4 |
| 2022 | Ordis: A Dynamic Order-Dispatch Algorithm for Ridehailing and Ridesharing in a Large Region
Juanjuan Zhao 0001, Yang Wang 0006, Cheng-Zhong Xu 0001 |
ICA3PP | 2 |
| 2022 | Distributed Data-Sharing Consensus in Cooperative Perception of Autonomous VehiclesabstractTo enable self-driving without a human driver, an autonomous vehicle needs to perceive its surrounding obstacles using onboard sensors, of which the perception accuracy might be limited by their own sensing range. An effective way to improve vehicles’ perception accuracy is to let nearby vehicles exchange their sensor data so that vehicles can detect obstacles beyond their own sensing ranges, called cooperative perception. The shared sensor data, however, might disclose the sensitive information of vehicles’ passengers, raising privacy and safety concerns (e.g. stalking or sensitive location leakage).In this paper, we propose a new data-sharing policy for the cooperative perception of autonomous vehicles, of which the objective is to minimize vehicles’ information disclosure without compromising their perception accuracy. Considering vehicles usually have different desires for data-sharing under different traffic environments, our policy provides vehicles autonomy to determine what types of sensor data to share based on their own needs. Moreover, given the dynamics of vehicles’ data-sharing decisions, the policy can be adjusted to incentivize vehicles’ decisions to converge to the desired decision field, such that a healthy cooperation environment can be maintained in a long term. To achieve such objectives, we analyze the dynamics of vehicles’ data-sharing decisions by resorting to the game theory model, and optimize the data-sharing ratio in the policy based on the analytic results. Finally, we carry out an extensive trace-driven simulation to test the performance of the proposed data-sharing policy. The experimental results demonstrate that our policy can help incentivize vehicles’ data-sharing decisions to the desired decision fields efficiently and effectively. Chenxi Qiu, Anna Cinzia Squicciarini, Qing Yang 0003, Song Fu, Juanjuan Zhao 0001, Cheng-Zhong Xu 0001 |
ICDCS | 6 |
| 2022 | CD-Guide: A Dispatching and Charging Approach for Electric TaxicabsabstractPrevious methods for passenger demand inference are unable to capture the effect of all possible random factors (e.g., accident and weather), hence resulting in insufficient accuracy. Moreover, due to the lack of charging optimization, existing taxicab dispatching methods cannot be applied to electric taxicabs directly. We propose CD-Guide, which provides Charging and Dispatching Guide for electric taxicabs based on customized selection and training of historical passenger demand data, multiobjective optimization, and reinforcement learning (RL). By analyzing a large-scale electric taxicab data set, we found that: 1) the histogram of passengers’ origin buildings is effective in illustrating the suitability of historical data for learning; 2) passenger demands in different regions vary a lot due to various random factors; and 3) charging time must be considered in dispatching electric taxicabs. We first develop a passenger demand inference model based on customized selection and training of suitable historical passenger demand data. Then, we develop two taxicab guidance methods that utilize multiobjective optimization and RL, respectively, to maximize the taxicab’s likelihood of finding passengers, maximally prevent the taxicab from missing passengers due to charging, and, meanwhile, maintain the continuous service of the taxicab. Extensive experiments on real-world data sets demonstrate that compared with the state of the art, CD-Guide increases the total number of served passengers by 100%, and the minimum State-of-Charge of all taxicabs by 75% during all time slots. Li Yan 0004, Haiying Shen, Liuwang Kang, Juanjuan Zhao 0001, Zhe Zhang 0048, Cheng-Zhong Xu 0001 |
IEEE Internet Things J. | 4 |
| 2022 | CatCharger: Deploying In-Motion Wireless Chargers in a Metropolitan Road Network via Categorization and Clustering of Vehicle TrafficabstractIn metropolitan areas with heavy transit demands, electric vehicles (EVs) are expected to be continuously driving without recharging downtime. Wireless power transfer (WPT) provides a promising solution for in-motion EV charging. Nevertheless, previous works are not directly applicable for the deployment of in-motion wireless chargers due to their different charging characteristics. The challenge of deploying in-motion wireless chargers to support the continuous driving of EVs in a metropolitan road network with the minimum cost remains unsolved. We proposeCatChargerto tackle this challenge. By analyzing a metropolitan-scale data set, we found that traffic attributes like vehicle passing speed, daily visit frequency at intersections (i.e., landmarks), and their variances are diverse, and these attributes are critical to in-motion wireless charging performance. Driven by these observations, we first group landmarks with similar attribute values using the entropy minimization clustering method, and select candidate landmarks from the groups with suitable attribute values. Then, we use the kernel density estimator (KDE) to deduce the expected vehicle residual energy at each candidate landmark and consider EV drivers’ routing choice behavior in charger deployment. Finally, we determine the deployment locations by formulating and solving a multiobjective optimization problem, which maximizes vehicle traffic flow at charger deployment positions while guaranteeing the continuous driving of EVs at each landmark. Trace-driven experiments demonstrate thatCatChargerincreases the ratio of driving EVs at the end of a day by 12.5% under the same deployment cost. Li Yan 0004, Haiying Shen, Juanjuan Zhao 0001, Cheng-Zhong Xu 0001, Feng Luo 0001, Chenxi Qiu, Zhe Zhang 0048, Shohaib Mahmud |
IEEE Internet Things J. | 3 |
| 2022 | How to Build a Graph-Based Deep Learning Architecture in Traffic Domain: A SurveyabstractIn recent years, various deep learning architectures have been proposed to solve complex challenges (e.g. spatial dependency, temporal dependency) in traffic domain, which have achieved satisfactory performance. These architectures are composed of multiple deep learning techniques in order to tackle various challenges in traffic tasks. Traditionally, convolution neural networks (CNNs) are utilized to model spatial dependency by decomposing the traffic network as grids. However, many traffic networks are graph-structured in nature. In order to utilize such spatial information fully, it’s more appropriate to formulate traffic networks as graphs mathematically. Recently, various novel deep learning techniques have been developed to process graph data, called graph neural networks (GNNs). More and more works combine GNNs with other deep learning techniques to construct an architecture dealing with various challenges in a complex traffic task, where GNNs are responsible for extracting spatial correlations in traffic network. These graph-based architectures have achieved state-of-the-art performance. To provide a comprehensive and clear picture of such emerging trend, this survey carefully examines various graph-based deep learning architectures in many traffic applications. We first give guidelines to formulate a traffic problem based on graph and construct graphs from various kinds of traffic datasets. Then we decompose these graph-based architectures to discuss their shared deep learning techniques, clarifying the utilization of each technique in traffic tasks. What’s more, we summarize some common traffic challenges and the corresponding graph-based deep learning solutions to each challenge. Finally, we provide benchmark datasets, open source codes and future research directions in this rapidly growing field. Jiexia Ye, Juanjuan Zhao 0001, Kejiang Ye, Cheng-Zhong Xu 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | MDLF: A Multi-View-Based Deep Learning Framework for Individual Trip Destination Prediction in Public Transportation SystemsabstractUnderstanding and predicting each individual’s real-time travel destination given the origin information in urban public transportation systems is crucial for personalized traveler recommendation, targeted demand management, dynamic traffic operations and so on. Existing methods are often based on modeling the regular travel patterns through analyzing the long-term personal travel information. They are suitable for destination prediction of individual regular trips with regular travel patterns, but may not work well for occasional trips with strong randomness and uncertainty, especially for the individuals with a few historical travel data. In this paper, we focus on more challenging issue about destination prediction of occasional trips. We design a general Multi-View Deep Learning Framework (MDLF) based on the data-driven insight that a location where a user will destine to is not only related to the user’s own travel preference to the location, but also influenced by crowd’s travel preference and the region’s characteristics of the location under certain spatiotemporal contexts. The destination of an individual’s occasional trip can be predicted by combining all these complementary influencing factors. The novelty of MDLF is mainly reflected in two aspects. The first is the effective feature extraction from multiple and complementary views. The second is that a CNN (Recurrent Neural Network) based deep learning component for predicting each occasional trip’s destination by calculating a moving trend score for each possible destination. We evaluate the MDLF based on two real-world smart card datasets collected by AFC (Automatic Fare Collection) Systems. The experimental results demonstrate the superiority of MDLF against other competitors. Juanjuan Zhao 0001, Liutao Zhang, Jiexia Ye, Cheng-Zhong Xu 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | GLTC: A Metro Passenger Identification Method Across AFC Data and Sparse WiFi DataabstractIn this paper, we investigate an efficient way for identifying passengers in a metro system across two heterogeneous but complementary trajectory data sources: AFC data recording two points per trip about when and where a passenger enters or leaves the metro system, and WiFi data recording a few points passed by a passenger in the way of some of his/her trips. The identification result can help us to complete individuals’ mobility, and benefits to lots of services, e.g., individual route choice analysis, epidemic case detecting and so on. The problem is similar to calculate the similarity between two trajectories from two data sources, where a trajectory refers to a sequence of points where a passenger appeared in a metro system on observed days. However, due to the small location space in a metro network, large number of passengers with similar travel pattern, and so on, there are lots of trip overlaps or point co-occurrences between different passengers. That results in a large number of passengers mismatched by existing trajectory similarity measurement. To address the problem, this paper proposes a novel global-local correlation based trajectory similarity measurement GLTC. Specifically, GLTC first extracts all overlapping trip pairs of two trajectories by considering the spatiotemporal inclusions from global level. Then it gets the similarity by aggregating each overlapping trip pair’s local similarity, which is calculated by considering some data-driven insights helpful to uniquely identify a passenger (e.g., uneven passenger flow distribution in different cross-sections of a metro network, the number of trips in same travel pattern of a trajectory, and so on). We evaluate GLTC based on real-world data, and the experimental result shows that GLTC outperforms other baselines. Juanjuan Zhao 0001, Liutao Zhang, Kejiang Ye, Jiexia Ye, Jun Zhang 0014, Fan Zhang 0019, Cheng-Zhong Xu 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | CoScal: Multifaceted Scaling of Microservices With Reinforcement LearningabstractThe emerging trend towards moving from monolithic applications to microservices has raised new performance challenges in cloud computing environments. Compared with traditional monolithic applications, the microservices are lightweight, fine-grained, and must be executed in a shorter time. Efficient scaling approaches are required to ensure microservices’ system performance under diverse workloads with strict Quality of Service (QoS) requirements and optimize resource provisioning. To solve this problem, we investigate the trade-offs between the dominant scaling techniques, including horizontal scaling, vertical scaling, and brownout in terms of execution cost and response time. We first present a prediction algorithm based on gradient recurrent units to accurately predict workloads assisting in scaling to achieve efficient scaling. Further, we propose a multi-faceted scaling approach using reinforcement learning called CoScal to learn the scaling techniques efficiently. The proposed CoScal approach takes full advantage of data-driven decisions and improves the system performance in terms of high communication cost and delay. We validate our proposed solution by implementing a containerized microservice prototype system and evaluated with two microservice applications. The extensive experiments demonstrate that CoScal reduces response time by 19%-29% and decreases the connection time of services by 16% when compared with the state-of-the-art scaling techniques for Sock Shop application. CoScal can also improve the number of successful transactions with 6%-10% for Stan’s Robot Shop application. Minxian Xu, Chenghao Song, Shashikant Ilager, Sukhpal Singh, Juanjuan Zhao 0001, Kejiang Ye, Cheng-Zhong Xu 0001 |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2021 | Dynamic traffic bottlenecks identification based on congestion diffusion model by influence maximization in metro-city scalesabstractSummary Traffic bottlenecks dynamically change with the variance of traffic demand. Identifying traffic bottlenecks plays an important role in traffic planning and provides decision making. However, traffic bottlenecks are difficult to identify because of the complexity of traffic road networks and many other factors. In this article, we propose an influence spreading based method to find the dynamic changed traffic bottlenecks, where the influence caused by bottlenecks is maximal. We first build a traffic congestion diffusion (TCD) model to capture traffic flow influence (TFI) spreading over traffic road networks. The bottlenecks identification problem based on TCD is modeled as an influence maximization problem, that is, selecting the most influential nodes such that the deterioration of traffic condition is maximal. With the proof of the submodularity of TFI spreading over traffic networks, a provably near‐optimal algorithm is used to solve the NP‐hard problem. With the exploration of unique properties of TFI spread, an approximate influence maximization method for TCD (TCD‐AIM) is proposed. To the best of our knowledge, this should be the first model for a metro‐city scale from the influence perspective. Experimental results show that TCD‐AIM finds bottlenecks with up to 130% congestion density increase in the future. Baoxin Zhao, Cheng-Zhong Xu 0001, Siyuan Liu 0001, Juanjuan Zhao 0001, Li Li 0064 |
Concurr. Comput. Pract. Exp. | 4 |
| 2020 | MobiCharger: Optimal Scheduling for Cooperative EV-to-EV Dynamic Wireless ChargingabstractWith ever increasing concerns on environmental issues caused by gasoline fuel based vehicles, electric vehicles (EVs) have attracted more and more attention from governments, industries, and customers [1] . The recent advancements in EVs have great potential to create a more environmentally friendly smart city. However, due to limited battery capacity, most current mainstream EVs still have quite limited driving range (e.g., 100 miles) [2] . How to ensure the continuous running of EVs on a large-scale road network (e.g., metropolitan city, interstate) becomes a major concern. Li Yan 0004, Haiying Shen, Liuwang Kang, Juanjuan Zhao 0001, Cheng-Zhong Xu 0001 |
ICDCS | 4 |
| 2020 | Multi-Graph Convolutional Network for Relationship-Driven Stock Movement PredictionabstractStock price movement prediction is commonly accepted as a very challenging task due to the volatile nature of financial markets. Previous works typically predict the stock price mainly based on its own information, neglecting the cross effect among involved stocks. However, it is well known that an individual stock price is correlated with prices of other stocks in complex ways. To take the cross effect into consideration, we propose a deep learning framework, called Multi-GCGRU, which comprises graph convolutional network (GCN) and gated recurrent unit (GRU) to predict stock movement. Specifically, we first encode multiple relationships among stocks into graphs based on financial domain knowledge and utilize GCN to extract the cross effect based on these pre-defined graphs. To further get rid of prior knowledge, we explore an adaptive relationship learned by data automatically. The cross-correlation features produced by GCN are concatenated with historical records and then fed into GRU to model the temporal dependency of stock prices. Experiments on two stock indexes in China market show that our model outperforms other baselines. Note that our model is rather feasible to incorporate more effective stock relationships containing expert knowledge, as well as learn data-driven relationship. Jiexia Ye, Juanjuan Zhao 0001, Kejiang Ye, Cheng-Zhong Xu 0001 |
ICPR | 2 |
| 2020 | Multi-STGCnet: A Graph Convolution Based Spatial-Temporal Framework for Subway Passenger Flow ForecastingabstractSubway passenger flow forecasting, an essential component of intelligent transportation system, is critical for traffic management, public safety, urban planning. However, it is very challenging due to the high nonlinearities and complex dynamic spatio-temporal dependencies of passenger flows. In this paper, we model the subway system as a directed weighted graph and propose a novel spatio-temporal deep learning framework, Multi-STGCnet, for forecasting short-term subway passenger flow at a station level. Specifically, Multi-STGCnet is mainly composed of two components, temporal component and spatial component. (1) The temporal component employs three long short-term memory network (LSTM)-based modules to capture three temporal properties of the target station, which are the interval closeness, daily periodicity, weekly trend. (2) The spatial component designs three spatial matrixes to extract spatial correlation of a target station with all other stations classified as near neighbors, middle neighbors and distant neighbors. Respectively, it adopts three graph convolution network (GCN) and LSTM combined modules to capture the spatio-temporal influences from different neighbors. Finally, the outputs of the two components are fused with different weights to generate prediction. We evaluate Multi-STGCnet on a real world dataset from the metro system in Shenzhen, China. Experiment results demonstrate that our model outperforms multiple baselines. Jiexia Ye, Juanjuan Zhao 0001, Kejiang Ye, Cheng-Zhong Xu 0001 |
IJCNN | 2 |
| 2020 | CD-Guide: A Reinforcement Learning based Dispatching and Charging Approach for Electric TaxicabsabstractPrevious passenger demand inference methods have insufficient accuracy because they fail to catch the influence of all random factors (e.g., weather, holiday). Also, existing taxicab dispatching methods are not directly applicable for electric taxicabs because they cannot optimize their charging. We present CD-Guide: an electric taxicab dispatching and charging approach based on customized training and Reinforcement Learning (RL). We studied a metropolitan-scale taxicab dataset, and found: histogram of passengers' origin buildings (i.e., where they come from) is useful for selecting suitable training data for inference model, passenger demand in different regions may be influenced by various unpredictable random factors, and taxicabs' charging time must be considered to avoid missing potential passengers. By saying suitable historical data, we mean the data that are under the influence of random factors similar as current time. Then, we develop a RL based method to guide a taxicab to maximize its probability of picking up a passenger, minimize the number of its missed passengers due to charging, and meanwhile avoid the taxicab from battery exhaustion. Our trace-driven experiments show that compared with previous methods, CD-Guide increases the total number of served passengers by 100%. Li Yan 0004, Haiying Shen, Liuwang Kang, Juanjuan Zhao 0001, Cheng-Zhong Xu 0001 |
MASS | 4 |
| 2020 | Reinforcement Learning based Scheduling for Cooperative EV-to-EV Dynamic Wireless ChargingabstractPrevious Electric Vehicle (EV) charging scheduling methods and EV route planning methods require EVs to spend extra waiting time and driving burden for a recharge. With the advancement of dynamic wireless charging for EVs, Mobile Energy Disseminator (MED), which can charge an EV in motion, becomes available. However, existing wireless charging scheduling methods for wireless sensors, which are the most related works to the deployment of MEDs, are not directly applicable for the scheduling of MEDs on city-scale road networks. We present MobiCharger: a Mobile wireless Charger guidance system that determines the number of serving MEDs, and the optimal routes of the MEDs periodically (e.g., every 30 minutes). Through analyzing a metropolitan-scale vehicle mobility dataset, we found that most vehicles have routines, and the temporal change of the number of driving vehicles changes during different time slots, which means the number of MEDs should adaptively change as well. Then, we propose a Reinforcement Learning based method to determine the number and the driving route of serving MEDs. Our experiments driven by the dataset demonstrate that MobiCharger increases the medium state-of-charge and the number of charges of all EVs by 50% and 100%, respectively. Li Yan 0004, Haiying Shen, Liuwang Kang, Juanjuan Zhao 0001, Cheng-Zhong Xu 0001 |
MASS | 4 |
| 2019 | A Congestion Diffusion Model with Influence Maximization for Traffic Bottlenecks Identification in Metrocity ScalesabstractTraffic bottlenecks identification plays an important role in traffic planning and provides decision-making for prevention of traffic congestion. Although traffic bottlenecks widely exist, they are difficult to predict because of the changing traffic condition and traffic demand. In this paper, we introduce a traffic congestion diffusion (TCD) model with traffic flow influence (TFI) to capture the traffic dynamics and give a panoramic view for the city by cross domain data fusion. We proposed novel definition of bottleneck from the perspective of influence spread under TCD. The bottlenecks identification problem is modeled as an influence maximization problem, i.e., selecting the top K influential nodes in road networks under certain traffic conditions. We establish the submodularity of influence spread and solve the NP-hard optimal seed selection problem by using an efficient heuristic algorithm (TCD-IM) with provable near-optimal performance guarantees. To the best of our knowledge, this should be the first model for a metro-city scale from the influence perspective. The TCD-IM model is able to identify the dynamic traffic bottlenecks. Baoxin Zhao, Cheng-Zhong Xu 0001, Siyuan Liu 0001, Juanjuan Zhao 0001, Li Li 0064 |
IEEE BigData | 4 |
| 2017 | CatCharger: Deploying wireless charging lanes in a metropolitan road network through categorization and clustering of vehicle trafficabstractThe future generation of transportation system will be featured by electrified public transportation. To fulfill metropolitan transit demands, electric vehicles (EVs) must be continuously operable without recharging downtime. Wireless Power Transfer (WPT) techniques for in-motion EV charging is a solution. It however brings up a challenge: how to deploy charging lanes in a metropolitan road network to minimize the deployment cost while enabling EVs' continuous operability. In this paper, we propose CatCharger, which is the first work that handles this challenge. From a metropolitan-scale dataset collected from multiple sources of vehicles, we observe the diversity of vehicle passing speed and daily visit frequency (called traffic attributes) at intersections (i.e., landmarks), which are important factors for charging lane deployment. To select landmarks for deployment, we first group landmarks with similar traffic attribute values using the entropy minimization clustering method, and choose better candidate landmarks from each group suitable for deployment. To determine the deployment locations from the candidate landmarks, we infer the expected vehicle residual energy at each landmark using a Kernel Density Estimator fed by the vehicles' mobility, and formulate and solve an optimization problem to minimize the total deployment cost while ensuring a certain level of expected residual energy of EVs at each landmark. Our trace-driven experiments demonstrate the superior performance of CatCharger over other methods. Li Yan 0004, Haiying Shen, Juanjuan Zhao 0001, Cheng-Zhong Xu 0001, Feng Luo 0001, Chenxi Qiu |
INFOCOM | 3 |
| 2017 | Last-Mile Transit Service with Urban Infrastructure DataabstractIn this article, we propose a transit service Feeder to tackle the last-mile problem, that is, passengers’ destinations lay beyond a walking distance from a public transit station. Feeder utilizes ridesharing-based vehicles (e.g., minibus) to deliver passengers from existing transit stations to selected stops closer to their destinations. We infer real-time passenger demand (e.g., exiting stations and times) for Feeder design by utilizing extreme-scale urban infrastructures, which consist of 10 million cellphones, 27 thousand vehicles, and 17 thousand smartcard readers for 16 million smartcards in a Chinese city, Shenzhen. Regarding these numerous devices as pervasive sensors, we mine both online and offline data for a two-end Feeder service: a back-end Feeder server to calculate service schedules and front-end customized Feeder devices in vehicles for real-time schedule downloading. We implement Feeder using a fleet of vehicles with customized hardware in a subway station of Shenzhen by collecting data for 30 days. The evaluation results show that compared to the ground truth, Feeder reduces last-mile distances by 68% and travel time by 56%, on average. Desheng Zhang 0002, Juanjuan Zhao 0001, Fan Zhang 0019, Ruobing Jiang, Tian He 0001, Nikolaos Papanikolopoulos |
ACM Trans. Cyber Phys. Syst. | 2 |
| 2017 | Spatio-Temporal Analysis of Passenger Travel Patterns in Massive Smart Card DataabstractMetro systems have become one of the most important public transit services in cities. It is important to understand individual metro passengers' spatio-temporal travel patterns. More specifically, for a specific passenger: what are the temporal patterns? what are the spatial patterns? is there any relationship between the temporal and spatial patterns? are the passenger's travel patterns normal or special? Answering all these questions can help to improve metro services, such as evacuation policy making and marketing. Given a set of massive smart card data over a long period, how to effectively and systematically identify and understand the travel patterns of individual passengers in terms of space and time is a very challenging task. This paper proposes an effective data-mining procedure to better understand the travel patterns of individual metro passengers in Shenzhen, a modern and big city in China. First, we investigate the travel patterns in individual level and devise the method to retrieve them based on raw smart card transaction data, then use statistical-based and unsupervised clustering-based methods, to understand the hidden regularities and anomalies of the travel patterns. From a statistical-based point of view, we look into the passenger travel distribution patterns and find out the abnormal passengers based on the empirical knowledge. From unsupervised clustering point of view, we classify passengers in terms of the similarity of their travel patterns. To interpret the group behaviors, we also employ the bus transaction data. Moreover, the abnormal passengers are detected based on the clustering results. At last, we provide case studies and findings to demonstrate the effectiveness of the proposed scheme. Juanjuan Zhao 0001, Qiang Qu 0001, Fan Zhang 0019, Cheng-Zhong Xu 0001, Siyuan Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2017 | Estimation of Passenger Route Choice Pattern Using Smart Card Data for Complex Metro SystemsabstractMetro systems play an important role in meeting the demand for urban transportation in large cities. The understanding of passenger route choice is critical for public transit management. The wide deployment of automated fare collection (AFC) systems opens up a new opportunity. However, only each trip's tap-in and tap-out time stamp and stations can be directly obtained from AFC system records; the train and route chosen by a passenger are unknown, information necessary to solve our problem. While existing methods work well in some specific situations, they hardly work for complicated situations. In this paper, we propose a solution that needs no additional equipment or human involvement than the AFC systems. We develop a probabilistic model that can estimate from empirical analysis how the passenger flows are dispatched to different routes and trains. We validate our approach using a large-scale data set collected from the Shenzhen Metro system. The measured results provide us with useful input when building the passenger path choice model. Juanjuan Zhao 0001, Fan Zhang 0019, Lai Tu, Cheng-Zhong Xu 0001, Dayong Shen, Chen Tian 0001, Xiang-Yang Li 0001, Zhengxi Li |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2016 | CatCharge: Deploying wireless charging lane in metropolitan scale through categorization and clustering of vehicle mobilityabstractThe future generation transportation system will be featured by electrified public transportation. To fulfill metropolitan transit demands, electric vehicles (EVs) must be continuously operable without recharging downtime. Wireless Power Transfer (WPT) techniques for in-motion EV charging is a solution [1], [2]. It however brings up a challenge: how to deploy charging lanes in a metropolitan road network to minimize the deployment cost while enabling EVs' continuous operability. Li Yan 0004, Juanjuan Zhao 0001, Haiying Shen, Cheng-Zhong Xu 0001, Feng Luo 0001 |
ICNP | 2 |
| 2016 | Heterogeneous Model Integration for Multi-Source Urban Infrastructure DataabstractData-driven modeling usually suffers from data sparsity, especially for large-scale modeling for urban phenomena based on single-source urban-infrastructure data under fine-grained spatial-temporal contexts. To address this challenge, we motivate, design, and implement UrbanCPS, a cyber-physical system with heterogeneous model integration, based on extremely-large multi-source infrastructures in the Chinese city Shenzhen, involving 42,000 vehicles, 10 million residents, and 16 million smartcards. Based on temporal, spatial, and contextual contexts, we formulate an optimization problem about how to optimally integrate models based on highly diverse datasets under three practical issues, that is, heterogeneity of models, input data sparsity, or unknown ground truth. We further propose a real-world application called Speedometer, inferring real-time traffic speeds in urban areas. The evaluation results show that, compared to a state-of-the-art system, Speedometer increases the inference accuracy by 29% on average. Desheng Zhang 0002, Juanjuan Zhao 0001, Fan Zhang 0019, Tian He 0001, Haengju Lee, Sang Hyuk Son |
ACM Trans. Cyber Phys. Syst. | 2 |
| 2015 | coMobile: real-time human mobility modeling at urban scale using multi-view learningabstractReal-time human mobility modeling is essential to various urban applications. To model such human mobility, numerous data-driven techniques have been proposed. However, existing techniques are mostly driven by data from a single view, e.g., a transportation view or a cellphone view, which leads to over-fitting of these single-view models. To address this issue, we propose a human mobility modeling technique based on a generic multi-view learning framework called coMobile. In coMobile, we first improve the performance of single-view models based on tensor decomposition with correlated contexts, and then we integrate these improved single-view models together for multi-view learning to iteratively obtain mutually-reinforced knowledge for real-time human mobility at urban scale. We implement coMobile based on an extremely large dataset in the Chinese city Shenzhen, including data about taxi, bus and subway passengers along with cellphone users, capturing more than 27 thousand vehicles and 10 million urban residents. The evaluation results show that our approach outperforms a single-view model by 51% on average. Desheng Zhang 0002, Juanjuan Zhao 0001, Fan Zhang 0019, Tian He 0001 |
SIGSPATIAL/GIS | 2 |
| 2015 | Feeder: supporting last-mile transit with extreme-scale urban infrastructure dataabstractIn this paper, we propose a transit service Feeder to tackle the last-mile problem, i.e., passengers' destinations lay beyond a walking distance from a public transit station. Feeder utilizes ridesharing-based vehicles (e.g., minibus) to deliver passengers from existing transit stations to selected stops closer to their destinations. We infer real-time passenger demand (e.g., exiting stations and times) for Feeder design by utilizing extreme-scale urban infrastructures, which consist of 10 million cellphones, 27 thousand vehicles, and 17 thousand smartcard readers for 16 million smartcards in a Chinese city Shenzhen. Regarding these numerous devices as pervasive sensors, we mine both online and offline data for a two-end Feeder service: a back-end Feeder server to calculate service schedules; front-end customized Feeder devices in vehicles for real-time schedule downloading. The evaluation results show that compared to the ground truth, Feeder reduces last-mile distances by 68% and travel time by 52% on average. Desheng Zhang 0002, Juanjuan Zhao 0001, Fan Zhang 0019, Ruobing Jiang, Tian He 0001 |
IPSN | 2 |
| 2013 | A characterization of big data benchmarksabstractRecently, big data has been evolved into a buzzword from academia to industry all over the world. Benchmarks are important tools for evaluating an IT system. However, benchmarking big data systems is much more challenging than ever before. First, big data systems are still in their infant stage and consequently they are not well understood. Second, big data systems are more complicated compared to previous systems such as a single node computing platform. While some researchers started to design benchmarks for big data systems, they do not consider the redundancy between their benchmarks. Moreover, they use artificial input data sets rather than real world data for their benchmarks. It is therefore unclear whether these benchmarks can be used to precisely evaluate the performance of big data systems. In this paper, we first analyze the redundancy among benchmarks from ICTBench, HiBench and typical workloads from real world applications: spatio-temporal data analysis for Shenzhen transportation system. Subsequently, we present an initial idea of a big data benchmark suite for spatio-temporal data. There are three findings in this work: (1) redundancy exists in these pioneering benchmark suites and some of them can be removed safely. (2) The workload behavior of trajectory data analysis applications is dramatically affected by their input data sets. (3) The benchmarks created for academic research cannot represent the cases of real world applications. Zhibin Yu 0001, Zhendong Bei, Juanjuan Zhao 0001, Fan Zhang 0019, Yubin Zou, Ye Li 0002, Cheng-Zhong Xu 0001 |
IEEE BigData | 4 |