Lingyu Zhang 0001

dblp:35/10185-1 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-4651-7991ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 11 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 1 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Visualization-Oriented Progressive Time Series Transformation
abstract
Visual analysis of large time-series data often requires transformations over multivariate time series. Existing methods struggle to meet interactive response time requirements, relying on full transformations that incur high computation costs. We propose a visualization-oriented transformation system PIVOT that incrementally generates accurate visualizations by selectively transforming only essential data samples. At its core is a transformation-aware query mechanism that efficiently computes point-wise transformations by leveraging cached hierarchical data on the server. To support responsive interaction, we introduce a pixel-based error-bound guarantee that estimates the accuracy of intermediate visualizations without requiring a reference, enabling a balance between latency and visual fidelity. Experiments show that PIVOT achieves highly accurate visualizations with interactive response times, outperforming existing error-free methods by up to an order of magnitude on billion-scale datasets.
Xin Chen 0075, Lingyu Zhang 0001, Huaiwei Bao, Wei Lu 0015, Eugene Wu 0002, Xiaohui Yu 0001, Yunhai Wang
Proc. ACM Manag. Data2
2023 Easy Begun Is Half Done: Spatial-Temporal Graph Modeling with ST-Curriculum Dropout
abstract
Spatial-temporal (ST) graph modeling, such as traffic speed forecasting and taxi demand prediction, is an important task in deep learning area. However, for the nodes in the graph, their ST patterns can vary greatly in difficulties for modeling, owning to the heterogeneous nature of ST data. We argue that unveiling the nodes to the model in a meaningful order, from easy to complex, can provide performance improvements over traditional training procedure. The idea has its root in Curriculum Learning, which suggests in the early stage of training models can be sensitive to noise and difficult samples. In this paper, we propose ST-Curriculum Dropout, a novel and easy-to-implement strategy for spatial-temporal graph modeling. Specifically, we evaluate the learning difficulty of each node in high-level feature space and drop those difficult ones out to ensure the model only needs to handle fundamental ST relations at the beginning, before gradually moving to hard ones. Our strategy can be applied to any canonical deep learning architecture without extra trainable parameters, and extensive experiments on a wide range of datasets are conducted to illustrate that, by controlling the difficulty level of ST relations as the training progresses, the model is able to capture better representation of the data and thus yields better generalization.
Hongjun Wang 0007, Jiyuan Chen, Tong Pan, Zipei Fan, Xuan Song 0001, Renhe Jiang, Lingyu Zhang 0001, Boyuan Zhang 0005
AAAI7
2023 Multi-Task Weakly Supervised Learning for Origin-Destination Travel Time Estimation
abstract
Travel time estimation from GPS trips is of great importance to order duration, ridesharing, taxi dispatching, etc. However, the dense trajectory is not always available due to the limitation of data privacy and acquisition, while the origin-destination (OD) type of data, such as NYC taxi data, NYC bike data, and Capital Bikeshare data, is more accessible. To address this issue, this paper starts to estimate the OD trips travel time combined with the road network. Subsequently, aMulti-taskWeaklySupervisedLearning Framework forTravelTimeEstimation (MWSL-TTE) has been proposed to infer transition probability between roads segments, and the travel time on road segments and intersection simultaneously. Technically, given an OD pair, the transition probability intends to recover the most possible route. And then, the output of travel time is equal to the summation of all segments’ and intersections’ travel time in this route. A novel route recovery function has been proposed to iteratively maximize the current routes’ co-occurrence probability, and minimize the discrepancy between routes’ probability distribution and the inverse distribution of routes’ estimation loss. Moreover, the expected log-likelihood function based on a weakly-supervised framework has been deployed in optimizing the travel time from road segments and intersections concurrently. We conduct experiments on a wide range of real-world taxi datasets in Xi’an and Chengdu and demonstrate our method's effectiveness on route recovery and travel time estimation.
Hongjun Wang 0007, Zhiwen Zhang 0004, Zipei Fan, Jiyuan Chen, Lingyu Zhang 0001, Ryosuke Shibasaki, Xuan Song 0001
IEEE Trans. Knowl. Data Eng.5
2022 Pick-Up Point Recommendation Using Users' Historical Ride-Hailing Orders
Lingyu Zhang 0001, Zhijie He, Guobin Wu 0001, Ziqiang Yu, Minghao Ji, Yunhai Wang
WASA (2)1
2022 Users' Departure Time Prediction Based on Light Gradient Boosting Decision Tree
Lingyu Zhang 0001, Zhijie He, Guobin Wu 0001, Ziqiang Yu, Minghao Ji, Yunhai Wang
WASA (2)1
2021 Engaging Drivers in Ride Hailing via Competition: A Case Study with Arena
abstract
Sustained work enthusiasms of drivers are crucial for the success of large-scale ride-hailing platforms. In this paper, we conduct the first-of-its-kind exploration to encourage active participation of drivers via competition. We design Arena, a competition where drivers compete for prizes via completing more trips. Through a pilot study covering over 2,600 participants, we uncover the easy-win problem, an overlooked and serious issue in competition design for real-world drivers. It refers to situations where one competitor does not show up during competition whereas the other easily wins. To solve the easy-win problem without impairing motivation of drivers, we devise a novel prediction-based matchmaking framework. On observing that no-shows are highly correlated to the online time of drivers during competition, we propose to identify potential no-shows by predicting drivers' online time and avoid matching potential noshow drivers with drivers that will show up so as to reduce easy-wins. We conduct large-scale experiments based on real competition data involving over 10,000 drivers. The results show that our prediction-based matchmaking scheme can effectively reduce the ratio of easy-wins.
Shuyue Wei 0001, Lingyu Zhang 0001, Zimu Zhou, Yongxin Tong
MDM3
2020 Predicting Origin-Destination Flow via Multi-Perspective Graph Convolutional Network
abstract
Predicting Origin-Destination (OD) flow is a crucial problem for intelligent transportation. However, it is extremely challenging because of three reasons: first, correlations exist between both origins and destinations; second, the correlations are dynamic across the time; at last, there are multiple correlations from different aspects. To the best of our knowledge, existing models for OD flow prediction cannot tackle all of these three issues simultaneously. We propose Multi-Perspective Graph Convolutional Networks (MPGCN) to capture the complex dependencies. Our proposed model first utilizes long short-term memory (LSTM) network to extract temporal features for each OD pair and then learns the spatial dependency of origins and destinations by a two-dimensional graph convolutional network. Furthermore, we design a dynamic graph together with two static graphs to capture the complicated spatial dependencies and use an average strategy to obtain the final predicted OD flow. We conduct extensive experiments on two large-scale and real-world datasets, which not only demonstrate our design philosophy but also validate the effectiveness of the proposed model.
Hongzhi Shi, Quanming Yao, Lingyu Zhang 0001, Jieping Ye, Yong Li 0008, Yan Liu 0002
ICDE5
2020 Predicting Individual Treatment Effects of Large-scale Team Competitions in a Ride-sharing Economy
abstract
Millions of drivers worldwide have enjoyed financial benefits and work schedule flexibility through a ride-sharing economy, but meanwhile they have suffered from the lack of a sense of identity and career achievement. Equipped with social identity and contest theories, financially incentivized team competitions have been an effective instrument to increase drivers' productivity, job satisfaction, and retention, and to improve revenue over cost for ride-sharing platforms. While these competitions are overall effective, the decisive factors behind the treatment effects and how they affect the outcomes of individual drivers have been largely mysterious. In this study, we analyze data collected from more than 500 large-scale team competitions organized by a leading ride-sharing platform, building machine learning models to predict individual treatment effects. Through a careful investigation of features and predictors, we are able to reduce out-sample prediction error by more than 24%. Through interpreting the best-performing models, we discover many novel and actionable insights regarding how to optimize the design and the execution of team competitions on ride-sharing platforms. A simulated analysis demonstrates that by simply changing a few contest design options, the average treatment effect of a real competition is expected to increase by as much as 26%. Our procedure and findings shed light on how to analyze and optimize large-scale online field experiments in general.
Teng Ye, Wei Ai 0002, Lingyu Zhang 0001, Jieping Ye, Qiaozhu Mei
KDD3
2020 Two-sided online bipartite matching in spatial data: experiments and analysis
Jingzhi Fang, Yuxiang Zeng, Balz Maag, Yongxin Tong, Lingyu Zhang 0001
GeoInformatica6
2019 Spatiotemporal Multi-Graph Convolution Network for Ride-Hailing Demand Forecasting
abstract
Region-level demand forecasting is an essential task in ridehailing services. Accurate ride-hailing demand forecasting can guide vehicle dispatching, improve vehicle utilization, reduce the wait-time, and mitigate traffic congestion. This task is challenging due to the complicated spatiotemporal dependencies among regions. Existing approaches mainly focus on modeling the Euclidean correlations among spatially adjacent regions while we observe that non-Euclidean pair-wise correlations among possibly distant regions are also critical for accurate forecasting. In this paper, we propose the spatiotemporal multi-graph convolution network (ST-MGCN), a novel deep learning model for ride-hailing demand forecasting. We first encode the non-Euclidean pair-wise correlations among regions into multiple graphs and then explicitly model these correlations using multi-graph convolution. To utilize the global contextual information in modeling the temporal correlation, we further propose contextual gated recurrent neural network which augments recurrent neural network with a contextual-aware gating mechanism to re-weights different historical observations. We evaluate the proposed model on two real-world large scale ride-hailing demand datasets and observe consistent improvement of more than 10% over stateof-the-art baselines.
Xu Geng, Leye Wang, Lingyu Zhang 0001, Qiang Yang 0001, Jieping Ye, Yan Liu 0002
AAAI4
2019 CIKM 2019 Workshop on Artificial Intelligence in Transportation (AI in transportation)
abstract
Data-enabled smart transportation has attracted a surge of interest from machine learning and data mining researchers nowadays due to the bloom of online ride-hailing industry and rapid development of autonomous driving. Large-scale high quality route data and trading data (spatiotemporal data) have been generated every day, which makes AI an urgent need and preferred solution for the decision making in intelligent transportation systems. While a large of amount of work have been dedicated to traditional transportation problems, they are far from satisfactory for the rising need. We propose a half-day workshop at CIKM 2019 for the professionals, researchers, and practitioners who are interested in mining and understanding big and heterogeneous data generated in transportation, and AI applications to improve the transportation system. We plan to have several invited talks from both academia and industry. This workshop would be organized by Shanghai Jiao Tong University, Didi Chuxing and Pennsylvania State University.
Weinan Zhang 0001, Haiming Jin, Lingyu Zhang 0001, Hongtu Zhu, Zhenhui Jessie Li, Jieping Ye
CIKM3
2019 Recommendation-based Team Formation for On-demand Taxi-calling Platforms
abstract
On-demand taxi-calling platforms often ignore the social engagement of individual drivers. The lack of social incentives impairs the work enthusiasms of drivers and will affect the quality of service. In this paper, we propose to form teams among drivers to promote participation. A team consists of a leader and multiple members, which acts as the basis for various group-based incentives such as competition. We define the Recommendation-based Team Formation (RTF) problem to form as many teams as possible while accounting for the choices of drivers. The RTF problem is challenging. It needs both accurate recommendation and coordination among recommendations, since each driver can be in at most one team. To solve the RTF problem, we devise a Recommendation-Matrix-Based Framework (RMBF). It first estimates the acceptance probability of recommendations and then derives a recommendation matrix to maximize the number of formed teams from a global view. We conduct trace-driven simulations using real data covering over 64,000 drivers and deploy our solution on a large on-demand taxi-calling platform for online evaluations. Experimental results show that RMBF outperforms the greedy-based strategy by forming up to 20% and 12.4% teams in trace-driven simulations and online evaluations, and the drivers who form teams and are involved in the competition have more service time, number of finished orders and income.
Lingyu Zhang 0001, Tianshu Song, Yongxin Tong, Zimu Zhou, Wei Ai 0002, Guobin Wu 0001, Yan Liu 0002, Jieping Ye
CIKM1
2018 SIGIR 2018 Workshop on Intelligent Transportation Informatics
abstract
We propose a half-day workshop at SIGIR 2018 for the professionals, researchers, and practitioners who are interested in mining and understanding big and heterogeneous data generated in transportation to improve the transportation system. We plan to have both paper presentations and invited talks.
Yan Liu 0002, Zhenhui Li, Wei Ai 0002, Lingyu Zhang 0001
SIGIR4
2018 Taxi or Hitchhiking: Predicting Passenger's Preferred Service on Ride Sharing Platforms
abstract
Ride sharing apps like Uber and Didi Chuxing have played an important role in addressing the users' transportation needs, which come not only in huge volumes, but also in great variety. While some users prefer low-cost services such as carpooling or hitchhiking, others prefer more pricey options like taxi or premier services. Further analyses suggest that such preference may also be associated with different time and location. In this paper, we empirically analyze the preferred services and propose a recommender system which provides service recommendation based on temporal, spatial, and behavioral features. Offline simulations show that our system achieves a high prediction accuracy and reduces the user's effort in finding the desired service. Such a recommender system allows a more precise scheduling for the platform, and enables personalized promotions.
Lingyu Zhang 0001, Wei Ai 0002, Chuan Yuan, Jieping Ye
SIGIR1
2017 A Taxi Order Dispatch Model based On Combinatorial Optimization
abstract
Taxi-booking apps have been very popular all over the world as they provide convenience such as fast response time to the users. The key component of a taxi-booking app is the dispatch system which aims to provide optimal matches between drivers and riders. Traditional dispatch systems sequentially dispatch taxis to riders and aim to maximize the driver acceptance rate for each individual order. However, the traditional systems may lead to a low global success rate, which degrades the rider experience when using the app. In this paper, we propose a novel system that attempts to optimally dispatch taxis to serve multiple bookings. The proposed system aims to maximize the global success rate, thus it optimizes the overall travel efficiency, leading to enhanced user experience. To further enhance users' experience, we also propose a method to predict destinations of a user once the taxi-booking APP is started. The proposed method employs the Bayesian framework to model the distribution of a user's destination based on his/her travel histories.
Lingyu Zhang 0001, Yue Min, Guobin Wu 0001, Pengcheng Feng, Pinghua Gong, Jieping Ye
KDD1