EDBT 2026 Demo / reviewers in the wild / expert
Tian He 0001
dblp:h/TianHe
· DBLP profile ↗
30ranked-venue papers in the field
0as first author
21since 2021 · last 2026
0000-0001-6062-2619ORCID · conflict
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 10Database Systems & Data Management · 9Information Retrieval & Web Search · 9Big Data, Cloud & Distributed Data Systems · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GenHAR: Generalizing Cross-domain Human Activity Recognition for Last-mile Delivery
Zhiqing Hong, Xiubin Fan, Guang Yang 0028, Baoshen Guo, Haotian Wang 0008, Tian He 0001, Desheng Zhang 0002 |
KDD (1) | 7 |
| 2025 | D3-TR: Data-driven Daily Delivery Task Rescheduling for Cost-effective Last-mile DeliveryabstractIn last-mile logistics, couriers are typically assigned fixed zones to perform door-to-door deliveries. In practice, packages in some delivery zones might not be fulfilled on time due to couriers taking irregular leave for sickness or higher-priority task assignments, e.g., services for VIPs and regulatory training. Beyond the costly real-world practice, i.e., hiring temporary workers, analysis of historical data reveals that a daily delivery task rescheduling among on-duty couriers can be a cost-effective and efficient alternative. It involves individual workload assessments and delivery task assignments, both of which existing methods can not address adequately: (i) Existing courier workload assessment methods are not tailored for downstream optimization tasks, leading to poor performance. (ii) Efficiency-oriented task assignment methods may lead to unfair workload among the couriers. To address the above two limitations, in this paper, we propose D3-TR, a data-driven method for task reassignment among present couriers. Firstly, we design a consistency-guided predictor that can quickly and precisely predict the workload of couriers. Secondly, based on this predictor, we design a workload-aware genetic algorithm to solve the optimal task allocation problem. Experimental results underscore the superiority of our method over several baselines. Furthermore, real-world deployment on millions of orders demonstrates the effectiveness of our solution, yielding an average of 3.9% improvement in the on-time delivery rate. Lidi Zhang, Yinfeng Xiang, Wenjun Lyu, Zhiqing Hong, Haotian Wang 0008, Desheng Zhang 0002, Yunhuai Liu, Tian He 0001 |
CIKM | 8 |
| 2025 | LLM4HAR: Generalizable On-device Human Activity Recognition with Pretrained LLMsabstractA long-standing challenge for pushing sensor-based human activity recognition (HAR) to industrial usage is the distribution shift between training data and testing data: significant variations in data distribution lead to a notable decline in performance. Recently, Large Language Models (LLMs) have demonstrated exceptional generalization capability, which provides a new opportunity to mitigate the distribution shift problem of HAR. However, since LLMs are inherently designed and trained on textual data, their potential to enhance generalization in HAR applications remains an open question. In this paper, we introduce LLM4HAR, a novel LLM-based model to improve cross-domain HAR. LLM4HAR consists of three main modules: (i) the Sensor Data Adaptation module, which aligns IMU signals with LLMs via sensor embedding(ii) the Sensor Knowledge Learning module, which injects sensor knowledge into LLMs for activity recognition, and (iii) the Efficiency Enhancement module, which employs a partial training strategy and reduces the model size by more than 10 times. Extensive evaluations show that LLM4HAR outperforms the existing methods by 13.82% in average F1 score, demonstrating the feasibility and effectiveness of transferring knowledge from pretrained LLMs to enhance HAR. Further, LLM4HAR has been adopted by JD Logistics to support downstream applications such as Courier Welfare Improvement and Map Data Generation. Zhiqing Hong, Yiwei Song, Anlan Yu, Shuxin Zhong, Yi Ding 0011, Tian He 0001, Desheng Zhang 0002 |
KDD (2) | 7 |
| 2025 | A Transferable Spatio-temporal Learning Framework for Cross-city Logistics Demand PredictionabstractIn logistic systems, demand prediction is an essential task providing the basis for improving the quality of terminal services, such as pick-up and delivery efficiency. However, the geographical scope of operations across multiple cities brings challenges due to the sparsity of user behavior data, hindering accurate predictions. Despite cross-city prediction methods potentially solving this problem by relying on the label of overlapping users in different cities, annotating these overlapping users is expensive. Additionally, the dynamic and diverse nature of user behaviors complicates feature transfer between cities. In this work, we define the logistics demand prediction problem as forecasting pick-up and delivery demand for zones, the smallest operational units in logistics systems, in different cities. To address the challenge, we propose TSTL, a Transferable Spatio-Temporal Learning framework for cross-city logistics prediction with sparse user data. TSTL advances existing methods from two aspects: (1) User-level invariant representation module extracts consistent user representations for overlapping and non-overlapping users across cities. (2) User-zone graph aggregation module enhances user embeddings by integrating dynamic interactions, such as logistics behaviors, into inherent user relations. Finally, the multi-city transfer module fine-tunes model parameters for city-invariant knowledge adoption and predicts future logistics demand. We implement and evaluate TSTL on one of the largest logistics systems. Extensive offline experiments and real-world deployment demonstrate the effectiveness of TSTL. Kaiwen Xia, Li Lin 0011, Xinrui Zhang 0006, Haotian Wang 0008, Shuai Wang 0008, Tian He 0001 |
KDD (2) | 6 |
| 2025 | ProST: Prompt Future Snapshot on Dynamic Graphs for Spatio-Temporal PredictionabstractSpatio-temporal prediction focuses on jointly modeling spatial correlations and temporal evolution and has a wide range of applications. Due to the heterogeneity of spatio-temporal data, accurate prediction relies on effectively integrating topological structures and sequential patterns. Although recurrent graph learning methods excel at capturing dynamic graph patterns, explicitly inferring future snapshots from historical dynamic graphs remains a significant challenge. Recently, prompt-based graph learning has shown the potential to improve future snapshot inference by leveraging node or task-specific prompts. However, these methods fail to fully capture edge information resulting in incomplete and less accurate representations of future snapshot structures. To bridge this gap, we propose ProST, a framework that Prompts future snapshots on dynamic graphs for Spatio-Temporal prediction, which leverages dynamic graph pre-training to generate a premise graph containing historical graph information and then employs prompts on the premise graph to infer explicit future snapshots. Specifically, this framework comprises three steps: Firstly, dynamic graph pre-training is performed using multi-granularity evolution graph convolution to obtain the premise graph with both local and global features of dynamic graphs. Secondly, prompt subgraphs are used to prompt node pairs and edge features within the premise graph. The subgraph prompt aggregation mechanism propagates this information to generate future snapshots. Finally, we freeze the parameters of the pre-trained model and update the subgraph prompt parameters using meta-learning to adapt to downstream spatio-temporal prediction tasks. Extensive experiments on real-world datasets validate that ProST achieves state-of-the-art performance. Kaiwen Xia, Li Lin 0011, Shuai Wang 0008, Qi Zhang 0087, Shuai Wang 0021, Tian He 0001 |
KDD (1) | 6 |
| 2025 | Scalable Area Difficulty Assessment with Knowledge-enhanced AI for Nationwide Logistics Systems
Zejun Xie, Wenjun Lyu, Yiwei Song, Haotian Wang 0008, Guang Yang 0028, Yunhuai Liu, Tian He 0001, Desheng Zhang 0002, Guang Wang 0001 |
KDD (1) | 7 |
| 2025 | AddrLLM: Address Rewriting via Large Language Model on Nationwide Logistics DataabstractTextual description of a physical location, commonly known as an address, plays an important role in location-based services(LBS) such as on-demand delivery and navigation. However, the prevalence of abnormal addresses, those containing inaccuracies that fail to pinpoint a location, have led to significant costs. Address rewriting has emerged as a solution to rectify these abnormal addresses. Despite the critical need, existing address rewriting methods are limited, typically tailored to correct specific error types, or frequently require retraining to process new address data effectively. In this study, we introduce AddrLLM, an innovative framework for address rewriting that is built upon a retrieval augmented large language model. AddrLLM overcomes aforementioned limitations through a meticulously designed Supervised Fine-Tuning module, an Address-centric Retrieval Augmented Generation module and a Bias-free Objective Alignment module. To the best of our knowledge, this study pioneers the application of LLM-based address rewriting approach to solve the issue of abnormal addresses. Through comprehensive offline testing with real-world data on a national scale and subsequent online deployment, AddrLLM has demonstrated superior performance in integration with existing logistics system. It has significantly decreased the rate of parcel re-routing by approximately 43%, underscoring its exceptional efficacy in real-world applications. Qinchen Yang 0001, Zhiqing Hong, Dongjiang Cao, Haotian Wang 0008, Zejun Xie, Tian He 0001, Yunhuai Liu, Yu Yang 0010, Desheng Zhang 0002 |
KDD (1) | 6 |
| 2025 | InCo: Exploring Inter-Trip Cooperation for Efficient Last-mile DeliveryabstractAn efficient last-mile delivery scheme in logistics benefits customers, couriers, and the platform. In practice, the delivery scope of a delivery station is divided into multiple areas, each of which is covered by a courier. The long distances between the delivery station and areas limit the couriers' delivery efficiency given that they need to travel back and forth multiple times a day. To solve this problem, we explore an inter-trip cooperation scheme for last-mile delivery, in which couriers traveling to the delivery station and back to corresponding areas earlier can help to take others' orders back. Coordinating the courier cooperation is challenging because we need to consider the courier's status, e.g., locations, and vehicle capacity constraint simultaneously. In this work, we design an inter-trip cooperation-based last-mile delivery system, InCo, aiming to minimize the average order delivery time. InCo includes two components: i) a time-aware spanning tree algorithm to generate the cooperation result for a group of couriers; and ii) a capacity-constrained courier grouping algorithm to optimize the courier grouping result iteratively. Extensive evaluation results with real-world order data collected from one of the largest logistics companies show that InCo improves the average saved delivery time and reduces average travel time by up to 80.2% and 28.4%, respectively, compared to baseline methods. The deployment results show InCo improves the average courier working efficiency by 21.6% to the state-of-the-practice. Wenjun Lyu, Shuxin Zhong, Guang Yang 0028, Haotian Wang 0008, Yi Ding 0011, Shuai Wang 0008, Yunhuai Liu, Tian He 0001, Desheng Zhang 0002 |
WWW | 8 |
| 2025 | HPST-GT: Full-Link Delivery Time Estimation Via Heterogeneous Periodic Spatial-Temporal Graph TransformerabstractA warehouse-distribution integration (WDI) e-commerce platform is an approach that combines warehousing and distribution processes, which is increasingly adopted in industry to enhance business efficiency. In the WDI e-commerce, one of the most important problems is to estimate the full-link delivery time for decision-making. Traditional methods designed for separate warehouse-distribution models struggle to address challenges in integrated systems. The difficulties stem from two main factors: (i) the contextual influence exerted by neighboring units within heterogeneous delivery networks, and (ii) the uncertainty in delivery times caused by dynamic and periodic temporal factors such as fluctuations in online sales volumes and the varying characteristics of different delivery units (e.g., warehouses and sorting centers). To address these challenges, we propose a novel full-link delivery time estimation framework calledHeterogeneousPeriodicSpatial-TemporalGraphTransformer (HPST-GT). First, we develop heterogeneous graph transformers to capture the hierarchical and diverse information of the warehouse-distribution network. Next, we design spatial-temporal transformers based on heterogeneous features to analyze the correlation between spatial and temporal information. Finally, we create a heterogeneous spatial-temporal graph prediction module to estimate full-link delivery time. Our method, evaluated on a one-month dataset from a leading e-commerce platform, surpasses current benchmarks across multiple performance metrics. Shuai Wang 0008, Hai Wang 0019, Li Lin 0011, Xiaohui Zhao 0006, Tian He 0001, Dian Shen, Wei Xi 0003 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | Hierarchical Information Propagation and Aggregation in Disentangled Graph Networks for Audience ExpansionabstractWith the development of the logistics industry, the user base of logistics services has expanded swiftly. This rapid increase in user scale presents significant challenges for logistics business management. A fundamental issue in such scenarios is audience expansion, which aims to find users willing to sign long-term services with logistics companies to foster business growth. Existing methods in addressing audience expansion mainly assume user modeling is entangled and neglects the inherent community structure among users. Due to these limitations, the effectiveness of traditional methods in achieving accurate user expansion is often restricted. Our work introduces a novel heterogeneous graph-based model, named Hi-DGN, which concentrates on the Hierarchical information propagation and aggregation in Disentangled Graph Networks for audience expansion. It consists of three main components: (i) the disentangled embedding layer to decouple user representations into different aspects, enabling the extraction of differentiated features; (ii) the hierarchical information propagation module partitions individual nodes into distinct groups and propagates information from group nodes to individual nodes hierarchically to capture diverse granularity representations; and (iii) the aggregation module to fuse all relation-specific embeddings to generate global node embeddings. Extensive experiments on two real-world datasets demonstrate the effectiveness of our method in various evaluation settings. Li Lin 0011, Kaiwen Xia, Shuai Wang 0008, Desheng Zhang 0002, Tian He 0001 |
CIKM | 6 |
| 2024 | DIFN: A Dual Intention-aware Network for Repurchase Recommendation with Hierarchical Spatio-temporal FusionabstractRecommendation systems play a crucial role in both industrial applications and research fields, which target to understand user preferences and intentions to provide personalized services. Compared to conventional recommendations, repurchase recommendations aim to suggest suitable products to users that they used to buy based on their intention evolution. Existing research on product recommendation can mainly be divided into behavior sequence-based methods and graph-based methods. Although these methods represent user interests and preference features effectively, they still fail to model repurchase behaviors because (i) the environment causing repurchase intention change is neglected and (ii) the lack of feedback after purchasing makes it difficult to learn the impacts of diverse behaviors. To comprehensively consider these limitations, we design a D ual I ntention-aware F usion N etwork framework (DIFN) to understand the effects of environment and after-purchasing feedback on users' intentions. Firstly, a hierarchical graph-based multi-level relational attention module is designed to effectively extract basic user features and spatial features from complex environmental information. Then, we introduce a behavior intention module and a usage intention module for different types of feedback data. Finally, we propose a dual intention fusion network that effectively fuses user basic features with spatial attributes and user intention features with temporal attributes for recommendation. Comprehensive evaluations on real-world datasets show that our method exceeds state-of-the-art baselines, which show an average of 8.2% improvements in different metrics. Li Lin 0011, Hai Wang 0019, Tian He 0001, Desheng Zhang 0002, Shuai Wang 0008 |
CIKM | 4 |
| 2024 | DECO: Cooperative Order Dispatching for On-Demand Delivery with Real-Time Encounter DetectionabstractIn on-demand delivery,online orders are delivered by couriers from merchants to customers within a short time (e.g., 45 minutes). An important task is to provide an efficient order dispatching solution. Existing studies focus on scenarios with stable routing behavior using pre-determined courier-order matching before delivery while ignoring real-time dynamics during delivery. In this work, we leverage courier-courier encounter events as an opportunity to enable cooperative order dispatching (i.e., conducting order transfers among couriers during delivery) for better delivery efficiency. However, it is non-trivial to conduct encounter-aware cooperative order dispatching in real-time dynamics due to two major challenges: (i) the dynamic nature of encounters in diverse real-world scenarios, and (ii) global delivery efficiency optimization by local order transfers. To address the above challenges, we design a detection-driven cooperative dispatching framework, called DECO. Specifically, we design (i) a Received Signal Strength Indicator (RSSI) variance-based state encoder to model encounter dynamics, (ii) an encounter event selector to choose encounter scenarios, (iii) a time-constrained order mask module to filter unsuitable orders, and (iv) an encounter-aware order transfer scheduler to make detailed order transfer decisions. Extensive experiments on real-world data from two large companies (i.e., JD Logistics, Eleme) show that DECO outperforms other baselines.Real-world deployment results at JD Logistics show that DECO improves the order overdue rate by 4.8%. Shuai Wang 0008, Yu Yang 0010, Hai Wang 0019, Baoshen Guo, Desheng Zhang 0002, Shuai Wang 0021, Tian He 0001 |
CIKM | 8 |
| 2024 | Multi-task Conditional Attention Network for Conversion Prediction in Logistics AdvertisingabstractLogistics advertising is an emerging task in online-to-offline logistics systems, where logistics companies expand parcel shipping services to new users through advertisements on shopping websites. Compared to existing online e-commerce advertising, logistics advertising has two significant new characteristics: (i) the complex factors in logistics advertising considering both users' offline logistics preference and online purchasing profiles; and (ii) data sparsity and mutual relations among multiple steps due to longer advertising conversion processes. To address these challenges, we design MCAC, a Multi-task Conditional Attention network-based logistics advertising Conversion prediction framework, which consists of (i) an offline shipping preference extraction model to extract the user's offline logistics preference from historical shipping records, and (ii) a multi-task conditional attention-based conversion rate prediction module to model mutual relations among multiple steps in logistics advertising conversion processes. We evaluate and deploy MCAC on one of the largest e-commerce platforms in China for logistics advertising. Extensive offline experiments show that our method outperforms state-of-the-art baselines in various metrics. Moreover, the conversion rate prediction results of large-scale online A/B testing show that MCAC achieves a 15.22% improvement compared to existing industrial practices, which demonstrates the effectiveness of the proposed framework. Baoshen Guo, Xining Song, Shuai Wang 0008, Wei Gong 0001, Tian He 0001, Xue (Steve) Liu |
KDD | 5 |
| 2023 | AutoBuild: Automatic Community Building Labeling for Last-mile DeliveryabstractFine-grained community-building information, such as building names and accurate geographical coordinates, is critical for a range of practical applications like navigation and door-to-door services (e.g., on-demand delivery and last-mile delivery). A common practice of traditional methods to gather community-building information usually relies on manual collection, which is typically labor-intensive and time-consuming. To address these issues, we utilize the massive data generated from e-commerce delivery services and design a framework, AutoBuild, for fine-grained large-scale community-building labeling. AutoBuild consists of two main components: (i) a Location Candidate Detection Module that identifies potential building names and coordinates from multi-source delivery data, and (ii) a Progressive Building Matching Model that employs trajectory modeling, human behavior analysis, and heterogeneous graph alignment to match building names and coordinates. To evaluate the performance of AutoBuild, we applied it to two real-world multi-modal datasets from Beijing City and Chengdu City. The results reveal that AutoBuild significantly outperforms multiple baseline models by 50-meter accuracy of 81.8% and 100-meter accuracy of 95.9% in Beijing City. More importantly, we conduct a real-world case study to demonstrate the practical impact of AutoBuild in last-mile delivery. Zhiqing Hong, Dongjiang Cao, Haotian Wang 0008, Guang Wang 0001, Tian He 0001, Desheng Zhang 0002 |
CIKM | 5 |
| 2023 | Urban-scale POI Updating with Crowd IntelligenceabstractPoints of Interest (POIs), such as entertainment, dining, and living, are crucial for urban planning and location-based services. However, the high dynamics and expensive updating costs of POIs pose a key roadblock for their urban applications. This is especially true for developing countries, where active economic activities lead to frequent POI updates (e.g., merchants closing down and new ones opening). Therefore, POI updating, i.e., detecting new POIs and different names of the same POIs (alias) to update the POI database, has become an urgent but challenging problem to address. In this paper, we attempt to answer the research question of how to detect and update large-scale POIs via a low-cost approach. To do so, we propose a novel framework called UrbanPOI, which formulates the POI updating problem as a tagging and detection problem based on multi-modal logistics delivery data. UrbanPOI consists of two key modules: (i) a hierarchical POI candidate generation module based on the POINet model that detects POIs from shipping addresses; and (ii) a new POI detection module based on the Siamese Attention Network that models multi-modal data and crowd intelligence. We evaluate our framework on real-world logistics delivery datasets from two Chinese cities. Extensive results show that our model outperforms state-of-the-art models in Beijing City by 26.2% in precision and 10.7% in F1-score, respectively. Zhiqing Hong, Haotian Wang 0008, Wenjun Lyu, Hai Wang 0019, Yunhuai Liu, Guang Wang 0001, Tian He 0001, Desheng Zhang 0002 |
CIKM | 7 |
| 2023 | HST-GT: Heterogeneous Spatial-Temporal Graph Transformer for Delivery Time Estimation in Warehouse-Distribution Integration E-CommerceabstractWarehouse-distribution integration has been adopted by many e-commerce retailers (e.g., Amazon, TAOBAO, and JD) as an efficient business mode. In warehouse-distribution integration e-commerce, one of the most important problems is to estimate the full-link delivery time for better decision-making. Existing solutions for traditional warehouse-distribution separation mode are challenging to address this problem due to two unique features in the integration mode including (i) contextual influence caused by neighbor units in heterogeneous delivery networks, (ii) uncertain delivery time caused by the dynamic temporal data (e.g., online sales volume) and heterogeneity of delivery units. To incorporate these new factors, we propose Heterogeneous Spatial-Temporal Graph Transformer (HST-GT), a novel full-link delivery time estimation method under the warehouse-distribution integration mode, where we (i) develop heterogeneous graph transformers to capture hierarchical heterogeneous information; and (ii) design a set of spatial-temporal transformers based on heterogeneous features to fully exploit the correlation of spatial and temporal information. We extensively evaluate our method based on one-month real-world data consisting of hundreds of warehouses and sorting centers, and millions of historical orders collected from one of the largest e-commerce retailers in the world. Experimental results demonstrate that our method outperforms state-of-the-art baselines in various metrics. Xiaohui Zhao 0006, Shuai Wang 0008, Hai Wang 0019, Tian He 0001, Desheng Zhang 0002, Guang Wang 0001 |
CIKM | 4 |
| 2023 | Towards Equitable Assignment: Data-Driven Delivery Zone Partition at Last-mile LogisticsabstractThe popularity of online e-commerce has promoted the rapid development of last-mile logistics in recent years. In last-mile services, to ensure delivery efficiency and enhance user experience, the delivery zone is proposed to perform delivery task assignment, which is a fundamental part of last-mile delivery. Each courier is responsible for one delivery zone. Couriers will collect orders belonging to their delivery zones from the delivery station and deliver orders to customers. Existing delivery zone partition practices in last-mile logistics consist of manual experience-based and static optimization-based methods, which perform order amount balancing among different zone but suffer from dissatisfaction and inefficiency because of two limitations: (i) using order amount is not always a good balancing metric considering deliveries' various difficulties (e.g., residence or industrial park, with or without elevators); (ii) less considering couriers' familiarity and preference behaviors. To generate delivery zone partition with equitable workload assignment, in this paper, we propose E-partition, a data-driven delivery zone partition framework to achieve equitable workload assignment in last-mile logistics. We first design a learning-based workload prediction model to estimate service time given a partition plan that consists of unseen courier-zone matching scenarios. Then, a delivery zone partition algorithm is proposed to iterative optimize couriers' core-AOI (i.e., area of interest) generation and AOI assignment process. Extensive offline experimental results show that our model outperforms baselines in working time prediction and workload balancing performances. Real-world deployment results at JD Logistics also verify the effectiveness of equitable-assignment aware delivery zone partition, with a 2.2% increase in service on-time rate compared to state-of-practice partition solutions. Baoshen Guo, Shuai Wang 0008, Haotian Wang 0008, Yunhuai Liu, Fanshuo Kong, Desheng Zhang 0002, Tian He 0001 |
KDD | 7 |
| 2023 | A Predict-Then-Optimize Couriers Allocation Framework for Emergency Last-mile LogisticsabstractIn recent years, emergency last-mile logistics (ELML) have played an essential role in urban emergencies. The efficient allocation of couriers in ELML is of practical significance to ensure the supply of essential materials, especially in public health emergencies (PHEs). However, couriers allocation becomes challenging due to the instability of demand, dynamic supply comprehension, and the evolutional delivery environment for ELML caused by PHEs. While existing work has delved into couriers allocation, the impact of PHEs on demand-supply-delivery has yet to be considered. In this work, we design PTOCA, a Predict-Then-Optimize Couriers Allocation framework. Specifically, in the prediction stage, we design a resource-aware prediction module that performs spatio-temporal modeling of unstable demand characteristics using a variational graph GRU encoder and builds a task-resource regressor to predict demand accurately. In the optimization stage, firstly, the priority ranking module solves the matching of delivery resources under demand-supply imbalance. Then the multi-factor task allocation module is used to model the dynamic evolutional environment and reasonably assign the delivery tasks of couriers. We evaluate PTOCA using real-world data covering 170 delivery zones, more than 10,000 couriers, and 100 million delivery tasks. The data is collected from JD Logistics, one of the largest logistics service companies. Extensive experimental results show that our method outperforms the baseline in task delivery rate and on-time delivery rate. Kaiwen Xia, Li Lin 0011, Shuai Wang 0008, Haotian Wang 0008, Desheng Zhang 0002, Tian He 0001 |
KDD | 6 |
| 2023 | eShare+: A Data-Driven Balancing Mechanism for Bike Sharing Systems Considering Both Quality of Service and MaintenanceabstractWith the rapid development of sharing economy, we have access to massive sharing systems such as Uber, Airbnb, and bike sharing nowadays. The sharing economy, at its core, is to achieve efficient use of resources. However, the actual usage of shared resources is still unclear to us. Little measurement or analysis, if any, has been conducted to investigate the resource usage patterns with the large-scale data collected from these sharing systems. In this paper, we first analyze the shared bike usage patterns in three typical bike sharing systems based on 140-month multi-event data. From our data-driven analysis, we found that the most used 20% of shared bikes account for 45% of total usage, while the least used 20% of bikes account for less than 1% of usage. To efficiently utilize shared bikes, we propose a usage balancing design called eShare+ to improve the bike sharing systems by considering both the quality of service and bike maintenance, which includes three key components: (i) a statistical model based on archived data to infer historical usage; (ii) an entropy and contextual LSTM-based prediction model with both real-time and archived data to infer future usage; (iii) a model-driven optimal calibration engine for bike selection to dynamically balance usage. We develop an ID swapping-based evaluation methodology to measure the efficiency of eShare+ with data from three large-scale bike sharing systems including 84,000 bikes and 3,300 service stations. Our results show that eShare+ not only fully utilizes shared bikes with efficient maintenance but also improves the quality of service. In addition, eShare+ also has the potential to be applicable to different fleet sizes. Shuai Wang 0008, Xin Zhu 0007, Guang Wang 0001, Yunhuai Liu, Tian He 0001, Desheng Zhang 0002 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2023 | $\mathrm{W}^{2}$Parking: A Data-Driven Win-Win Contract Parking Sharing Mechanism Under Both Supply and Demand UncertaintiesabstractWith the rapid growth of the number of private vehicles, searching for accessible parking spaces becomes intractable for drivers, especially during high-demand hours. In recent years, we are witnessing a number of sharing economy services. Contract parking sharing, as an innovative sharing economy mode, has the potential to alleviate the difficult parking issue and make full use of the urban parking resources. However, the uncertainties of both drivers’ parking demand and owners’ sharing supply make it challenging to achieve efficient sharing. Thanks to IoT technology, many current parking lots now record vehicles’ fine-grained parking data for billing purposes. Leveraging these fine-grained parking data, we exploit available contract parking spaces to share them with drivers that have temporary parking demand. Specifically, we propose$\mathrm{W^{2}}$Parking, awin-win contractparkingsharing system, which includes two key components: (i) an idle time prediction model to estimate available periods of parking spaces and (ii) a parking sharing model to schedule temporary users to have access to these available parking spaces under both demand and supply uncertainties using dynamic programming combined with a 2-approximation algorithm with performance-bound guarantees. we evaluate our system on seven-month real-world parking data from 368 parking lots with 14,704 parking spaces. Extensive experimental results show that our$\mathrm{W^{2}}$Parking achieves more than 90% of accuracy in parking time prediction, and the utilization rate of contract parking spaces is improved by 35%. Shuai Wang 0008, Xin Zhu 0007, Guang Wang 0001, Desheng Zhang 0002, Lai Tu, Tian He 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2022 | $O^{2}$-SiteRec: Store Site Recommendation under the O2O Model via Multi-graph Attention NetworksabstractThe emergence of Online-to-Offline (O2O) stores based on delivery platforms (e.g., Uber Eats, DoorDash, and Eleme) provides great convenience to people's lives. In the O2O model, one of the essential problems for merchants is to select a suitable store site, i.e., store site recommendation problem. We argue that the existing works for the traditional brick-and-mortar stores cannot address this problem due to two unique factors in the O2O model including (i) dynamic supply caused by courier capacity and dispatching strategies and (ii) various customer demands caused by delivery distance and customer preferences. To incorporate these new factors, we design$O^{2}$SiteRec, a store site recommendation method under the O2O model via multi-graph attention networks, which consists of (i) a courier capacity model based on a multi-semantic relation graph attention network to capture courier capacity; (ii) a heterogeneous multi-graph based recommendation model, where the courier capacity, customer preferences, and context features are fused. We evaluate our method based on one-month real-world data consisting of 39,465 stores and 23.6 million orders from one of the largest O2O platforms in China. Experimental results demonstrate that our method outperforms state-of-the-art baselines in various metrics. Shuai Wang 0008, Yu Yang 0010, Baoshen Guo, Tian He 0001, Desheng Zhang 0002 |
ICDE | 5 |
| 2019 | Towards Efficient Sharing: A Usage Balancing Mechanism for Bike Sharing SystemsabstractWith the rapid development of sharing economy, massive sharing systems such as Uber, Airbnb, and bikeshare have percolated into people's daily life. The sharing economy, at its core, is to achieve efficient use of resources. The actual usage of shared resources, however, is unclear to us. Little measurement or analysis, if any, has been conducted to investigate the resource usage status with the large-scale data collected from these sharing systems. In this paper, we analyze the bike usage status in three typical bikeshare systems based on 140-month multi-event data. Our analysis shows that the most used 20% of bikes account for 45% of usage, while the least used 20% of bikes account for less than 1% of usage. To efficiently utilize shared bikes, we propose a usage balancing design called eShare which has three components: (i) a statistical model based on archived data to infer historical usage; (ii) an entropy-based prediction model based on both real-time and archived data to infer future usage; (iii) a model-driven optimal calibration engine for bike selection to dynamically balance usage. We develop an ID swapping based evaluation methodology and measure the efficiency of eShare with data from three systems including the world's largest bikeshare system with 84,000 bikes and 3,300 stations. Our results show that eShare not only fully utilizes shared bikes but also improves service quality. Shuai Wang 0008, Tian He 0001, Desheng Zhang 0002, Yunhuai Liu, Sang Hyuk Son |
WWW | 2 |
| 2018 | Real-Time Human Mobility Modeling with Multi-View LearningabstractReal-time human mobility modeling is essential to various urban applications. To model such human mobility, numerous data-driven techniques have been proposed. However, existing techniques are mostly driven by data from a single view, for example, a transportation view or a cellphone view, which leads to over-fitting of these single-view models. To address this issue, we propose a human mobility modeling technique based on a generic multi-view learning framework called coMobile. In coMobile, we first improve the performance of single-view models based on tensor decomposition with correlated contexts, and then we integrate these improved single-view models together for multi-view learning to iteratively obtain mutually reinforced knowledge for real-time human mobility at urban scale. We implement coMobile based on an extremely large dataset in the Chinese city Shenzhen, including data about taxi, bus, and subway passengers along with cellphone users, capturing more than 27 thousand vehicles and 10 million urban residents. The evaluation results show that our approach outperforms a single-view model by 51% on average. More importantly, we design a novel application where urban taxis are dispatched based on unaccounted mobility demand inferred by coMobile. Desheng Zhang 0002, Tian He 0001, Fan Zhang 0019 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2017 | Hybrid Bicycle Allocation for Usage Load Balancing and Lifetime Optimization in Bike-Sharing SystemsabstractNowadays, public bike-sharing systems are broadly adopted and deployed in many major cities, however, as public facilities, bicycles will be prone to damage and need to be replaced frequently, which results in high system maintenance costs. One of the root causes of bicycle damages is the serious load-unbalance of bicycle usage. In this paper, we propose a hybrid bicycle allocation strategy for bicycle lifetime optimization, which can effectively reduce the degree of imbalance of system load. First, we analyze and verify the load-unbalance status of bicycle usage in the existing bike-sharing system. Then, a hybrid bicycle allocation strategy is proposed, which is evaluated on Washington D.C. bike-sharing system. Furthermore, according to a bicycle lifetime model based on Weibull distribution, the proposed bicycle allocation strategy could significantly cut down the percentage of the bicycles need to be replaced in a certain period of time. Xiawen Yao, Xingfa Shen, Landi Wang, Tian He 0001 |
MDM | 4 |
| 2016 | MultiCalib: national-scale traffic model calibration in real time with multi-source incomplete dataabstractReal-time traffic modeling at national scale is essential to many applications, but its calibration is extremely challenging due to its large spatial and fine temporal coverage. The existing work mostly is focused on urban-scale calibration with complete field data from single data sources (e.g., loop sensors or taxis), which cannot be generalized to national scale, because complete single-source field data at national scale are almost impossible to obtain. To address this challenge, in this paper, we design MultiCalib, a model calibration framework to optimize traffic models based on multiple incomplete data sources at national scale in real time. Instead of naively combining multi-source data, we theoretically formulate a multi-source model calibration problem based on real-world contexts and multi-view learning. More importantly, we implement and evaluate MultiCalib with two heterogeneous nationwide vehicle networks with 340,000 vehicles to infer traffic conditions on 36 expressways and 119 highways, along with 4 cities across China. The results show that MultiCalib outperforms state-of-the- art calibration by 25% on average with same input data. Desheng Zhang 0002, Fan Zhang 0019, Tian He 0001 |
SIGSPATIAL/GIS | 3 |
| 2015 | EveryoneCounts: Data-driven digital advertising with uncertain demand model in metro networksabstractNowadays most metro advertising systems schedule advertising slots on digital advertising screens to achieve the maximum exposure to passengers by exploring passenger demand models. However, our empirical results show that these passenger demand models experience uncertainty at fine temporal granularity (e.g., per min). As a result, for fine-grained advertisements (shorter than one minute), a scheduling based on these demand models cannot achieve the maximum advertisement exposure. To address this issue, we propose an online advertising approach, called EveryoneCounts, based on an uncertain passenger demand model. It combines coarse-grained statistical demand modeling and fine-grained Bayesian demand modeling by leveraging realtime card-swiping records along with both passenger mobility patterns and travel periods within metro systems. Based on this uncertain demand model, it schedules advertising time online based on robust receding horizon control to maximize the advertisement exposure. We evaluate the proposed approach based on an one-month sample from our 530 GB real-world metro fare dataset with 16 million cards. The results show that our approach provides a 61.5% lower traffic prediction error and 20% improvement on advertising efficiency on average. Desheng Zhang 0002, Ruobing Jiang, Shuai Wang 0008, Yanmin Zhu 0006, Bo Yang 0006, Jian Cao 0001, Fan Zhang 0019, Tian He 0001 |
IEEE BigData | 8 |
| 2015 | coMobile: real-time human mobility modeling at urban scale using multi-view learningabstractReal-time human mobility modeling is essential to various urban applications. To model such human mobility, numerous data-driven techniques have been proposed. However, existing techniques are mostly driven by data from a single view, e.g., a transportation view or a cellphone view, which leads to over-fitting of these single-view models. To address this issue, we propose a human mobility modeling technique based on a generic multi-view learning framework called coMobile. In coMobile, we first improve the performance of single-view models based on tensor decomposition with correlated contexts, and then we integrate these improved single-view models together for multi-view learning to iteratively obtain mutually-reinforced knowledge for real-time human mobility at urban scale. We implement coMobile based on an extremely large dataset in the Chinese city Shenzhen, including data about taxi, bus and subway passengers along with cellphone users, capturing more than 27 thousand vehicles and 10 million urban residents. The evaluation results show that our approach outperforms a single-view model by 51% on average. Desheng Zhang 0002, Juanjuan Zhao 0001, Fan Zhang 0019, Tian He 0001 |
SIGSPATIAL/GIS | 4 |
| 2013 | CallCab: A unified recommendation system for carpooling and regular taxicab servicesabstractCarpooling taxicab services hold the promise of providing additional transportation supply, especially in extreme weather or rush hour when regular taxicab services are insufficient. Although many recommendation systems about regular taxicab services have been proposed recently, little research, if any, has been done to assist passengers to find a successful taxicab ride with carpooling. In this paper, we present the first systematic work to design a unified recommendation system for both regular and carpooling services, called CallCab, based on a data driven approach. In response to a passenger's request, CallCab aims to recommend either (i) a vacant taxicab for a regular service with no detour, or (ii) an occupied taxicab heading to the similar direction for a carpooling service with less detour, yet without assuming any knowledge of destinations of passengers already on occupied taxicabs. To analyze these unknown destinations of occupied taxicabs, CallCab generates and refines taxicab trip distributions based on GPS datasets and context information collected in the existing taxicab infrastructure. To improve CallCab's efficiency to process such a big dataset, we augment the efficient MapReduce model with a Measure phase tailored for our application. We evaluate CallCab with a real world dataset of 14,000 taxicabs, and results show that compared to ground truth, CallCab can reduce 64% of the total mileage to deliver all passengers and 63% of passenger's waiting time. Desheng Zhang 0002, Tian He 0001, Yunhuai Liu, John A. Stankovic |
IEEE BigData | 2 |
| 2009 | Aggregate Location Monitoring for Wireless Sensor Networks: A Histogram-Based ApproachabstractLocation monitoring systems are used to detect human activities and provide monitoring services, e.g., aggregate queries. In this paper, we consider an aggregate location monitoring system where wireless sensor nodes are counting sensors that are only capable of detecting the number of objects within their sensing areas. As traditional query processors rely on the knowledge of users' exact locations, they cannot provide any monitoring services based on the readings reported from counting sensors. To this end, we propose an adaptive spatio-temporal histogram to enable monitoring services without the need of users' exact locations. The main idea of the histogram is to keep statistics about the distribution of moving objects. At the core of the histogram, we propose three techniques, memorization, locality awareness and packing, to improve monitoring accuracy and efficiency. Furthermore, the histogram is designed in a way that achieves a trade-off between the energy and bandwidth consumption of the sensor network and the accuracy of monitoring services. Experimental results show that the proposed histogram provides high-quality location monitoring services (i.e., 90% accuracy for both skewed and uniform mobility patterns) and outperforms a basic histogram and the state-of-the-art spatio-temporal histogram by two orders of magnitude in most cases. Chi-Yin Chow, Mohamed F. Mokbel, Tian He 0001 |
Mobile Data Management | 3 |
| 2008 | Tinycasper: a privacy-preserving aggregate location monitoring system in wireless sensor networksabstractThis demo presents a privacy-preserving aggregate location monitoring system, namely, TinyCasper, in which we can monitor moving objects in wireless sensor networks while preserving their location privacy. TinyCasper consists of two main modules, in-network location anonymization and aggregate query processing over anonymized locations. In the first module, trusted wireless sensor nodes collaborate with each other to anonymize users' exact locations by a cloaked spatial region that satisfies a prespecified privacy requirement. On the other side, the aggregate query processing module collects and analyzes the cloaked spatial regions reported from the wireless sensor nodes to support aggregate and alarm queries over anonymized locations. The prototype of TinyCasper is implemented on a physical test-bed on the TinyOS/Mote platform with 39 MICAz motes. Chi-Yin Chow, Mohamed F. Mokbel, Tian He 0001 |
SIGMOD Conference | 3 |