VLDB 2026 Research / reviewers in the wild / expert
Hongtu Zhu
dblp:03/5683
· DBLP profile ↗
15ranked-venue papers in the field
0as first author
9since 2021 · last 2025
0000-0002-6781-2690ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 10Information Retrieval & Web Search · 3Big Data, Cloud & Distributed Data Systems · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TSMO 2025: Two-sided Marketplace Optimization: Search, Discovery, Matching, Pricing & GrowthabstractIn recent years, two-sided marketplaces have emerged as viable business models in many real-world applications. In particular, we have moved from the social network paradigm to a network with two distinct types of participants representing the supply and demand of a specific good. Examples of industries include but are not limited to accommodation (Airbnb, Booking.com), video content (YouTube, Instagram, TikTok), ridesharing (Uber, Lyft), online shops (Etsy, Ebay, Facebook Marketplace), music (Spotify, Amazon), app stores (Apple App Store, Google App Store) or job sites (LinkedIn). The traditional research in most of these industries focused on satisfying the demand. OTAs would sell hotel accommodation, TV networks would broadcast their own content, or taxi companies would own their own vehicle fleet. In modern examples like Airbnb, YouTube, Instagram, or Uber, the platforms operate by outsourcing the service they provide to their users, whether they are hosts, content creators or drivers, and have to develop their models considering their needs and goals. Mihajlo Grbovic, Vladan Radosavljevic, Rui Song 0006, Minmin Chen, Zhiwei (Tony) Qin, Katerina Iliakopoulou-Zanos, Thanasis Noulas, Hongtu Zhu, Fabrizio Silvestri |
KDD (2) | 9 |
| 2025 | Sampling-guided Heterogeneous Graph Neural Network with Temporal Smoothing for Scalable Longitudinal Data ImputationabstractIn this paper, we propose a novel framework, the Sampling-guided Heterogeneous Graph Neural Network (HT-GNN), to effectively tackle the challenge of missing data imputation in longitudinal studies. Unlike traditional methods, which often require extensive preprocessing to handle irregular or inconsistent missing data, our approach accommodates arbitrary missing data patterns while maintaining computational efficiency. HT-GNN models both observations and covariates as distinct node types, connecting observation nodes at successive time points through subject-specific longitudinal subnetworks, while covariate-observation interactions are represented by attributed edges within bipartite graphs. By leveraging subject-wise mini-batch sampling and a multi-layer temporal smoothing mechanism, HT-GNN efficiently scales to large datasets, while effectively learning node representations and imputing missing data. Extensive experiments on both synthetic and real-world datasets, including the Alzheimer's Disease Neuroimaging Initiative (DNI) dataset, demonstrate that HT-GNN significantly outperforms existing imputation methods, even with high missing data rates (e.g., 80%). The empirical results highlight HT-GNN's robust imputation capabilities and superior performance, particularly in the context of complex, large-scale longitudinal data. Ziqi Chen 0002, Qiao Liu 0008, Jinhan Xie, Hongtu Zhu |
KDD (2) | 5 |
| 2023 | KDD-2023 Workshop on Decision Intelligence and Analytics for Online MarketplacesabstractOnline marketplace is a digital platform that connects buyers (demand) and sellers (supply) and provides exposure opportunities that individual participants would not otherwise have access to. Online marketplaces exist in a diverse set of domains and industries, for example, rideshare (Lyft, DiDi, Uber), house rental (Airbnb), real estate (Beke), online retail (Amazon, Ebay), and food ordering and delivery (Doordash, Meituan). Besides academia, many companies and institutions are researching on topics specific to their particular domains. The fundamental mechanism of an online marketplace is to match supply and demand to generate transactions, with objectives considering service quality, participants experience, financial and operational efficiency. It is valuable to bring together researchers and practitioners from different application domains to discuss their experiences, challenges, and opportunities to leverage cross-domain knowledge. The goal of this workshop is to offer an opportunity to appreciate the diversity in applications, to draw connections to inform decision optimization across different industries, and to discover new problems that are fundamental to marketplaces of different domains. The previous version of this workshop at KDD-2022 was a tremendous success in terms of participation, technical contribution, and community interest. This updated version of the workshop is especially timely to cover the issues and algorithms pertinent to general online marketplaces, specific problems and applications arising from those diverse domains, as well as emerging topics such as competition and resilience to market condition shifts. Zhiwei (Tony) Qin, Rui Song 0006, Jieping Ye, Hongtu Zhu, Michael I. Jordan |
KDD | 4 |
| 2023 | DNet: Distributional Network for Distributional Individualized Treatment EffectsabstractThere is a growing interest in developing methods to estimate individualized treatment effects (ITEs) for various real-world applications, such as e-commerce and public health. This paper presents a novel architecture, called DNet, to infer distributional ITEs. DNet can learn the entire outcome distribution for each treatment, whereas most existing methods primarily focus on the conditional average treatment effect and ignore the conditional variance around its expectation. Additionally, our method excels in settings with heavy-tailed outcomes and outperforms state-of-the-art methods in extensive experiments on benchmark and real-world datasets. DNet has also been successfully deployed in a widely used mobile app with millions of daily active users. Guojun Wu, Xiaoxiang Lv, Shikai Luo, Chengchun Shi, Hongtu Zhu |
KDD | 6 |
| 2022 | Decision Intelligence and Analytics for Online Marketplaces: Jobs, Ridesharing, Retail and BeyondabstractOnline marketplace is a digital platform that connects buyers (demand) and sellers (supply) and provides exposure opportunities that individual participants would not otherwise have access to. Online marketplaces exist in a diverse set of domains and industries, for example, rideshare (Lyft, DiDi, Uber), house rental (Airbnb), real estate (Beke), online retail (Amazon, Ebay), job search (LinkedIn, Indeed.com, CareerBuilder), and food ordering and delivery (Doordash, Meituan). Besides academia, many companies and institutions are researching on topics specific to their particular domains. The fundamental mechanism of an online marketplace is to match supply and demand to generate transactions, with objectives considering service quality, participants experience, financial and operational efficiency. It is valuable to bring together researchers and practitioners from different application domains to discuss their experiences, challenges, and opportunities to leverage cross-domain knowledge. The goal of this workshop is to offer an opportunity to appreciate the diversity in applications, to draw connections to inform decision optimization across different industries, and to discover new problems that are fundamental to marketplaces of different domains. This workshop will follow a dual-track format. Track 1 covers the issues and algorithms pertinent to general online marketplaces as well as specific problems and applications arising from those diverse domains, such as ridesharing, online retail, food delivery, house rental, real estate, and more. Track 2 focuses on the state of the art advances in the computational jobs marketplace. Interesting challenges in this domain include the drastic increase of work from home or remote work, the imbalance between the demand and supply of the job market, the popularity of independent workers, the capability of helping job seekers on their whole job seeking journey and career development, the different objectives and behaviors of all major stakeholders in the ecosystem, e.g. job seekers, employers, recruiters and job agents. Zhiwei (Tony) Qin, Liangjie Hong, Rui Song 0006, Hongtu Zhu, Mohammed Korayem, Haiyan Luo, Michael I. Jordan |
KDD | 4 |
| 2022 | Intrinsic partial linear models for manifold-valued data
Shihui Ying, Hongtu Zhu |
Inf. Process. Manag. | 3 |
| 2021 | Optimizing Bike-Share Repositioning: Networked Inventory Management with Spatiotemporal ModelingabstractIn a bike-sharing system, demand loss is primarily due to out-of-stock stations. One solution to tackle this problem is to rebalance the bike inventory of the stations through repositioning. In a large-scale bike-sharing system, bike repositioning typically works in two steps: The platform generates the reposition tasks, and then the operators execute those tasks. In this paper, we focus on the problem of determining the reposition tasks, which is a core problem for the platform operations. We model the bike-sharing system as a networked inventory system and propose a select-and-match method to generate optimal repositions, seamlessly combining inventory management and spatiotemporal value learning techniques. We compute the optimal (s, S) policy to select the stations with excess and deficit bike stocks. Subsequently, we solve a minimum cost maximum flow problem for rebalance assignment while considering the long-term effects of repositioning, characterized by a bike transition value. We design a spatiotemporal value network to learn the values. We evaluate the performance of the proposed method in a simulator based on real bike-sharing data, and the results demonstrate the superiority of our approach over the other alternatives. Chunyi Liu, Zhiwei (Tony) Qin, Hongtu Zhu |
IEEE BigData | 5 |
| 2021 | Multi-Objective Distributional Reinforcement Learning for Large-Scale Order DispatchingabstractThe aim of this paper is to develop a multi-objective distributional reinforcement learning framework for improving order dispatching on large-scale ride-hailing platforms. Compared with traditional RL-based approaches that focus on drivers’ income, the proposed framework also accounts for the spatiotemporal difference between the supply and demand networks. Specifically, we model the dispatching problem as a two-objective Semi-Markov Decision Process (SMDP) and estimate the relative importance of the two objectives under some unknown existing policy via Inverse Reinforcement Learning (IRL). Then, we combine Implicit Quantile Networks (IQN) with the traditional Deep Q-Networks (DQN) to jointly learn the two return distributions and adjusting their weights to refine the old policy through on-line planning and achieve a higher supply-demand coherence of the platform. We conduct large-scale dispatching experiments to demonstrate the remarkable improvement of proposed approach on the platform’s efficiency. Fan Zhou 0003, Chenfan Lu, Xiaocheng Tang, Fan Zhang 0098, Zhiwei (Tony) Qin, Jieping Ye, Hongtu Zhu |
ICDM | 7 |
| 2021 | Value Function is All You Need: A Unified Learning Framework for Ride Hailing PlatformsabstractLarge ride-hailing platforms, such as DiDi, Uber and Lyft, connect tens of thousands of vehicles in a city to millions of ride demands throughout the day, providing great promises for improving transportation efficiency through the tasks of order dispatching and vehicle repositioning. Existing studies, however, usually consider the two tasks in simplified settings that hardly address the complex interactions between the two, the real-time fluctuations between supply and demand, and the necessary coordinations due to the large-scale nature of the problem. In this paper we propose a unified value-based dynamic learning framework (V1D3) for tackling both tasks. At the center of the framework is a globally shared value function that is updated continuously using online experiences generated from real-time platform transactions. To improve the sample-efficiency and the robustness, we further propose a novel periodic ensemble method combining the fast online learning with a large-scale offline training scheme that leverages the abundant historical driver trajectory data. This allows the proposed framework to adapt quickly to the highly dynamic environment, to generalize robustly to recurrent patterns and to drive implicit coordinations among the population of managed vehicles. Extensive experiments based on real-world datasets show considerably improvements over other recently proposed methods on both tasks. Particularly, V1D3 outperforms the first prize winners of both dispatching and repositioning tracks in the KDD Cup 2020 RL competition, achieving state-of-the-art results on improving both total driver income and user experience related metrics. Xiaocheng Tang, Fan Zhang 0098, Zhiwei (Tony) Qin, Yansheng Wang, Dingyuan Shi, Bingchen Song, Yongxin Tong, Hongtu Zhu, Jieping Ye |
KDD | 8 |
| 2020 | A Joint Inverse Reinforcement Learning and Deep Learning Model for Drivers' Behavioral PredictionabstractUsers' behavioral predictions are crucially important for many domains including major e-commerce companies, ride-hailing platforms, social networking, and education. The success of such prediction strongly depends on the development of representation learning that can effectively model the dynamic evolution of user's behavior. This paper aims to develop a joint framework of combining inverse reinforcement learning (IRL) with deep learning (DL) regression model, called IRL-DL, to predict drivers' future behavior in ride-hailing platforms. Specifically, we formulate the dynamic evolution of each driver as a sequential decision-making problem and then employ IRL as representation learning to learn the preference vector of each driver. Then, we integrate drivers' preference vector with their static features (e.g., age, gender) and other attributes to build a regression model (e.g., LTSM-neural network) to predict drivers' future behavior. We use an extensive driver data set obtained from a ride-sharing platform to verify the effectiveness and efficiency of our IRL-DL framework, and results show that our IRL-DL framework can achieve consistent and remarkable improvements over models without drivers' preference vectors. Guojun Wu, Shikai Luo, Jieping Ye, Xiaohu Qie, Hongtu Zhu |
CIKM | 9 |
| 2019 | Origin-destination Flow Prediction with Vehicle Trajectory Data and Semi-supervised Recurrent Neural NetworkabstractOrigin-Destination (OD) flow data is an important instrument for traffic study and management. So far traditional ways like surveys or detectors are costly and only give limited availability of OD flows. Various statistical and stochastic models for OD flow estimation and prediction based on limited link volume data or automatic vehicle identification (AVI) data have been developed. However, smartphone-generated trajectory data has not been as much leveraged in this field, though the usage of smartphones in traveling is emerging in recent years. In this paper, we propose a semi-supervised deep learning based model that appropriately combines both AVI and smartphone trajectory data during training and is able to generate predictions of OD flows in an urban network solely based on the smartphone trajectory data at inference time. Our model can provide OD estimation and prediction services on larger spatial areas beyond the limited spatial coverage of AVI data. Tests of our model using real data have shown promising results, compared with an AVI input-dependent Kalman filter model. Potentially, our model can easily be embedded to a trajectory collecting platform and generate continuous real-time OD flow predictions online. Yintai Ma, Zhiwei (Tony) Qin, Henry X. Liu, Hongtu Zhu, Jieping Ye |
IEEE BigData | 6 |
| 2019 | CIKM 2019 Workshop on Artificial Intelligence in Transportation (AI in transportation)abstractData-enabled smart transportation has attracted a surge of interest from machine learning and data mining researchers nowadays due to the bloom of online ride-hailing industry and rapid development of autonomous driving. Large-scale high quality route data and trading data (spatiotemporal data) have been generated every day, which makes AI an urgent need and preferred solution for the decision making in intelligent transportation systems. While a large of amount of work have been dedicated to traditional transportation problems, they are far from satisfactory for the rising need. We propose a half-day workshop at CIKM 2019 for the professionals, researchers, and practitioners who are interested in mining and understanding big and heterogeneous data generated in transportation, and AI applications to improve the transportation system. We plan to have several invited talks from both academia and industry. This workshop would be organized by Shanghai Jiao Tong University, Didi Chuxing and Pennsylvania State University. Weinan Zhang 0001, Haiming Jin, Lingyu Zhang 0001, Hongtu Zhu, Zhenhui Jessie Li, Jieping Ye |
CIKM | 4 |
| 2019 | InBEDE: Integrating Contextual Bandit with TD Learning for Joint Pricing and Dispatch of Ride-Hailing PlatformsabstractFor both the traditional street-hailing taxi industry and the recently emerged on-line ride-hailing, it has been a major challenge to improve the ride-hailing marketplace efficiency due to spatio-temporal imbalance between the supply and demand, among other factors. Despite the numerous approaches to improve marketplace efficiency using pricing and dispatch strategies, they usually optimize pricing or dispatch separately. In this paper, we show that these two processes are in fact intrinsically interrelated. Motivated by this observation, we make an attempt to simultaneously optimize pricing and dispatch strategies. However, such a joint optimization is extremely challenging due to the inherent huge scale and lack of a uniform model of the problem. To handle the high complexity brought by the new problem, we propose InBEDE (Integrating contextual Bandit with tEmporal DiffErence learning), a learning framework where pricing strategies are learned via a contextual bandit algorithm, and the dispatch strategies are optimized with the help of temporal difference learning. The two learning components proceed in a mutual bootstrapping manner, in the sense that the policy evaluations of the two components are inter-dependent. Evaluated with real-world datasets of two Chinese cities from Didi Chuxing, an online ride-hailing platform, we show that the market efficiency of the ride-hailing platform can be significantly improved using InBEDE. Haipeng Chen 0001, Yan Jiao, Zhiwei (Tony) Qin, Xiaocheng Tang, Bo An 0001, Hongtu Zhu, Jieping Ye |
ICDM | 7 |
| 2019 | A Deep Value-network Based Approach for Multi-Driver Order DispatchingabstractRecent works on ride-sharing order dispatching have highlighted the importance of taking into account both the spatial and temporal dynamics in the dispatching process for improving the transportation system efficiency. At the same time, deep reinforcement learning has advanced to the point where it achieves superhuman performance in a number of fields. In this work, we propose a deep reinforcement learning based solution for order dispatching and we conduct large scale online A/B tests on DiDi's ride-dispatching platform to show that the proposed method achieves significant improvement on both total driver income and user experience related metrics. Xiaocheng Tang, Zhiwei (Tony) Qin, Fan Zhang 0098, Zhaodong Wang, Yintai Ma, Hongtu Zhu, Jieping Ye |
KDD | 7 |
| 2018 | Deep Reinforcement Learning with Knowledge Transfer for Online Rides Order DispatchingabstractRide dispatching is a central operation task on a ride-sharing platform to continuously match drivers to trip-requesting passengers. In this work, we model the ride dispatching problem as a Markov Decision Process and propose learning solutions based on deep Q-networks with action search to optimize the dispatching policy for drivers on ride-sharing platforms. We train and evaluate dispatching agents for this challenging decision task using real-world spatio-temporal trip data from the DiDi ride-sharing platform. A large-scale dispatching system typically supports many geographical locations with diverse demand-supply settings. To increase learning adaptability and efficiency, we propose a new transfer learning method Correlated Feature Progressive Transfer, along with two existing methods, enabling knowledge transfer in both spatial and temporal spaces. Through an extensive set of experiments, we demonstrate the learning and optimization capabilities of our deep reinforcement learning algorithms. We further show that dispatching policies learned by transferring knowledge from a source city to target cities or across temporal space within the same city significantly outperform those without transfer learning. Zhaodong Wang, Zhiwei (Tony) Qin, Xiaocheng Tang, Jieping Ye, Hongtu Zhu |
ICDM | 5 |