EDBT 2026 Demo / reviewers in the wild / expert
Ruohan Zhan
dblp:187/7770
· DBLP profile ↗
14ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-3426-2784ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Booking Funnel and Substitution-Aware User Behavior Modeling for Demand Prediction and Joint Room PricingabstractUser interactions on online travel platforms (OTPs) naturally follow a two-stage booking funnel, spanning hotel click from a multi-hotel listing page and booking conversion through room selection within the clicked hotel. Crucially, booking decisions are set-dependent: users select among multiple room types within the same hotel, where price changes of one option can shift demand to others. Two challenges arise in modeling such behavior: (i) the cascading booking-funnel behavior from hotel click to booking conversion, and (ii) intra-hotel substitution across room types. While accurate modeling of such behavior is essential for demand prediction and downstream optimization, existing methods typically ignore this structure by collapsing user behavior into a single stage and overlooking intra-hotel substitution effects, resulting in biased demand estimations and suboptimal decisions. In this paper, we propose BFSNet, a Booking Funnel and Substitution-aware neural network that decomposes booking demand into hotel click and booking conversion, enabling interpretable and stage-wise learning of user behavior. To capture the intra-hotel substitution effect, we introduce a mixing network-based substitution effect module that incorporates economic structural information by enforcing monotonic substitution relationships between the demand of a focal room and the prices of competing room types. To bridge prediction and downstream decision-making, we further develop a lightweight pricing procedure that distills the learned prediction model to efficiently evaluate candidate price vectors under operational constraints. Experiments on large-scale production data from a major online travel platform demonstrate consistent improvements in demand prediction accuracy and significant gains in downstream revenue. Zhikang Fan 0001, Pin Gao 0001, Ruohan Zhan, Shaowen Zhang, Xingrui Li, Jianpei Wen, Su Zhao, Wei Lin 0022 |
SIGIR | 4 |
| 2025 | Distributionally Robust Policy Learning under Concept DriftsabstractDistributionally robust policy learning aims to find a policy that performs well
under the worst-case distributional shift, and yet most existing methods for
robust policy learning consider the worst-case *joint* distribution of
the covariate and the outcome. The joint-modeling strategy can be unnecessarily conservative
when we have more information on the source of distributional shifts. This paper studies
a more nuanced problem --- robust policy learning under the *concept drift*,
when only the conditional relationship between the outcome and the covariate changes.
To this end, we first provide a doubly-robust estimator for evaluating
the worst-case average reward of a given policy under a set of perturbed conditional distributions.
We show that the policy value estimator enjoys asymptotic normality even if the nuisance parameters
are estimated with a slower-than-root-$n$ rate.
We then propose a learning algorithm that outputs the policy maximizing the
estimated policy value within a given policy class $\Pi$, and show
that the sub-optimality gap of the proposed algorithm is of the order
$\kappa(\Pi)n^{-1/2}$, where $\kappa(\Pi)$ is the entropy integral of $\Pi$ under the Hamming distance
and $n$ is the sample size. A matching lower bound is provided to show the optimality of the rate.
The proposed methods are implemented and evaluated in numerical studies,
demonstrating substantial improvement compared with existing benchmarks. Zhimei Ren, Ruohan Zhan, Zhengyuan Zhou |
ICML | 3 |
| 2024 | Statistical Properties of Robust SatisficingabstractThe Robust Satisficing (RS) model is an emerging approach to robust optimization, offering streamlined procedures and robust generalization across various applications. However, the statistical theory of RS remains unexplored in the literature. This paper fills in the gap by comprehensively analyzing the theoretical properties of the RS model. Notably, the RS structure offers a more straightforward path to deriving statistical guarantees compared to the seminal Distributionally Robust Optimization (DRO), resulting in a richer set of results. In particular, we establish two-sided confidence intervals for the optimal loss without the need to solve a minimax optimization problem explicitly. We further provide finite-sample generalization error bounds for the RS optimizer. Importantly, our results extend to scenarios involving distribution shifts, where discrepancies exist between the sampling and target distributions. Our numerical experiments show that the RS model consistently outperforms the baseline empirical risk minimization in small-sample regimes and under distribution shifts. Furthermore, compared to the DRO model, the RS model exhibits lower sensitivity to hyperparameter tuning, highlighting its practicability for robustness considerations. Yunbei Xu, Ruohan Zhan |
ICML | 3 |
| 2024 | Adaptively Learning to Select-Rank in Online PlatformsabstractRanking algorithms are fundamental to various online platforms across e-commerce sites to content streaming services. Our research addresses the challenge of adaptively ranking items from a candidate pool for heterogeneous users, a key component in personalizing user experience. We develop a user response model that considers diverse user preferences and the varying effects of item positions, aiming to optimize overall user satisfaction with the ranked list. We frame this problem within a contextual bandits framework, with each ranked list as an action. Our approach incorporates an upper confidence bound to adjust predicted user satisfaction scores and selects the ranking action that maximizes these adjusted scores, efficiently solved via maximum weight imperfect matching. We demonstrate that our algorithm achieves a cumulative regret bound of $O(d\sqrt{NKT})$ for ranking $K$ out of $N$ items in a $d$-dimensional context space over $T$ rounds, under the assumption that user responses follow a generalized linear model. This regret alleviates dependence on the ambient action space, whose cardinality grows exponentially with $N$ and $K$ (thus rendering direct application of existing adaptive learning algorithms – such as UCB or Thompson sampling – infeasible). Experiments conducted on both simulated and real-world datasets demonstrate our algorithm outperforms the baseline. Perry Dong, Ruohan Zhan, Zhengyuan Zhou |
ICML | 4 |
| 2024 | Estimating Treatment Effects under Recommender Interference: A Structured Neural Networks ApproachabstractRecommender systems are essential for content-sharing platforms by curating personalized content. To evaluate updates of recommender systems targeting content creators, platforms frequently engage in creatorside randomized experiments to assess their performance. These experiments help estimate treatment effects, defined as the difference in outcomes when a new (vs. the status quo) algorithm is deployed on the platform. We show that the standard difference-in-means estimator can lead to a biased treatment effect estimate. This bias can occur in either direction and may even cause a reversal in the estimated sign, leading to incorrect decision-making. This bias arises because of recommender interference, which occurs when treated and control creators compete for exposure through the recommender system. We propose a "recommender choice model" that captures how an item is chosen among a pool comprised of both treated and control content items. By combining a structural choice model with neural networks, the framework directly models the interference pathway in a microfounded way while accounting for rich viewer-content heterogeneity. This model enables counterfactual evaluations of treatment effects under alternative treatment assignments (e.g., all treated or all control) for both the entire population and specific subgroups. Within this modeling framework, we further construct a double/debiased estimator of the treatment effect that is consistent and asymptotically normal under regularity conditions. We demonstrate its empirical performance with a field experiment on Weixin short-video platform. Besides the standard creator-side experiment, we implement a costly blocked double-sided randomization design to obtain a benchmark estimate of the treatment effect without interference bias. We show that the proposed estimator significantly reduces bias in treatment effect estimates compared to the standard difference-in-means estimator. The full paper is available at https://arxiv.org/abs/2406.14380. Ruohan Zhan, Shichao Han, Zhenling Jiang |
EC | 1 |
| 2023 | ResAct: Reinforcing Long-term Engagement in Sequential Recommendation with Residual Actor
Wanqi Xue, Qingpeng Cai 0001, Ruohan Zhan, Peng Jiang 0002, Kun Gai, Bo An 0001 |
ICLR | 3 |
| 2023 | Fragility Index: A New Approach for Binary ClassificationabstractIn binary classification problems, many performance metrics evaluate the probability that some error exceeds a threshold. Nevertheless, they focus more on the probability and fail to capture the magnitude of the error, which evaluates how large this error exceeds the threshold. Capturing the magnitude of error is desired in many applications. For example, in detecting disease and predicting credit default, the magnitude of error illustrates the confidence in making the wrong prediction. We propose a novel metric, the Fragility Index (FI), to evaluate the performance of binary classifiers by capturing the magnitude of the error. FI alleviates the risk of misclassification by penalizing the large error greatly, which is seldom considered by standard metrics. Moreover, to strengthen the generalization ability and handle unseen samples, we adopt the framework of distributionally robust optimization and robust satisficing, which allows us to derive and control the maximum degree of fragility of the classifier when the distribution of samples shifts. We show that FI can be easily calculated and optimized for common probabilistic distance measures. Experiments with real datasets demonstrate the new insights brought by FI and the advantages of classifiers selected under FI, which always improve the robustness and reduce the risk of large errors as compared to classifiers selected by alternative metrics. Chen Yang 0025, Bo Cao 0007, Daniel Zhuoyu Long, Feng Wang 0023, Ruohan Zhan |
KDD | 10 |
| 2023 | Proportional Response: Contextual Bandits for Simple and Cumulative Regret MinimizationabstractIn many applications, e.g. in healthcare and e-commerce, the goal of a contextual bandit may be to learn an optimal treatment assignment policy at the end of the experiment. That is, to minimize simple regret. However, this objective remains understudied. We propose a new family of computationally efficient bandit algorithms for the stochastic contextual bandit setting, where a tuning parameter determines the weight placed on cumulative regret minimization (where we establish near-optimal minimax guarantees) versus simple regret minimization (where we establish state-of-the-art guarantees). Our algorithms work with any function class, are robust to model misspecification, and can be used in continuous arm settings. This flexibility comes from constructing and relying on “conformal arm sets" (CASs). CASs provide a set of arms for every context, encompassing the context-specific optimal arm with a certain probability across the context distribution. Our positive results on simple and cumulative regret guarantees are contrasted with a negative result, which shows that no algorithm can achieve instance-dependent simple regret guarantees while simultaneously achieving minimax optimal cumulative regret guarantees. Sanath Kumar Krishna Murthy, Ruohan Zhan, Susan Athey, Emma Brunskill |
NeurIPS | 2 |
| 2023 | Two-Stage Constrained Actor-Critic for Short Video RecommendationabstractThe wide popularity of short videos on social media poses new opportunities and challenges to optimize recommender systems on the video-sharing platforms. Users sequentially interact with the system and provide complex and multi-faceted responses, including WatchTime and various types of interactions with multiple videos. On the one hand, the platforms aim at optimizing the users’ cumulative WatchTime (main goal) in the long term, which can be effectively optimized by Reinforcement Learning. On the other hand, the platforms also need to satisfy the constraint of accommodating the responses of multiple user interactions (auxiliary goals) such as Like, Follow, Share, etc. In this paper, we formulate the problem of short video recommendation as a Constrained Markov Decision Process (CMDP). We find that traditional constrained reinforcement learning algorithms fail to work well in this setting. We propose a novel two-stage constrained actor-critic method: At stage one, we learn individual policies to optimize each auxiliary signal. In stage two, we learn a policy to (i) optimize the main signal and (ii) stay close to policies learned in the first stage, which effectively guarantees the performance of this main policy on the auxiliaries. Through extensive offline evaluations, we demonstrate the effectiveness of our method over alternatives in both optimizing the main goal as well as balancing the others. We further show the advantage of our method in live experiments of short video recommendations, where it significantly outperforms other baselines in terms of both WatchTime and interactions. Our approach has been fully launched in the production system to optimize user experiences on the platform. Qingpeng Cai 0001, Zhenghai Xue, Wanqi Xue, Shuchang Liu 0001, Ruohan Zhan, Tianyou Zuo, Wentao Xie 0002, Peng Jiang 0002, Kun Gai |
WWW | 6 |
| 2022 | Deconfounding Duration Bias in Watch-time Prediction for Video RecommendationabstractWatch-time prediction remains to be a key factor in reinforcing user engagement via video recommendations. It has become increasingly important given the ever-growing popularity of online videos. However, prediction of watch time not only depends on the match between the user and the video but is often mislead by the duration of the video itself. With the goal of improving watch time, recommendation is always biased towards videos with long duration. Models trained on this imbalanced data face the risk of bias amplification, which misguides platforms to over-recommend videos with long duration but overlook the underlying user interests. This paper presents the first work to study duration bias in watch-time prediction for video recommendation. We employ a causal graph illuminating that duration is a confounding factor that concurrently affects video exposure and watch-time prediction---the first effect on video causes the bias issue and should be eliminated, while the second effect on watch time originates from video intrinsic characteristics and should be preserved. To remove the undesired bias but leverage the natural effect, we propose a Duration-Deconfounded Quantile-based (D2Q) watch-time prediction framework, which allows for scalability to perform on industry production systems. Through extensive offline evaluation and live experiments, we showcase the effectiveness of this duration-deconfounding framework by significantly outperforming the state-of-the-art baselines. We have fully launched our approach on Kuaishou App, which has substantially improved real-time video consumption due to more accurate watch-time predictions. Ruohan Zhan, Changhua Pei, Jianfeng Wen, Guanyu Mu, Peng Jiang 0002, Kun Gai |
KDD | 1 |
| 2021 | Off-Policy Evaluation via Adaptive Weighting with Data from Contextual BanditsabstractIt has become increasingly common for data to be collected adaptively, for example using contextual bandits. Historical data of this type can be used to evaluate other treatment assignment policies to guide future innovation or experiments. However, policy evaluation is challenging if the target policy differs from the one used to collect data, and popular estimators, including doubly robust (DR) estimators, can be plagued by bias, excessive variance, or both. In particular, when the pattern of treatment assignment in the collected data looks little like the pattern generated by the policy to be evaluated, the importance weights used in DR estimators explode, leading to excessive variance. Ruohan Zhan, Vitor Hadad, David A. Hirshberg, Susan Athey |
KDD | 1 |
| 2021 | Towards Content Provider Aware Recommender Systems: A Simulation Study on the Interplay between User and Provider UtilitiesabstractMost existing recommender systems focus primarily on matching users (content consumers) to content which maximizes user satisfaction on the platform. It is increasingly obvious, however, that content providers have a critical influence on user satisfaction through content creation, largely determining the content pool available for recommendation. A natural question thus arises: can we design recommenders taking into account the long-term utility of both users and content providers? By doing so, we hope to sustain more content providers and a more diverse content pool for long-term user satisfaction. Understanding the full impact of recommendations on both user and content provider groups is challenging. This paper aims to serve as a research investigation of one approach toward building a content provider aware recommender, and evaluating its impact in a simulated setup. Ruohan Zhan, Konstantina Christakopoulou, Ya Le, Jayden Ooi, Martin Mladenov, Alex Beutel, Craig Boutilier, Ed H. Chi, Minmin Chen |
WWW | 1 |
| 2020 | Distortion Agnostic Deep WatermarkingabstractWatermarking is the process of embedding information into an image that can survive under distortions, while requiring the encoded image to have little or no perceptual difference with the original image. Recently, deep learning-based methods achieved impressive results in both visual quality and message payload under a wide variety of image distortions. However, these methods all require differentiable models for the image distortions at training time, and may generalize poorly to unknown distortions. This is undesirable since the types of distortions applied to watermarked images are usually unknown and non-differentiable. In this paper, we propose a new framework for distortion-agnostic watermarking, where the image distortion is not explicitly modeled during training. Instead, the robustness of our system comes from two sources: adversarial training and channel coding. Compared to training on a fixed set of distortions and noise levels, our method achieves comparable or better results on distortions available during training, and better performance overall on unknown distortions. Xiyang Luo, Ruohan Zhan, Huiwen Chang, Feng Yang 0008, Peyman Milanfar |
CVPR | 2 |
| 2016 | CT Image Reconstruction by Spatial-Radon Domain Data-Driven Tight Frame RegularizationabstractThis paper proposes a spatial-Radon domain computed tomography (CT) image reconstruction model based on data-driven tight frames (SRD-DDTF). The proposed SRD-DDTF model combines the idea of the joint image and Radon domain inpainting model of Dong, Li, and Shen [J. Sci. Comput., 54 (2013), pp. 333--349] and that of the data-driven tight frames for image denoising [J.-F. Cai, H. Ji, Z. Shen, and G.-B. Ye, Appl. Comput. Harmon. Anal., 37 (2014), p. 89--105]. It is different from existing models in that both the CT image and its corresponding high quality projection image are reconstructed simultaneously using sparsity priors by tight frames that are adaptively learned from the data to provide optimal sparse approximations. An alternative minimization algorithm is designed to solve the proposed model, which is nonsmooth and nonconvex. Convergence analysis of the algorithm is provided. Numerical experiments show that the SRD-DDTF model is superior to the model of Dong, Li, and Shen [J. Sci. Comput., 54 (2013), pp. 333--349] especially in recovering some subtle structures in the images. Ruohan Zhan |
SIAM J. Imaging Sci. | 1 |