VLDB 2026 Research / reviewers in the wild / expert
Ruihao Zhu
dblp:131/9596
· DBLP profile ↗
18ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 7 since 2021Computer networks · 5 · 4 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evolutionary Contrastive Ensemble With Conditional Redundancy Fitness Evaluation for Goal-Conditioned Humanoid LocomotionabstractGoal-conditioned humanoid locomotion in reinforcement learning (RL) remains challenging due to sparse reward signals and the single-goal overfitting problem. Although contrastive reinforcement learning (CRL) has achieved considerable success in this setting, it can suffer from pronounced estimation variance, since epistemic uncertainty is difficult to reduce given limited task-specific information and model capacity. Ensemble-based critics can partially alleviate this issue. However, sufficient ensemble diversity and accurate individual estimates are not necessarily guaranteed during training, resulting in unstructured exploration. To address these challenges, we propose Conditional Redundancy-Guided Evolutionary Contrastive Ensemble with Direct Preference Optimization weighting (CRECE-DPO), which augments CRL with a vectorized critic ensemble and refines the ensemble via an evolutionary algorithm guided by a tailored fitness metric. Specifically, we design a DPO-weighted conditional redundancy fitness score, to prune redundant representations while promoting effective exploration of the parameter space. Simulation results on challenging goal-conditioned benchmarks, including humanoid locomotion, demonstrate consistent improvements over CRL and other baselines. Zhiyi Shi, Haoyu Pan, Ruihao Zhu, Changyu Li, Shuai Wu 0004, Qi Wu 0003 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Satisficing Regret Minimization in BanditsabstractMotivated by the concept of satisficing in decision-making, we consider the problem of satisficing exploration in bandit optimization. In this setting, the learner aims at finding a satisficing arm whose mean reward exceeds a certain threshold. The performance is measured by satisficing regret, which is the cumulative deficit of the chosen arm's mean reward compared to the threshold. We propose $\texttt{SELECT}$, a general algorithmic template for Satisficing REgret Minimization via SampLing and LowEr Confidence bound Testing, that attains constant satisficing regret for a wide variety of bandit optimization problems in the realizable case (i.e., whenever a satisficing arm exists). Specifically, given a class of bandit optimization problems and a corresponding learning oracle with sub-linear (standard) regret upper bound, $\texttt{SELECT}$ iteratively makes use of the oracle to identify a potential satisficing arm. Then, it collects data samples from this arm, and continuously compares the lower confidence bound of the identified arm's mean reward against the threshold value to determine if it is a satisficing arm. As a complement, $\texttt{SELECT}$ also enjoys the same (standard) regret guarantee as the oracle in the non-realizable case. Finally, we conduct numerical experiments to validate the performance of $\texttt{SELECT}$ for several popular bandit optimization settings. Ruihao Zhu |
ICLR | 3 |
| 2025 | Optimizing the Battery-Swapping Problem in Urban E-Bike Systems with Reinforcement LearningabstractE-bikes (EBs) are a key transportation mode in urban area, especially for couriers of delivery platforms, but underdeveloped EB systems can hinder courier's productivity due to limited battery capacity. Battery-swapping stations address this issue by enabling riders to exchange depleted batteries for fully charged ones. However, managing supply and demand (SnD) imbalances at these stations has become increasingly complex. To address this, we introduce a new approach that formulates the Battery-Swapping Problem (BSP) as a discrete-time Markov Decision Process (MDP) to capture the dynamics of SnD imbalances. Building on it, we propose a Wasserstein-enhanced Proximal Policy Optimization (W-PPO) algorithm, which integrates Wasserstein distance with reinforcement learning to improve the robustness against uncertainty in forecasting SnD. W-PPO provides a BSP-specific, accurate loss function that reflects reward variations between two policies under real-world simulation. The algorithm’s effectiveness is assessed using key metrics: Shared Battery Utilization Ratio (SBUR) and Battery Supply Ratio (BSR). Simulations on real-world datasets show that W-PPO achieves a 30.59% improvement in SBUR and a 16.09% increase in BSR ensures practical applicability. By optimizing battery utilization and improving EB delivery systems, this work highlights the potential of AI for creating efficient and sustainable urban transportation solutions. Zhao Li 0007, Xuanwu Liu, Ruihao Zhu, Zhenzhe Zheng 0001, Fan Wu 0006 |
IJCAI | 4 |
| 2025 | Contextual Online Pricing with (Biased) Offline DataabstractWe study contextual online pricing with biased offline data. For the scalar price elasticity case, we identify the instance-dependent quantity $\delta^2$ that measures how far the offline data lies from the (unknown) online optimum. We show that the time length $T$, bias bound $V$, size $N$ and dispersion $\lambda_{\min}(\hat{\Sigma})$ of the offline data, and $\delta^2$ jointly determine the statistical complexity. An Optimism‑in‑the‑Face‑of‑Uncertainty (OFU) policy achieves a minimax-optimal, instance-dependent regret bound $\tilde{\mathcal{O}}\big(d\sqrt{T} \wedge (V^2T + \frac{dT }{\lambda_{\min}(\hat{\Sigma}) + (N \wedge T) \delta^2})\big)$. For general price elasticity, we establish a worst‑case, minimax-optimal rate $\tilde{\mathcal{O}}\big(d\sqrt{T} \wedge (V^2T + \frac{dT }{\lambda_{\min}(\hat{\Sigma})})\big)$ and provide a generalized OFU algorithm that attains it. When the bias bound $V$ is unknown, we design a robust variant that always guarantees sub‑linear regret and strictly improves on purely online methods whenever the exact bias is small. These results deliver the first tight regret guarantees for contextual pricing in the presence of biased offline data. Our techniques also transfer verbatim to stochastic linear bandits with biased offline data, yielding analogous bounds. Ruihao Zhu, Qiaomin Xie |
NeurIPS | 2 |
| 2024 | User Experience Design Professionals' Perceptions of Generative Artificial IntelligenceabstractAmong creative professionals, Generative Artificial Intelligence (GenAI) has sparked excitement over its capabilities and fear over unanticipated consequences. How does GenAI impact User Experience Design (UXD) practice, and are fears warranted? We interviewed 20 UX Designers, with diverse experience and across companies (startups to large enterprises). We probed them to characterize their practices, and sample their attitudes, concerns, and expectations. We found that experienced designers are confident in their originality, creativity, and empathic skills, and find GenAI’s role as assistive. They emphasized the unique human factors of “enjoyment” and “agency”, where humans remain the arbiters of “AI alignment’’. However, skill degradation, job replacement, and creativity exhaustion can adversely impact junior designers. We discuss implications for human-GenAI collaboration, specifically copyright and ownership, human creativity and agency, and AI literacy and access. Through the lens of responsible and participatory AI, we contribute a deeper understanding of GenAI fears and opportunities for UXD. Jie Li 0064, Hancheng Cao, Laura Lin, Youyang Hou, Ruihao Zhu, Abdallah El Ali |
CHI | 5 |
| 2023 | MERIT: A Merchant Incentive Ranking Model for Hotel Search & RankingabstractOnline Travel Platforms (OTPs) have been working on improving their hotel Search & Ranking (S&R) systems that facilitate efficient matching between consumers and hotels. Existing OTPs focus on improving platform revenue. In this work, we take a first step in incorporating hotel merchants' objectives into the design of hotel S&R systems to achieve an incentive loop: the OTP tilts impressions and better-ranked positions to merchants with high service quality, and in return, the merchants provide better service to consumers. Three critical design challenges need to be resolved to achieve this incentive loop: Matthew Effect in the consumer feedback-loop, unclear relation between hotel service quality and performance, and conflicts between platform revenue and consumer experience. Shigang Quan, Zhenzhe Zheng 0001, Ruihao Zhu, Liangyue Li, Fan Wu 0006 |
CIKM | 5 |
| 2023 | Temporal Fairness in Learning and Earning: Price Protection Guarantee and Phase TransitionsabstractMotivated by the prevalence of "price protection guarantee", which helps to promote temporal fairness in dynamic pricing, we study the impact of such policy on the design of online learning algorithm for data-driven dynamic pricing with initially unknown customer demand. Under the price protection guarantee, a customer who purchased a product in the past can receive a refund from the seller during the so-called price protection period (typically defined as a certain time window after the purchase date) in case the seller decides to lower the price. We consider a setting where a firm sells a product over a horizon of T time steps. For this setting, we characterize how the value of M, the length of price protection period, can affect the optimal regret of the learning process. Our contributions can be summarized as follows: Ruihao Zhu, Stefanus Jasin |
EC | 2 |
| 2023 | LINet: A Location and Intention-Aware Neural Network for Hotel Group RecommendationabstractMotivated by the collaboration with Fliggy1, a leading Online Travel Platform (OTP), we investigate an important but less explored research topic about optimizing the quality of hotel supply, namely selecting potential profitable hotels in advance to build up adequate room inventory. We formulate a WWW problem, i.e., within a specific time period (When) and potential travel area (Where), which hotels should be recommended to a certain group of users with similar travel intentions (Why). We identify three critical challenges in solving the WWW problem: user groups generation, travel data sparsity and utilization of hotel recommendation information (e.g., period, location and intention). To this end, we propose LINet, a Location and Intention-aware neural Network for hotel group recommendation. Specifically, LINet first identifies user travel intentions for user groups generalization, and then characterizes the group preferences by jointly considering historical user-hotel interaction and spatio-temporal features of hotels. For data sparsity, we develop a graph neural network, which employs long-term data, and further design an auxiliary loss function of location that efficiently exploits data within the same and across different locations. Both offline and online experiments demonstrate the effectiveness of LINet when compared with state-of-the-art methods. LINet has been successfully deployed on Fliggy to retrieve high quality hotels for business development, serving hundreds of hotel operation scenarios and thousands of hotel operators. Ruitao Zhu, Detao Lv, Ruihao Zhu, Zhenzhe Zheng 0001, Ke Bu, Fan Wu 0006 |
WWW | 4 |
| 2022 | Safe Optimal Design with Applications in Off-Policy LearningabstractMotivated by practical needs in online experimentation and off-policy learning, we study the problem of safe optimal design, where we develop a data logging policy that efficiently explores while achieving competitive rewards with a baseline production policy. We first show, perhaps surprisingly, that a common practice of mixing the production policy with uniform exploration, despite being safe, is sub-optimal in maximizing information gain. Then we propose a safe optimal logging policy for the case when no side information about the actions’ expected rewards is available. We improve upon this design by considering side information and also extend both approaches to a large number of actions with a linear reward model. We analyze how our data logging policies impact errors in off-policy learning. Finally, we empirically validate the benefit of our designs by conducting extensive experiments. Ruihao Zhu, Branislav Kveton |
AISTATS | 1 |
| 2021 | Near-Optimal Model-Free Reinforcement Learning in Non-Stationary Episodic MDPsabstractWe consider model-free reinforcement learning (RL) in non-stationary Markov decision processes. Both the reward functions and the state transition functions are allowed to vary arbitrarily over time as long as their cumulative variations do not exceed certain variation budgets. We propose Restarted Q-Learning with Upper Confidence Bounds (RestartQ-UCB), the first model-free algorithm for non-stationary RL, and show that it outperforms existing solutions in terms of dynamic regret. Specifically, RestartQ-UCB with Freedman-type bonus terms achieves a dynamic regret bound of $\widetilde{O}(S^{\frac{1}{3}} A^{\frac{1}{3}} \Delta^{\frac{1}{3}} H T^{\frac{2}{3}})$, where $S$ and $A$ are the numbers of states and actions, respectively, $\Delta>0$ is the variation budget, $H$ is the number of time steps per episode, and $T$ is the total number of time steps. We further show that our algorithm is \emph{nearly optimal} by establishing an information-theoretical lower bound of $\Omega(S^{\frac{1}{3}} A^{\frac{1}{3}} \Delta^{\frac{1}{3}} H^{\frac{2}{3}} T^{\frac{2}{3}})$, the first lower bound in non-stationary RL. Numerical experiments validate the advantages of RestartQ-UCB in terms of both cumulative rewards and computational efficiency. We further demonstrate the power of our results in the context of multi-agent RL, where non-stationarity is a key challenge. Weichao Mao, Kaiqing Zhang, Ruihao Zhu, David Simchi-Levi, Tamer Basar |
ICML | 3 |
| 2020 | Reinforcement Learning for Non-Stationary Markov Decision Processes: The Blessing of (More) OptimismabstractWe consider un-discounted reinforcement learning (RL) in Markov decision processes (MDPs) under drifting non-stationarity, \ie, both the reward and state transition distributions are allowed to evolve over time, as long as their respective total variations, quantified by suitable metrics, do not exceed certain \emph{variation budgets}. We first develop the Sliding Window Upper-Confidence bound for Reinforcement Learning with Confidence Widening (\texttt{SWUCRL2-CW}) algorithm, and establish its dynamic regret bound when the variation budgets are known. In addition, we propose the Bandit-over-Reinforcement Learning (\texttt{BORL}) algorithm to adaptively tune the \sw to achieve the same dynamic regret bound, but in a \emph{parameter-free} manner, \ie, without knowing the variation budgets. Notably, learning drifting MDPs via conventional optimistic exploration presents a unique challenge absent in existing (non-stationary) bandit learning settings. We overcome the challenge by a novel confidence widening technique that incorporates additional optimism. Wang Chi Cheung, David Simchi-Levi, Ruihao Zhu |
ICML | 3 |
| 2019 | Learning to Optimize under Non-StationarityabstractWe introduce algorithms that achieve state-of-the-art dynamic regret bounds for non-stationary linear stochastic bandit setting. It captures natural applications such as dynamic pricing and ads allocation in a changing environment. We show how the difficulty posed by the non-stationarity can be overcome by a novel marriage between stochastic and adversarial bandits learning algorithms. Our main contributions are the tuned Sliding Window UCB (SW-UCB) algorithm with optimal dynamic regret, and the tuning free bandit-over-bandit (BOB) framework built on top of the SW-UCB algorithm with best (compared to existing literature) dynamic regret. Wang Chi Cheung, David Simchi-Levi, Ruihao Zhu |
AISTATS | 3 |
| 2017 | Coresets for differentially private k-means clustering and applications to privacy in mobile sensor networksabstractMobile sensor networks are a great source of data. By collecting data with mobile sensor nodes from individuals in a user community, e.g. using their smartphones, we can learn global information such as traffic congestion patterns in the city, location of key community facilities, and locations of gathering places. Can we publish and run queries on mobile sensor network databases without disclosing information about individual nodes? Dan Feldman, Chongyuan Xiang, Ruihao Zhu, Daniela Rus |
IPSN | 3 |
| 2016 | Threshold Bandits, With and Without Censored FeedbackabstractWe consider the \emph{Threshold Bandit} setting, a variant of the classical multi-armed bandit problem in which the reward on each round depends on a piece of side information known as a \emph{threshold value}. The learner selects one of $K$ actions (arms), this action generates a random sample from a fixed distribution, and the action then receives a unit payoff in the event that this sample exceeds the threshold value. We consider two versions of this problem, the \emph{uncensored} and \emph{censored} case, that determine whether the sample is always observed or only when the threshold is not met. Using new tools to understand the popular UCB algorithm, we show that the uncensored case is essentially no more difficult than the classical multi-armed bandit setting. Finally we show that the censored case exhibits more challenges, but we give guarantees in the event that the sequence of threshold values is generated optimistically. Jacob D. Abernethy, Kareem Amin 0002, Ruihao Zhu |
NIPS | 3 |
| 2015 | Differentially private and strategy-proof spectrum auction with approximate revenue maximizationabstractThe rapid growth of wireless mobile users and applications has led to high demand of spectrum. Auction is a powerful tool to improve the utilization of spectrum resource, and many auction mechanisms have been proposed thus far. However, none of them has considered both the privacy of bidders and the revenue gain of the auctioneer together. In this paper, we study the design of privacy-preserving auction mechanisms. We first propose a differentially private auction mechanism which can achieve strategy-proofness and a near optimal expected revenue based on the concept of virtual valuation. Assuming the knowledge of the bidders' valuation distributions, the near optimal differentially private and strategy-proof auction mechanism uses the generalized Vickrey-Clarke-Groves auction payment scheme to achieve high revenue with a high probability. To tackle its high computational complexity, we also propose an approximate differentially PrivAte, Strategy-proof, and polynomially tractable Spectrum (PASS) auction mechanism that can achieve a suboptimal revenue. PASS uses a monotone allocation algorithm and the critical payment scheme to achieve strategy-proofness. We also evaluate PASS extensively via simulation, showing that it can generate more revenue than existing mechanisms in the spectrum auction markets. Ruihao Zhu, Kang G. Shin |
INFOCOM | 1 |
| 2014 | Differentially private spectrum auction with approximate revenue maximizationabstractDynamic spectrum redistribution---under which spectrum owners lease out under-utilized spectrum to users for financial gain---is an effective way to improve spectrum utilization. Auction is a natural way to incentivize spectrum owners to share their idle resources. In recent years, a number of strategy-proof auction mechanisms have been proposed to stimulate bidders to truthfully reveal their valuations. However, it has been shown that truthfulness is not a necessary condition for revenue maximization. Furthermore, in most existing spectrum auction mechanisms, bidders may infer the valuations---which are private information---of the other bidders from the auction outcome. In this paper, we propose a Differentially privatE spectrum auction mechanism with Approximate Revenue maximization (DEAR). We theoretically prove that DEAR achieves approximate truthfulness, privacy preservation, and approximate revenue maximization. Our extensive evaluations show that DEAR achieves good performance in terms of both revenue and privacy preservation. Ruihao Zhu, Zhijing Li 0001, Fan Wu 0006, Kang G. Shin, Guihai Chen |
MobiHoc | 1 |
| 2013 | STAMP: A Strategy-proof Approximation auction Mechanism for Spatially reusable Items in wireless networksabstractThe advent of participatory sensing markets and spectrum markets based on the wireless networks have led to a new kind of auction dealing with spatially reusable items, which can be shared by multiple parties that are geographically far apart enough from each other. Simply applying traditional auctions to spatially reusable items is vulnerable to bid manipulation, and may lead to low allocation efficiency. In this paper, we study the problem of auctioning spatially reusable items. We propose STAMP, which is a STrategy-proof Approximation auction Mechanism for sPatially reusable items in wireless networks. STAMP can be implemented with any existing maximum independent set algorithm, and can guarantee the allocation efficiency as high as the algorithm based on. Evaluation results show that STAMP achieves much better performance than existing mechanisms, in terms of allocation efficiency. Ruihao Zhu, Fan Wu 0006, Guihai Chen |
GLOBECOM | 1 |
| 2013 | SAFE: A Strategy-Proof Auction Mechanism for Multi-radio, Multi-channel Spectrum Allocation
Ruihao Zhu, Fan Wu 0006, Guihai Chen |
WASA | 1 |