EDBT 2026 Demo / reviewers in the wild / expert
Shentao Yang
dblp:314/8076
· DBLP profile ↗
6ranked-venue papers
5as first author
6since 2021 · last 2025
0009-0009-8058-3149ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 57% Language models and text generation · 20% Generative modeling · 15% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › alignment
preference alignment |
1.4 | 2 | 2024 | A Dense Reward View on Aligning Text-to-Image Diffusion with Preference · ICML 2024 Preference-grounded Token-level Guidance for Language Model Fine-tuning · NeurIPS 2023 |
Machine learning › Reinforcement learning
offline reinforcement learning |
1.1 | 2 | 2022 | A Unified Framework for Alternating Offline Model Training and Policy Learning · NeurIPS 2022 Regularizing a Model-based Policy Stationary Distribution to Stabilize Offline Reinforcement Learning · ICML 2022 |
Recommender systems
video recommendation |
0.9 | 1 | 2025 | SWaT: Statistical Modeling of Video Watch Time through User Behavior Analysis · KDD (1) 2025 |
Recommender systems › video recommendation
watch-time prediction |
0.9 | 1 | 2025 | SWaT: Statistical Modeling of Video Watch Time through User Behavior Analysis · KDD (1) 2025 |
Machine learning › Reinforcement learning › reward design › reward shaping
dense reward shaping |
0.8 | 1 | 2024 | A Dense Reward View on Aligning Text-to-Image Diffusion with Preference · ICML 2024 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | A Dense Reward View on Aligning Text-to-Image Diffusion with Preference · ICML 2024 |
Machine learning › Reinforcement learning
reward design |
0.8 | 1 | 2024 | A Dense Reward View on Aligning Text-to-Image Diffusion with Preference · ICML 2024 |
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model |
0.8 | 1 | 2024 | A Dense Reward View on Aligning Text-to-Image Diffusion with Preference · ICML 2024 |
Natural language and speech › Language models and text generation
large language model fine-tuning |
0.7 | 1 | 2023 | Preference-grounded Token-level Guidance for Language Model Fine-tuning · NeurIPS 2023 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.7 | 1 | 2023 | Preference-grounded Token-level Guidance for Language Model Fine-tuning · NeurIPS 2023 |
Machine learning › Reinforcement learning
reward learning |
0.7 | 1 | 2023 | Fantastic Rewards and How to Tame Them: A Case Study on Reward Learning for Task-oriented Dialogue Systems · ICLR 2023 |
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue |
0.7 | 1 | 2023 | Fantastic Rewards and How to Tame Them: A Case Study on Reward Learning for Task-oriented Dialogue Systems · ICLR 2023 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.6 | 1 | 2022 | A Unified Framework for Alternating Offline Model Training and Policy Learning · NeurIPS 2022 |
Machine learning › Reinforcement learning
policy learning |
0.6 | 1 | 2022 | A Unified Framework for Alternating Offline Model Training and Policy Learning · NeurIPS 2022 |
Machine learning › Reinforcement learning
policy optimization |
0.6 | 1 | 2022 | Regularizing a Model-based Policy Stationary Distribution to Stabilize Offline Reinforcement Learning · ICML 2022 |
Visualization and visual analytics
user behavior analysis |
0.3 | 1 | 2025 | SWaT: Statistical Modeling of Video Watch Time through User Behavior Analysis · KDD (1) 2025 |
Machine learning › Reinforcement learning › reward design
reward shaping |
0.2 | 1 | 2023 | Fantastic Rewards and How to Tame Them: A Case Study on Reward Learning for Task-oriented Dialogue Systems · ICLR 2023 |
Methods — techniques the papers use, named apart from their topics
statistical modeling · 1.7regression · 1.7classification · 1.7bucketization · 1.7temporal discounting · 0.8direct preference optimization · 0.8reward learning · 0.7pairwise preference learning · 0.7imitation learning · 0.7case study · 0.7lower bound maximization · 0.6dynamics model · 0.6dynamic model training · 0.6distribution matching · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SWaT: Statistical Modeling of Video Watch Time through User Behavior AnalysisabstractThe significance of estimating video watch time has been highlighted by the rising importance of (short) video recommendation, which has become a core product of mainstream social media platforms. Modeling video watch time, however, has been challenged by the complexity of user-video interaction, such as different user behavior modes in watching the recommended videos and varying watching probability over the video progress bar. Despite the importance and challenges, existing literature on modeling video watch time mostly focuses on relatively black-box mechanical enhancement of the classical regression/classification losses, without factoring in user behavior in a principled manner. In this paper, we for the first time take on a user-centric perspective to model video watch time, from which we propose a white-box statistical framework that directly translates various user behavior assumptions in watching (short) videos into statistical watch time models. These behavior assumptions are portrayed by our domain knowledge on users' behavior modes in video watching. We further employ bucketization to cope with user's non-stationary watching probability over the video progress bar, which additionally helps to respect the constraint of video length and facilitate the practical compatibility between the continuous regression event of watch time and other binary classification events. We test our models extensively on two public datasets, a large-scale offline industrial dataset, and an online A/B test on a short video platform with hundreds of millions of daily-active users. On all experiments, our models perform competitively against strong relevant baselines, demonstrating the efficacy of our user-centric perspective and proposed framework. Shentao Yang, Haichuan Yang, Linna Du, Adithya Ganesh, Bo Peng 0009, Boying Liu, Serena Li, Ji Liu 0002 |
KDD (1) | 1 |
| 2024 | A Dense Reward View on Aligning Text-to-Image Diffusion with PreferenceabstractAligning text-to-image diffusion model (T2I) with preference has been gaining increasing research attention. While prior works exist on directly optimizing T2I by preference data, these methods are developed under the bandit assumption of a latent reward on the entire diffusion reverse chain, while ignoring the sequential nature of the generation process. This may harm the efficacy and efficiency of preference alignment. In this paper, we take on a finer dense reward perspective and derive a tractable alignment objective that emphasizes the initial steps of the T2I reverse chain. In particular, we introduce temporal discounting into DPO-style explicit-reward-free objectives, to break the temporal symmetry therein and suit the T2I generation hierarchy. In experiments on single and multiple prompt generation, our method is competitive with strong relevant baselines, both quantitatively and qualitatively. Further investigations are conducted to illustrate the insight of our approach. Source code is available at https://github.com/Shentao-YANG/Dense_Reward_T2I . Shentao Yang, Mingyuan Zhou |
ICML | 1 |
| 2023 | Fantastic Rewards and How to Tame Them: A Case Study on Reward Learning for Task-oriented Dialogue Systems
Yihao Feng, Shentao Yang, Shujian Zhang, Jianguo Zhang 0005, Caiming Xiong, Mingyuan Zhou, Huan Wang 0016 |
ICLR | 2 |
| 2023 | Preference-grounded Token-level Guidance for Language Model Fine-tuningabstractAligning language models (LMs) with preferences is an important problem in natural language generation. A key challenge is that preferences are typically provided at the *sequence level* while LM training and generation both occur at the *token level*. There is, therefore, a *granularity mismatch* between the preference and the LM training losses, which may complicate the learning problem. In this paper, we address this issue by developing an alternate training process, where we iterate between grounding the sequence-level preference into token-level training guidance, and improving the LM with the learned guidance. For guidance learning, we design a framework that extends the pairwise-preference learning in imitation learning to both variable-length LM generation and the utilization of the preference among multiple generations. For LM training, based on the amount of supervised data, we present two *minimalist* learning objectives that utilize the learned guidance. In experiments, our method performs competitively on two distinct representative LM tasks --- discrete-prompt generation and text summarization. Shentao Yang, Shujian Zhang, Congying Xia, Yihao Feng, Caiming Xiong, Mingyuan Zhou |
NeurIPS | 1 |
| 2022 | Regularizing a Model-based Policy Stationary Distribution to Stabilize Offline Reinforcement LearningabstractOffline reinforcement learning (RL) extends the paradigm of classical RL algorithms to purely learning from static datasets, without interacting with the underlying environment during the learning process. A key challenge of offline RL is the instability of policy training, caused by the mismatch between the distribution of the offline data and the undiscounted stationary state-action distribution of the learned policy. To avoid the detrimental impact of distribution mismatch, we regularize the undiscounted stationary distribution of the current policy towards the offline data during the policy optimization process. Further, we train a dynamics model to both implement this regularization and better estimate the stationary distribution of the current policy, reducing the error induced by distribution mismatch. On a wide range of continuous-control offline RL datasets, our method indicates competitive performance, which validates our algorithm. The code is publicly available. Shentao Yang, Yihao Feng, Shujian Zhang, Mingyuan Zhou |
ICML | 1 |
| 2022 | A Unified Framework for Alternating Offline Model Training and Policy LearningabstractIn offline model-based reinforcement learning (offline MBRL), we learn a dynamic model from historically collected data, and subsequently utilize the learned model and fixed datasets for policy learning, without further interacting with the environment. Offline MBRL algorithms can improve the efficiency and stability of policy learning over the model-free algorithms. However, in most of the existing offline MBRL algorithms, the learning objectives for the dynamic models and the policies are isolated from each other. Such an objective mismatch may lead to inferior performance of the learned agents. In this paper, we address this issue by developing an iterative offline MBRL framework, where we maximize a lower bound of the true expected return, by alternating between dynamic-model training and policy learning. With the proposed unified model-policy learning framework, we achieve competitive performance on a wide range of continuous-control offline reinforcement learning datasets. Source code is released at https://github.com/Shentao-YANG/AMPL_NeurIPS2022. Shentao Yang, Shujian Zhang, Yihao Feng, Mingyuan Zhou |
NeurIPS | 1 |