EDBT 2026 Demo / reviewers in the wild / expert
Tian-Shuo Liu
dblp:374/3346
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Reinforcement learning · 56% Generative modeling · 28% Motion planning and robot control · 7% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
offline reinforcement learning |
2.4 | 3 | 2025 | Semantic Temporal Abstraction via Vision-Language Model Guidance for Efficient Reinforcement Learning · ICLR 2025 Offline Transition Modeling via Contrastive Energy Learning · ICML 2024 Flow to Better: Offline Preference-based Reinforcement Learning via Preferred Trajectory Generation · ICLR 2024 |
Machine learning › Generative modeling
diffusion model |
1.5 | 2 | 2024 | Energy-Guided Diffusion Sampling for Offline-to-Online Reinforcement Learning · ICML 2024 Flow to Better: Offline Preference-based Reinforcement Learning via Preferred Trajectory Generation · ICLR 2024 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning › skill learning
skill discovery |
0.9 | 1 | 2025 | Semantic Temporal Abstraction via Vision-Language Model Guidance for Efficient Reinforcement Learning · ICLR 2025 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
temporal abstraction |
0.9 | 1 | 2025 | Semantic Temporal Abstraction via Vision-Language Model Guidance for Efficient Reinforcement Learning · ICLR 2025 |
Machine learning › Generative modeling › diffusion model
diffusion sampling |
0.8 | 1 | 2024 | Energy-Guided Diffusion Sampling for Offline-to-Online Reinforcement Learning · ICML 2024 |
Machine learning › Trustworthy machine learning › robustness
distribution shift |
0.8 | 1 | 2024 | Energy-Guided Diffusion Sampling for Offline-to-Online Reinforcement Learning · ICML 2024 |
Machine learning › Generative modeling
energy-based model |
0.8 | 1 | 2024 | Offline Transition Modeling via Contrastive Energy Learning · ICML 2024 |
Machine learning › Reinforcement learning › offline reinforcement learning
offline-to-online reinforcement learning |
0.8 | 1 | 2024 | Energy-Guided Diffusion Sampling for Offline-to-Online Reinforcement Learning · ICML 2024 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
preference-based reinforcement learning |
0.8 | 1 | 2024 | Flow to Better: Offline Preference-based Reinforcement Learning via Preferred Trajectory Generation · ICLR 2024 |
Robotics › Motion planning and robot control
trajectory planning |
0.8 | 1 | 2024 | Flow to Better: Offline Preference-based Reinforcement Learning via Preferred Trajectory Generation · ICLR 2024 |
Computer vision › Vision and language › vision-language model
vision-language model guidance |
0.3 | 1 | 2025 | Semantic Temporal Abstraction via Vision-Language Model Guidance for Efficient Reinforcement Learning · ICLR 2025 |
Machine learning › Reinforcement learning
imitation learning |
0.2 | 1 | 2024 | Flow to Better: Offline Preference-based Reinforcement Learning via Preferred Trajectory Generation · ICLR 2024 |
Machine learning › Reinforcement learning
off-policy evaluation |
0.2 | 1 | 2024 | Offline Transition Modeling via Contrastive Energy Learning · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
diffusion model · 1.5vision-language model · 0.9vector quantization · 0.9skill relabeling · 0.9preference augmentation · 0.8energy-based modeling · 0.8energy function · 0.8contrastive energy learning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Semantic Temporal Abstraction via Vision-Language Model Guidance for Efficient Reinforcement LearningabstractExtracting temporally extended skills can significantly improve the efficiency of reinforcement learning (RL) by breaking down complex decision-making problems with sparse rewards into simpler subtasks and enabling more effective credit assignment. However, existing abstraction methods either discover skills in an unsupervised manner, which often lacks semantic information and leads to erroneous or scattered skill extraction results, or require substantial human intervention. In this work, we propose to leverage the extensive knowledge in pretrained Vision-Language Models (VLMs) to progressively guide the latent space after vector quantization to be more semantically meaningful through relabeling each skill. This approach, termed **V**ision-l**an**guage model guided **T**emporal **A**bstraction (**VanTA**), facilitates the discovery of more interpretable and task-relevant temporal segmentations from offline data without the need for extensive manual intervention or heuristics. By leveraging the rich information in VLMs, our method can significantly outperform existing offline RL approaches that depend only on limited training data. From a theory perspective, we demonstrate that stronger internal sequential correlations within each sub-task, induced by VanTA, effectively reduces suboptimality in policy learning. We validate the effectiveness of our approach through extensive experiments on diverse environments, including Franka Kitchen, Minigrid, and Crafter. These experiments show that our method surpasses existing approaches in long-horizon offline reinforcement learning scenarios with both proprioceptive and visual observations. Tian-Shuo Liu, Xu-Hui Liu, Ruifeng Chen 0003, Lixuan Jin, Yang Yu 0001 |
ICLR | 1 |
| 2025 | InCLET: Large Language Model In-context Learning can Improve Embodied Instruction-following
Peng-Yuan Wang, Jing-Cheng Pang, Xu-Hui Liu, Tian-Shuo Liu, Si-Hang Yang, Hong Qian |
AAMAS | 5 |
| 2024 | Flow to Better: Offline Preference-based Reinforcement Learning via Preferred Trajectory GenerationabstractOffline preference-based reinforcement learning (PbRL) offers an effective solution to overcome the challenges associated with designing rewards and the high costs of online interactions. In offline PbRL, agents are provided with a fixed dataset containing human preferences between pairs of trajectories. Previous studies mainly focus on recovering the rewards from the preferences, followed by policy optimization with an off-the-shelf offline RL algorithm. However, given that preference label in PbRL is inherently trajectory-based, accurately learning transition-wise rewards from such label can be challenging, potentially leading to misguidance during subsequent offline RL training. To address this issue, we introduce our method named $\textit{Flow-to-Better (FTB)}$, which leverages the pairwise preference relationship to guide a generative model in producing preferred trajectories, avoiding Temporal Difference (TD) learning with inaccurate rewards. Conditioning on a low-preference trajectory, $\textit{FTB}$ uses a diffusion model to generate a better one with a higher preference, achieving high-fidelity full-horizon trajectory improvement. During diffusion training, we propose a technique called $\textit{Preference Augmentation}$ to alleviate the problem of insufficient preference data. As a result, we surprisingly find that the model-generated trajectories not only exhibit increased preference and consistency with the real transition but also introduce elements of $\textit{novelty}$ and $\textit{diversity}$, from which we can derive a desirable policy through imitation learning. Experimental results on D4RL benchmarks demonstrate that FTB achieves a remarkable improvement compared to state-of-the-art offline PbRL methods. Furthermore, we show that FTB can also serve as an effective data augmentation method for offline RL. Junyin Ye, Tian-Shuo Liu, Yang Yu 0001 |
ICLR | 4 |
| 2024 | Offline Transition Modeling via Contrastive Energy LearningabstractLearning a high-quality transition model is of great importance for sequential decision-making tasks, especially in offline settings. Nevertheless, the complex behaviors of transition dynamics in real-world environments pose challenges for the standard forward models because of their inductive bias towards smooth regressors, conflicting with the inherent nature of transitions such as discontinuity or large curvature. In this work, we propose to model the transition probability implicitly through a scalar-value energy function, which enables not only flexible distribution prediction but also capturing complex transition behaviors. The Energy-based Transition Models (ETM) are shown to accurately fit the discontinuous transition functions and better generalize to out-of-distribution transition data. Furthermore, we demonstrate that energy-based transition models improve the evaluation accuracy and significantly outperform other off-policy evaluation methods in DOPE benchmark. Finally, we show that energy-based transition models also benefit reinforcement learning and outperform prior offline RL algorithms in D4RL Gym-Mujoco tasks. Ruifeng Chen 0003, Chengxing Jia, Zefang Huang, Tian-Shuo Liu, Xu-Hui Liu, Yang Yu 0001 |
ICML | 4 |
| 2024 | Energy-Guided Diffusion Sampling for Offline-to-Online Reinforcement LearningabstractCombining offline and online reinforcement learning (RL) techniques is indeed crucial for achieving efficient and safe learning where data acquisition is expensive. Existing methods replay offline data directly in the online phase, resulting in a significant challenge of data distribution shift and subsequently causing inefficiency in online fine-tuning. To address this issue, we introduce an innovative approach, Energy-guided DIffusion Sampling (EDIS), which utilizes a diffusion model to extract prior knowledge from the offline dataset and employs energy functions to distill this knowledge for enhanced data generation in the online phase. The theoretical analysis demonstrates that EDIS exhibits reduced suboptimality compared to solely utilizing online data or directly reusing offline data. EDIS is a plug-in approach and can be combined with existing methods in offline-to-online RL setting. By implementing EDIS to off-the-shelf methods Cal-QL and IQL, we observe a notable 20% average improvement in empirical performance on MuJoCo, AntMaze, and Adroit environments. Code is available at https://github.com/liuxhym/EDIS. Xu-Hui Liu, Tian-Shuo Liu, Shengyi Jiang, Ruifeng Chen 0003, Yang Yu 0001 |
ICML | 2 |