VLDB 2026 Research / reviewers in the wild / expert
Shaofeng Yin
dblp:221/1497
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Reinforcement learning · 37% Image recognition and object detection · 23% Vision and language · 12% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
human-object interaction detection |
2.4 | 3 | 2025 | Open-Vocabulary Hoi Detection With Interaction-Aware Prompt and Concept Calibration · ICCV 2025 Exploring Conditional Multi-modal Prompts for Zero-Shot HOI Detection · ECCV (82) 2024 Exploring the Potential of Large Foundation Models for Open-Vocabulary HOI Detection · CVPR 2024 |
Machine learning › Reinforcement learning › model-based reinforcement learning
world model |
1.7 | 2 | 2025 | RLVR-World: Training World Models with Reinforcement Learning · NeurIPS 2025 Trajectory World Models for Heterogeneous Environments · ICML 2025 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
1.6 | 2 | 2025 | Trajectory World Models for Heterogeneous Environments · ICML 2025 iVideoGPT: Interactive VideoGPTs are Scalable World Models · NeurIPS 2024 |
Machine learning › Generative modeling
video generation |
1.1 | 2 | 2025 | RLVR-World: Training World Models with Reinforcement Learning · NeurIPS 2025 iVideoGPT: Interactive VideoGPTs are Scalable World Models · NeurIPS 2024 |
Natural language and speech › Language models and text generation › large language model reasoning
multi-step reasoning |
0.9 | 1 | 2025 | ToolVQA: A Dataset for Multi-Step Reasoning VQA with External Tools · ICCV 2025 |
Machine learning › Reinforcement learning
off-policy evaluation |
0.9 | 1 | 2025 | Trajectory World Models for Heterogeneous Environments · ICML 2025 |
Machine learning › Reinforcement learning › reward design
reinforcement learning with verifiable rewards |
0.9 | 1 | 2025 | RLVR-World: Training World Models with Reinforcement Learning · NeurIPS 2025 |
Computer vision › Vision and language
visual question answering |
0.9 | 1 | 2025 | ToolVQA: A Dataset for Multi-Step Reasoning VQA with External Tools · ICCV 2025 |
Machine learning › Transfer learning and domain adaptation
zero-shot learning |
0.9 | 1 | 2025 | Open-Vocabulary Hoi Detection With Interaction-Aware Prompt and Concept Calibration · ICCV 2025 |
Computer vision › Vision and language
vision-language model |
0.8 | 1 | 2024 | Exploring the Potential of Large Foundation Models for Open-Vocabulary HOI Detection · CVPR 2024 |
Computer vision › Image recognition and object detection › object detection › open-vocabulary object detection
zero-shot object detection |
0.8 | 1 | 2024 | Exploring Conditional Multi-modal Prompts for Zero-Shot HOI Detection · ECCV (82) 2024 |
Robotics › Motion planning and robot control › robot control
model predictive control |
0.3 | 1 | 2025 | Trajectory World Models for Heterogeneous Environments · ICML 2025 |
Natural language and speech › Language models and text generation › neural language model
autoregressive transformer |
0.2 | 1 | 2024 | iVideoGPT: Interactive VideoGPTs are Scalable World Models · NeurIPS 2024 |
Computer vision › Video understanding and tracking › dynamic scene analysis › video scene understanding
human-centric scene understanding |
0.2 | 1 | 2024 | Exploring the Potential of Large Foundation Models for Open-Vocabulary HOI Detection · CVPR 2024 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 0.9reinforcement learning with verifiable rewards · 0.9pre-training · 0.9maximum likelihood estimation · 0.9interaction-aware prompting · 0.9in-context learning · 0.9external tools · 0.9concept calibration · 0.9large language model · 0.8bipartite matching · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Open-Vocabulary Hoi Detection With Interaction-Aware Prompt and Concept Calibration
Ting Lei 0001, Shaofeng Yin, Qingchao Chen, Yuxin Peng 0001, Yang Liu 0105 |
ICCV | 2 |
| 2025 | ToolVQA: A Dataset for Multi-Step Reasoning VQA with External Tools
Shaofeng Yin |
ICCV | 1 |
| 2025 | Trajectory World Models for Heterogeneous EnvironmentsabstractHeterogeneity in sensors and actuators across environments poses a significant challenge to building large-scale pre-trained world models on top of this low-dimensional sensor information. In this work, we explore pre-training world models for heterogeneous environments by addressing key transfer barriers in both data diversity and model flexibility. We introduce UniTraj, a unified dataset comprising over one million trajectories from 80 environments, designed to scale data while preserving critical diversity. Additionally, we propose TrajWorld, a novel architecture capable of flexibly handling varying sensor and actuator information and capturing environment dynamics in-context. Pre-training TrajWorld on UniTraj yields substantial gains in transition prediction, achieves a new state-of-the-art for off-policy evaluation, and also delivers superior online performance of model predictive control. To the best of our knowledge, this work, for the first time, demonstrates the transfer benefits of world models across heterogeneous and complex control environments. Code and data are available at https://github.com/thuml/TrajWorld. Shaofeng Yin, Jialong Wu 0001, Siqiao Huang, Xingjian Su, Jianye Hao, Mingsheng Long |
ICML | 1 |
| 2025 | RLVR-World: Training World Models with Reinforcement LearningabstractWorld models predict state transitions in response to actions and are increasingly developed across diverse modalities. However, standard training objectives such as maximum likelihood estimation (MLE) often misalign with task-specific goals of world models, i.e., transition prediction metrics like accuracy or perceptual quality. In this paper, we present RLVR-World, a unified framework that leverages reinforcement learning with verifiable rewards (RLVR) to directly optimize world models for such metrics. Despite formulating world modeling as autoregressive prediction of tokenized sequences, RLVR-World evaluates metrics of decoded predictions as verifiable rewards. We demonstrate substantial performance gains on both language- and video-based world models across domains, including text games, web navigation, and robot manipulation. Our work indicates that, beyond recent advances in reasoning language models, RLVR offers a promising post-training paradigm for enhancing the utility of generative models more broadly. Code, datasets, models, and video samples are available at the project website: https://thuml.github.io/RLVR-World. Jialong Wu 0001, Shaofeng Yin, Ningya Feng, Mingsheng Long |
NeurIPS | 2 |
| 2024 | Exploring the Potential of Large Foundation Models for Open-Vocabulary HOI DetectionabstractOpen-vocabulary human-object interaction (HOI) detection, which is concerned with the problem of detecting novel HOIs guided by natural language, is crucial for understanding human-centric scenes. However, prior zero-shot HOI detectors often employ the same levels of feature maps to model HOIs with varying distances, leading to suboptimal performance in scenes containing human-object pairs with a wide range of distances. In addition, these detectors primarily rely on category names and over-look the rich contextual information that language can provide, which is essential for capturing open vocabulary concepts that are typically rare and not well-represented by category names alone. In this paper, we introduce a novel end-to-end open vocabulary HOI detection framework with conditional multi-level decoding and fine-grained semantic enhancement (CMD-SE), harnessing the potential of Visual-Language Models (VLMs). Specifically, we propose to model human-object pairs with different distances with different levels of feature maps by incorporating a soft constraint during the bipartite matching process. Furthermore, by leveraging large language models (LLMs) such as GPT models, we exploit their extensive world knowledge to generate descriptions of human body part states for various interactions. Then we integrate the generalizable and fine- grained semantics of human body parts to improve interaction recognition. Experimental results on two datasets, SWIG-HOI and HICO-DET, demonstrate that our proposed method achieves state-of-the-art results in open vocabulary HOI detection. The code and models are available at https://github.com/ltttpku/CMD-SE-release. Ting Lei 0001, Shaofeng Yin, Yang Liu 0105 |
CVPR | 2 |
| 2024 | Exploring Conditional Multi-modal Prompts for Zero-Shot HOI Detection
Ting Lei 0001, Shaofeng Yin, Yuxin Peng 0001, Yang Liu 0105 |
ECCV (82) | 2 |
| 2024 | iVideoGPT: Interactive VideoGPTs are Scalable World ModelsabstractWorld models empower model-based agents to interactively explore, reason, and plan within imagined environments for real-world decision-making. However, the high demand for interactivity poses challenges in harnessing recent advancements in video generative models for developing world models at scale. This work introduces Interactive VideoGPT (iVideoGPT), a scalable autoregressive transformer framework that integrates multimodal signals—visual observations, actions, and rewards—into a sequence of tokens, facilitating an interactive experience of agents via next-token prediction. iVideoGPT features a novel compressive tokenization technique that efficiently discretizes high-dimensional visual observations. Leveraging its scalable architecture, we are able to pre-train iVideoGPT on millions of human and robotic manipulation trajectories, establishing a versatile foundation that is adaptable to serve as interactive world models for a wide range of downstream tasks. These include action-conditioned video prediction, visual planning, and model-based reinforcement learning, where iVideoGPT achieves competitive performance compared with state-of-the-art methods. Our work advances the development of interactive general world models, bridging the gap between generative video models and practical model-based reinforcement learning applications. Code and pre-trained models are available at https://thuml.github.io/iVideoGPT. Jialong Wu 0001, Shaofeng Yin, Ningya Feng, Dong Li 0016, Jianye Hao, Mingsheng Long |
NeurIPS | 2 |