EDBT 2026 Demo / reviewers in the wild / expert
Yinpei Dai
dblp:209/9564
· DBLP profile ↗
14ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-5715-1833ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Question answering and dialogue systems · 41% Language models and text generation · 13% Reinforcement learning · 12% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% |
Topics — the 25 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue |
3.4 | 6 | 2023 | SpokenWOZ: A Large-Scale Speech-Text Benchmark for Spoken Task-Oriented Dialogue Agents · NeurIPS 2023 Task-Oriented Dialogue System as Natural Language Generation · SIGIR 2022 Unified Dialog Model Pre-training for Task-Oriented Dialog Understanding and Generation · SIGIR 2022 |
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue state tracking |
1.7 | 3 | 2023 | SpokenWOZ: A Large-Scale Speech-Text Benchmark for Spoken Task-Oriented Dialogue Agents · NeurIPS 2023 Unified Dialog Model Pre-training for Task-Oriented Dialog Understanding and Generation · SIGIR 2022 Unsupervised Learning of Deterministic Dialogue Structure with Edge-Enhanced Graph Auto-Encoder · AAAI 2021 |
Natural language and speech › Language models and text generation › pre-trained language model › conversational language models
dialogue pre-training |
1.1 | 2 | 2022 | Unified Dialog Model Pre-training for Task-Oriented Dialog Understanding and Generation · SIGIR 2022 GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-supervised Learning and Explicit Policy Injection · AAAI 2022 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › plan execution
failure recovery |
0.9 | 1 | 2025 | RACER: Rich Language-Guided Failure Recovery Policies for Imitation Learning · ICRA 2025 |
Machine learning › Reinforcement learning
imitation learning |
0.9 | 1 | 2025 | RACER: Rich Language-Guided Failure Recovery Policies for Imitation Learning · ICRA 2025 |
Robotics › Motion planning and robot control › robot learning › manipulation learning
language-conditioned manipulation |
0.9 | 1 | 2025 | RACER: Rich Language-Guided Failure Recovery Policies for Imitation Learning · ICRA 2025 |
Robotics › Motion planning and robot control
robot learning |
0.9 | 1 | 2025 | RACER: Rich Language-Guided Failure Recovery Policies for Imitation Learning · ICRA 2025 |
Machine learning › Reinforcement learning
embodied agent training |
0.8 | 1 | 2024 | Teaching Embodied Reinforcement Learning Agents: Informativeness and Diversity of Language Use · EMNLP 2024 |
Robotics › Robot navigation and mapping › social navigation › human-aware navigation
interactive navigation |
0.8 | 1 | 2024 | Think, Act, and Ask: Open-World Interactive Personalized Robot Navigation · ICRA 2024 |
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
language-conditioned reinforcement learning |
0.8 | 1 | 2024 | Teaching Embodied Reinforcement Learning Agents: Informativeness and Diversity of Language Use · EMNLP 2024 |
Robotics › Robot navigation and mapping
object goal navigation |
0.8 | 1 | 2024 | Think, Act, and Ask: Open-World Interactive Personalized Robot Navigation · ICRA 2024 |
Robotics › Robot navigation and mapping › object goal navigation
zero-shot object navigation |
0.8 | 1 | 2024 | Think, Act, and Ask: Open-World Interactive Personalized Robot Navigation · ICRA 2024 |
Human-AI interaction › large language model interaction › language-based interaction
natural language interface |
0.8 | 1 | 2024 | Think, Act, and Ask: Open-World Interactive Personalized Robot Navigation · ICRA 2024 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.7 | 2 | 2022 | GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-supervised Learning and Explicit Policy Injection · AAAI 2022 Unified Dialog Model Pre-training for Task-Oriented Dialog Understanding and Generation · SIGIR 2022 |
Natural language and speech › Question answering and dialogue systems
spoken dialogue systems |
0.7 | 1 | 2023 | SpokenWOZ: A Large-Scale Speech-Text Benchmark for Spoken Task-Oriented Dialogue Agents · NeurIPS 2023 |
Natural language and speech › Speech recognition and synthesis
spoken language understanding |
0.7 | 1 | 2023 | SpokenWOZ: A Large-Scale Speech-Text Benchmark for Spoken Task-Oriented Dialogue Agents · NeurIPS 2023 |
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
adapter tuning |
0.6 | 1 | 2022 | Task-Oriented Dialogue System as Natural Language Generation · SIGIR 2022 |
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue policy learning |
0.6 | 1 | 2022 | GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-supervised Learning and Explicit Policy Injection · AAAI 2022 |
Natural language and speech › Question answering and dialogue systems › dialogue generation
dialogue response generation |
0.6 | 1 | 2022 | Task-Oriented Dialogue System as Natural Language Generation · SIGIR 2022 |
Robotics › Autonomous driving
intention prediction |
0.6 | 1 | 2022 | Unified Dialog Model Pre-training for Task-Oriented Dialog Understanding and Generation · SIGIR 2022 |
Natural language and speech › Language models and text generation › large language model › large language model adaptation
pre-trained language model fine-tuning |
0.6 | 1 | 2022 | Task-Oriented Dialogue System as Natural Language Generation · SIGIR 2022 |
Natural language and speech › Question answering and dialogue systems › dialogue modeling
dialogue structure modeling |
0.5 | 1 | 2021 | Unsupervised Learning of Deterministic Dialogue Structure with Edge-Enhanced Graph Auto-Encoder · AAAI 2021 |
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
user simulation |
0.5 | 1 | 2021 | Transferable Dialogue Systems and User Simulators · ACL/IJCNLP (1) 2021 |
Natural language and speech › Language models and text generation › instruction following
instruction-following language models |
0.2 | 1 | 2024 | Teaching Embodied Reinforcement Learning Agents: Informativeness and Diversity of Language Use · EMNLP 2024 |
Machine learning › Learning paradigms
semi-supervised learning |
0.2 | 1 | 2022 | GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-supervised Learning and Explicit Policy Injection · AAAI 2022 |
Methods — techniques the papers use, named apart from their topics
large language model · 1.5transfer learning · 1.1vision-language model · 0.9supervisor-actor framework · 0.9reinforcement learning · 0.8dual-modal models · 0.7baseline · 0.7LLM · 0.7dialog act prediction · 0.6consistency regularization · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RACER: Rich Language-Guided Failure Recovery Policies for Imitation LearningabstractDeveloping robust and correctable visuomotor policies for robotic manipulation is challenging due to the lack of self-recovery mechanisms from failures and the limitations of simple language instructions in guiding robot actions. To address these issues, we propose a scalable data generation pipeline that automatically augments expert demonstrations with failure recovery trajectories and fine-grained language annotations for training. We then introduce Rich languAge-guided failure reCovERy (RACER), a supervisor-actor frame-work, which combines failure recovery data with rich language descriptions to enhance robot control. RACER features a vision-language model (VLM) that acts as an online supervisor, providing detailed language guidance for error correction and task execution, and a language-conditioned visuomotor policy as an actor to predict the next actions. Our experimental results show that RACER outperforms the state-of-the-art Robotic View Transformer (RVT) on RLbench across various evaluation settings, including standard long-horizon tasks, dynamic goal-change tasks and zero-shot unseen tasks, achieving superior performance in both simulated and real world environments. Videos and code are available at: https://rich-language-failure-recovery.github.io. Yinpei Dai, Jayjun Lee, Nima Fazeli, Joyce Y. Chai |
ICRA | 1 |
| 2024 | Teaching Embodied Reinforcement Learning Agents: Informativeness and Diversity of Language UseabstractIn real-world scenarios, it is desirable for embodied agents to have the ability to leverage human language to gain explicit or implicit knowledge for learning tasks.Despite recent progress, most previous approaches adopt simple low-level instructions as language inputs, which may not reflect natural human communication.It's not clear how to incorporate rich language use to facilitate task learning.To address this question, this paper studies different types of language inputs in facilitating reinforcement learning (RL) embodied agents.More specifically, we examine how different levels of language informativeness (i.e., feedback on past behaviors and future guidance) and diversity (i.e., variation of language expressions) impact agent learning and inference.Our empirical results based on four RL benchmarks demonstrate that agents trained with diverse and informative language feedback can achieve enhanced generalization and fast adaptation to new tasks.These findings highlight the pivotal role of language use in teaching embodied agents new tasks in an open world. 1 Jiajun Xi, Yinong He, Yinpei Dai, Joyce Y. Chai |
EMNLP | 4 |
| 2024 | Think, Act, and Ask: Open-World Interactive Personalized Robot NavigationabstractZero-Shot Object Navigation (ZSON) enables agents to navigate towards open-vocabulary objects in unknown environments. The existing works of ZSON mainly focus on following individual instructions to find generic object classes, neglecting the utilization of natural language interaction and the complexities of identifying user-specific objects. To address these limitations, we introduce Zero-shot Interactive Personalized Object Navigation (ZIPON), where robots need to navigate to personalized goal objects while engaging in conversations with users. To solve ZIPON, we propose a new framework termed Open-woRld Interactive persOnalized Navigation (ORION)1, which uses Large Language Models (LLMs) to make sequential decisions to manipulate different modules for perception, navigation and communication. Experimental results show that the performance of interactive agents that can leverage user feedback exhibits significant improvement. However, obtaining a good balance between task completion and the efficiency of navigation and interaction remains challenging for all methods. We further provide more findings on the impact of diverse user feedback forms on the agents’ performance. Yinpei Dai, Run Peng, Sikai Li, Joyce Y. Chai |
ICRA | 1 |
| 2023 | SpokenWOZ: A Large-Scale Speech-Text Benchmark for Spoken Task-Oriented Dialogue AgentsabstractTask-oriented dialogue (TOD) models have made significant progress in recent years. However, previous studies primarily focus on datasets written by annotators, which has resulted in a gap between academic research and real-world spoken con- versation scenarios. While several small-scale spoken TOD datasets are proposed to address robustness issues such as ASR errors, they ignore the unique challenges in spoken conversation. To tackle the limitations, we introduce SpokenWOZ, a large-scale speech-text dataset for spoken TOD, containing 8 domains, 203k turns, 5.7k dialogues and 249 hours of audios from human-to-human spoken conversations. SpokenWOZ further incorporates common spoken characteristics such as word-by-word processing and reasoning in spoken language. Based on these characteristics, we present cross-turn slot and reasoning slot detection as new challenges. We conduct experiments on various baselines, including text-modal models, newly proposed dual-modal models, and LLMs, e.g., ChatGPT. The results show that the current models still have substantial room for improvement in spoken conversation, where the most advanced dialogue state tracker only achieves 25.65% in joint goal accuracy and the SOTA end-to-end model only correctly completes the user request in 52.1% of dialogues. Our dataset, code, and leaderboard are available at https://spokenwoz.github.io/SpokenWOZ-github.io/. Shuzheng Si, Yuchuan Wu, Ting-En Lin, Yinpei Dai, Hangyu Li 0003, Rui Yan 0001, Fei Huang 0002, Yongbin Li 0001 |
NeurIPS | 6 |
| 2022 | GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-supervised Learning and Explicit Policy InjectionabstractPre-trained models have proved to be powerful in enhancing task-oriented dialog systems. However, current pre-training methods mainly focus on enhancing dialog understanding and generation tasks while neglecting the exploitation of dialog policy. In this paper, we propose GALAXY, a novel pre-trained dialog model that explicitly learns dialog policy from limited labeled dialogs and large-scale unlabeled dialog corpora via semi-supervised learning. Specifically, we introduce a dialog act prediction task for policy optimization during pre-training and employ a consistency regularization term to refine the learned representation with the help of unlabeled dialogs. We also implement a gating mechanism to weigh suitable unlabeled dialog samples. Empirical results show that GALAXY substantially improves the performance of task-oriented dialog systems, and achieves new state-of-the-art results on benchmark datasets: In-Car, MultiWOZ2.0 and MultiWOZ2.1, improving their end-to-end combined scores by 2.5, 5.3 and 5.5 points, respectively. We also show that GALAXY has a stronger few-shot ability than existing models under various low-resource settings. For reproducibility, we release the code and data at https://github.com/siat-nlp/GALAXY. Wanwei He, Yinpei Dai, Yinhe Zheng, Yuchuan Wu, Zheng Cao 0003, Dermot Liu, Min Yang 0007, Fei Huang 0002, Luo Si, Jian Sun 0021, Yongbin Li 0001 |
AAAI | 2 |
| 2022 | SPACE-2: Tree-Structured Semi-Supervised Contrastive Pre-training for Task-Oriented Dialog UnderstandingabstractPre-training methods with contrastive learning objectives have shown remarkable success in dialog understanding tasks. However, current contrastive learning solely considers the self-augmented dialog samples as positive samples and treats all other dialog samples as negative ones, which enforces dissimilar representations even for dialogs that are semantically related. In this paper, we propose SPACE-2, a tree-structured pre-trained conversation model, which learns dialog representations from limited labeled dialogs and large-scale unlabeled dialog corpora via semi-supervised contrastive pre-training. Concretely, we first define a general semantic tree structure (STS) to unify the inconsistent annotation schema across different dialog datasets, so that the rich structural information stored in all labeled data can be exploited. Then we propose a novel multi-view score function to increase the relevance of all possible dialogs that share similar STSs and only push away other completely different dialogs during supervised contrastive pre-training. To fully exploit unlabeled dialogs, a basic self-supervised contrastive loss is also added to refine the learned representations. Experiments show that our method can achieve new state-of-the-art results on the DialoGLUE benchmark consisting of seven datasets and four popular dialog understanding tasks. Wanwei He, Yinpei Dai, Binyuan Hui, Min Yang 0007, Zheng Cao 0003, Jianbo Dong, Fei Huang 0002, Luo Si, Yongbin Li 0001 |
COLING | 2 |
| 2022 | CGoDial: A Large-Scale Benchmark for Chinese Goal-oriented Dialog EvaluationabstractPractical dialog systems need to deal with various knowledge sources, noisy user expressions, and the shortage of annotated data.To better solve the above problems, we propose CGoDial 1 , a new challenging and comprehensive Chinese benchmark for multi-domain Goal-oriented Dialog evaluation.It contains 96,763 dialog sessions, and 574,949 dialog turns totally, covering three datasets with different knowledge sources: 1) a slot-based dialog (SBD) dataset with table-formed knowledge, 2) a flow-based dialog (FBD) dataset with treeformed knowledge, and a retrieval-based dialog (RBD) dataset with candidate-formed knowledge.To bridge the gap between academic benchmarks and spoken dialog scenarios, we either collect data from real conversations or add spoken features to existing datasets via crowdsourcing.The proposed experimental settings include the combinations of training with either the entire training set or a few-shot training set, and testing with either the standard test set or a hard test subset, which can assess model capabilities in terms of general prediction, fast adaptability and reliable robustness. Yinpei Dai, Wanwei He, Bowen Li 0002, Yuchuan Wu, Zheng Cao 0003, Zhongqi An, Jian Sun 0021, Yongbin Li 0001 |
EMNLP | 1 |
| 2022 | Unified Dialog Model Pre-training for Task-Oriented Dialog Understanding and GenerationabstractRecently, pre-training methods have shown remarkable success in task-oriented dialog (TOD) systems. However, most existing pre-trained models for TOD focus on either dialog understanding or dialog generation, but not both. In this paper, we propose SPACE, a novel unified pre-trained dialog model learning from large-scale dialog corpora with limited annotations, which can be effectively fine-tuned on a wide range of downstream dialog tasks. Specifically, SPACE consists of four successive components in a single transformer to maintain a task-flow in TOD systems: (i) a dialog encoding module to encode dialog history, (ii) a dialog understanding module to extract semantic vectors from either user queries or system responses, (iii) a dialog policy module to generate a policy vector that contains high-level semantics of the response, and (iv) a dialog generation module to produce appropriate responses. We design a dedicated pre-training objective for each component. Concretely, we pre-train the dialog encoding module with span mask language modeling to learn contextualized dialog information. To capture the structured dialog semantics, we pre-train the dialog understanding module via a novel tree-induced semi-supervised contrastive learning objective with the help of extra dialog annotations. In addition, we pre-train the dialog policy module by minimizing the ℒ2 distance between its output policy vector and the semantic vector of the response for policy optimization. Finally, the dialog generation model is pre-trained by language modeling. Results show that SPACE achieves state-of-the-art performance on eight downstream dialog benchmarks, including intent prediction, dialog state tracking, and end-to-end dialog modeling. We also show that SPACE has a stronger few-shot ability than existing models under the low-resource setting. Wanwei He, Yinpei Dai, Min Yang 0007, Jian Sun 0021, Fei Huang 0002, Luo Si, Yongbin Li 0001 |
SIGIR | 2 |
| 2022 | Task-Oriented Dialogue System as Natural Language GenerationabstractIn this paper, we propose to formulate the task-oriented dialogue system as the purely natural language generation task, so as to fully leverage the large-scale pre-trained models like GPT-2 and simplify complicated delexicalization prepossessing. However, directly applying this method heavily suffers from the dialogue entity inconsistency caused by the removal of delexicalized tokens, as well as the catastrophic forgetting problem of the pre-trained model during fine-tuning, leading to unsatisfactory performance. To alleviate these problems, we design a novel GPT-Adapter-CopyNet network, which incorporates the lightweight adapter and CopyNet modules into GPT-2 to achieve better performance on transfer learning and dialogue entity generation. Experimental results conducted on the DSTC8 Track 1 benchmark and MultiWOZ dataset demonstrate that our proposed approach significantly outperforms baseline models with a remarkable performance on automatic and human evaluations. Weizhi Wang, Zhirui Zhang, Junliang Guo, Yinpei Dai, Boxing Chen, Weihua Luo |
SIGIR | 4 |
| 2021 | Unsupervised Learning of Deterministic Dialogue Structure with Edge-Enhanced Graph Auto-EncoderabstractIt is important for task-oriented dialogue systems to discover the dialogue structure (i.e. the general dialogue flow) from dialogue corpora automatically. Previous work models dialogue structure by extracting latent states for each utterance first and then calculating the transition probabilities among states. These two-stage methods ignore the contextual information when calculating the probabilities, which makes the transitions between the states ambiguous. This paper proposes a conversational graph (CG) to represent deterministic dialogue structure where nodes and edges represent the utterance and context information respectively. An unsupervised Edge-Enhanced Graph Auto-Encoder (EGAE) architecture is designed to model local-contextual and global-structural information for conversational graph learning. Furthermore, a self-supervised objective is introduced with the response selection task to guide the unsupervised learning of the dialogue structure. Experimental results on several public datasets demonstrate that the novel model outperforms several alternatives in aggregating utterances with similar semantics. The effectiveness of the learned dialogue structured is also verified by more than 5\% joint accuracy improvement in the downstream task of low resource dialogue state tracking. Yajing Sun, Yong Shan, Chengguang Tang, Yue Hu 0002, Yinpei Dai, Jing Yu 0007, Jian Sun 0021, Fei Huang 0002, Luo Si |
AAAI | 5 |
| 2021 | Transferable Dialogue Systems and User SimulatorsabstractBo-Hsiang Tseng, Yinpei Dai, Florian Kreyssig, Bill Byrne. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Bo-Hsiang Tseng, Yinpei Dai, Florian Kreyssig, William J. Byrne |
ACL/IJCNLP (1) | 2 |
| 2020 | Learning Low-Resource End-To-End Goal-Oriented Dialog for Fast and Reliable System DeploymentabstractExisting end-to-end dialog systems perform less effectively when data is scarce. To obtain an acceptable success in real-life online services with only a handful of training examples, both fast adaptability and reliable performance are highly desirable for dialog systems. In this paper, we propose the Meta-Dialog System (MDS), which combines the advantages of both meta-learning approaches and human-machine collaboration. We evaluate our methods on a new extended-bAbI dataset and a transformed MultiWOZ dataset for low-resource goal-oriented dialog learning. Experimental results show that MDS significantly outperforms non-meta-learning baselines and can achieve more than 90% per-turn accuracies with only 10 dialogs on the extended-bAbI dataset. Yinpei Dai, Hangyu Li 0003, Chengguang Tang, Yongbin Li 0001, Jian Sun 0021, Xiaodan Zhu 0001 |
ACL | 1 |
| 2020 | Improved Learning of Word Embeddings with Word Definitions and Semantic Injection
Yichi Zhang 0001, Yinpei Dai, Zhijian Ou, Huixin Wang, Junlan Feng |
INTERSPEECH | 2 |
| 2018 | Tracking of Enriched Dialog States for Flexible Conversational Information AccessabstractDialog state tracking (DST) is a crucial component in a task-oriented dialog system for conversational information access. A common practice in current dialog systems is to define the dialog state by a set of slot-value pairs. Such representation of dialog states and the slot-filling based DST have been widely employed, but suffer from three drawbacks. (1) The dialog state can contain only a single value for a slot, and (2) can contain only users' affirmative preference over the values for a slot. (3) Current task-based dialog systems mainly focus on the searching task, while the enquiring task is also very common in practice. The above observations motivate us to enrich current representation of dialog states and collect a brand new dialog dataset about movies, based upon which we build a new DST, called enriched DST (EDST), for flexible movie information access. The EDST supports the searching task, the enquiring task and their mixed task. We show that our new EDST method not only achieves good results on Iqiyi dataset, but also outperforms other state-of-the-art DST methods on the traditional dialog datasets, WOZ2.0 and DSTC2. Yinpei Dai, Zhijian Ou, Dawei Ren, Pengfei Yu 0001 |
ICASSP | 1 |