VLDB 2026 Research / reviewers in the wild / expert
Yujia Qin
dblp:126/2333
· DBLP profile ↗
23ranked-venue papers
8as first author
21since 2021 · last 2025
0000-0003-3608-5061ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 7 first-author · 19 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHubabstractBohan Lyu, Xin Cong, Heyang Yu, Pan Yang, Cheng Qian, Zihe Wang, Yujia Qin, Yining Ye, Yaxi Lu, Chen Qian, Zhong Zhang, Yukun Yan, Yankai Lin, Zhiyuan Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Bohan Lyu 0001, Xin Cong, Heyang Yu, Pan Yang 0022, Cheng Qian 0008, Yujia Qin, Yining Ye, Yaxi Lu, Zhong Zhang 0004, Yukun Yan, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 7 |
| 2025 | GUICourse: From General Vision Language Model to Versatile GUI AgentabstractUtilizing Graphic User Interfaces (GUIs) for human-computer interaction is essential for accessing various digital tools. Recent advancements in Vision Language Models (VLMs) reveal significant potential for developing versatile agents that assist humans in navigating GUIs. However, current VLMs face challenges related to fundamental abilities, such as OCR and grounding, as well as a lack of knowledge about GUI elements functionalities and control methods. These limitations hinder their effectiveness as practical GUI agents. To address these challenges, we introduce GUICourse, a series of datasets for training visual-based GUI agents using general VLMs. First, we enhance the OCR and grounding capabilities of VLMs using the GUIEnv dataset. Next, we enrich the GUI knowledge of VLMs using the GUIAct and GUIChat datasets. Our experiments demonstrate that even a small-sized GUI agent (with 3.1 billion parameters) performs effectively on both single-step and multi-step GUI tasks. We further finetune our GUI agents on other GUI tasks with different action spaces (AITW and Mind2Web), and the results show that our agents are better than their baseline VLMs. Additionally, we analyze the impact of OCR and grounding capabilities through an ablation study, revealing a positive correlation with GUI navigation ability. Wentong Chen, Junbo Cui, Jinyi Hu, Yujia Qin, Junjie Fang, Chongyi Wang, Guirong Chen, Yupeng Huo, Yuan Yao 0013, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 4 |
| 2025 | Rational Decision-Making Agent with Learning Internal Utility JudgmentabstractWith remarkable advancements, large language models (LLMs) have attracted significant efforts to develop LLM-based agents capable of executing intricate multi-step decision-making tasks. Existing approaches predominantly build upon the external performance measure to guide the decision-making process but the reliance on the external performance measure as prior is problematic in real-world scenarios, where such prior may be unavailable, flawed, or even erroneous. For genuine autonomous decision-making for LLM-based agents, it is imperative to develop rationality from their posterior experiences to judge the utility of each decision independently. In this work, we propose RaDAgent (Rational Decision-Making Agent), which fosters the development of its rationality through an iterative framework involving Experience Exploration and Utility Learning. Within this framework, Elo-based Utility Learning is devised to assign Elo scores to individual decision steps to judge their utilities via pairwise comparisons. Consequently, these Elo scores guide the decision-making process to derive optimal outcomes. Experimental results on the Game of 24, WebShop, ToolBench and RestBench datasets demonstrate RaDAgent’s superiority over baselines, achieving about 7.8% improvement on average. Besides, RaDAgent also can reduce costs (ChatGPT API calls), highlighting its effectiveness and efficiency. Yining Ye, Xin Cong, Shizuo Tian, Yujia Qin, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001 |
ICLR | 4 |
| 2024 | Tell Me More! Towards Implicit User Intention Understanding of Language Model Driven AgentsabstractCheng Qian, Bingxiang He, Zhong Zhuang, Jia Deng, Yujia Qin, Xin Cong, Zhong Zhang, Jie Zhou, Yankai Lin, Zhiyuan Liu, Maosong Sun. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Cheng Qian 0008, Bingxiang He, Zhong Zhuang, Yujia Qin, Xin Cong, Zhong Zhang 0004, Jie Zhou 0016, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 5 |
| 2024 | AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent BehaviorsabstractAutonomous agents empowered by Large Language Models (LLMs) have undergone significant improvements, enabling them to generalize across a broad spectrum of tasks. However, in real-world scenarios, cooperation among individuals is often required to enhance the efficiency and effectiveness of task accomplishment. Hence, inspired by human group dynamics, we propose a multi-agent framework AgentVerse that can effectively orchestrate a collaborative group of expert agents as a greater-than-the-sum-of-its-parts system. Our experiments demonstrate that AgentVerse can proficiently deploy multi-agent groups that outperform a single agent. Extensive experiments on text understanding, reasoning, coding, tool utilization, and embodied AI confirm the effectiveness of AgentVerse. Moreover, our analysis of agent interactions within AgentVerse reveals the emergence of specific collaborative behaviors, contributing to heightened group efficiency. We will release our codebase, AgentVerse, to further facilitate multi-agent research. Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang 0002, Chenfei Yuan, Chi-Min Chan, Heyang Yu, Yaxi Lu, Yi-Hsin Hung, Yujia Qin, Xin Cong, Ruobing Xie, Zhiyuan Liu 0001, Maosong Sun 0001, Jie Zhou 0016 |
ICLR | 11 |
| 2024 | ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsabstractDespite the advancements of open-source large language models (LLMs), e.g., LLaMA, they remain significantly limited in tool-use capabilities, i.e., using external tools (APIs) to fulfill human instructions. The reason is that current instruction tuning largely focuses on basic language tasks but ignores the tool-use domain. This is in contrast to the excellent tool-use capabilities of state-of-the-art (SOTA) closed-source LLMs, e.g., ChatGPT. To bridge this gap, we introduce ToolLLM, a general tool-use framework encompassing data construction, model training, and evaluation. We first present ToolBench, an instruction-tuning dataset for tool use, which is constructed automatically using ChatGPT. Specifically, the construction can be divided into three stages: (i) API collection: we collect 16,464 real-world RESTful APIs spanning 49 categories from RapidAPI Hub; (ii) instruction generation: we prompt ChatGPT to generate diverse instructions involving these APIs, covering both single-tool and multi-tool scenarios; (iii) solution path annotation: we use ChatGPT to search for a valid solution path (chain of API calls) for each instruction. To enhance the reasoning capabilities of LLMs, we develop a novel depth-first search-based decision tree algorithm. It enables LLMs to evaluate multiple reasoning traces and expand the search space. Moreover, to evaluate the tool-use capabilities of LLMs, we develop an automatic evaluator: ToolEval. Based on ToolBench, we fine-tune LLaMA to obtain an LLM ToolLLaMA, and equip it with a neural API retriever to recommend appropriate APIs for each instruction. Experiments show that ToolLLaMA demonstrates a remarkable ability to execute complex instructions and generalize to unseen APIs, and exhibits comparable performance to ChatGPT. Our ToolLLaMA also demonstrates strong zero-shot generalization ability in an out-of-distribution tool-use dataset: APIBench. Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin 0001, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou 0016, Mark Gerstein, Dahai Li, Zhiyuan Liu 0001, Maosong Sun 0001 |
ICLR | 1 |
| 2024 | Empowering Large Language Models: Tool Learning for Real-World InteractionabstractSince the advent of large language models (LLMs), the field of tool learning has remained very active in solving various tasks in practice, including but not limited to information retrieval. This half-day tutorial provides basic concepts of this field and an overview of recent advancements with several applications. In specific, we start with some foundational components and architecture of tool learning (i.e., cognitive tool and physical tool), and then we categorize existing studies in this field into tool-augmented learning and tool-oriented learning, and introduce various learning methods to empower LLMs this kind of capability. Furthermore, we provide several cases about when, what, and how to use tools in different applications. We end with some open challenges and several potential research directions for future studies. We believe this tutorial is suited for both researchers at different stages (introductory, intermediate, and advanced) and industry practitioners who are interested in LLMs and tool learning. Hongru Wang 0003, Yujia Qin, Yankai Lin 0001, Jeff Z. Pan, Kam-Fai Wong |
SIGIR | 2 |
| 2024 | Whole-genome bisulfite sequencing data analysis learning module on Google Cloud PlatformabstractThis study describes the development of a resource module that is part of a learning platform named 'NIGMS Sandbox for Cloud-based Learning' https://github.com/NIGMS/NIGMS-Sandbox. The overall genesis of the Sandbox is described in the editorial NIGMS Sandbox at the beginning of this Supplement. This module is designed to facilitate interactive learning of whole-genome bisulfite sequencing (WGBS) data analysis utilizing cloud-based tools in Google Cloud Platform, such as Cloud Storage, Vertex AI notebooks and Google Batch. WGBS is a powerful technique that can provide comprehensive insights into DNA methylation patterns at single cytosine resolution, essential for understanding epigenetic regulation across the genome. The designed learning module first provides step-by-step tutorials that guide learners through two main stages of WGBS data analysis, preprocessing and the identification of differentially methylated regions. And then, it provides a streamlined workflow and demonstrates how to effectively use it for large datasets given the power of cloud infrastructure. The integration of these interconnected submodules progressively deepens the user's understanding of the WGBS analysis process along with the use of cloud resources. Through this module, we can enhance the accessibility and adoption of cloud computing in epigenomic research, speeding up the advancements in the related field and beyond. This manuscript describes the development of a resource module that is part of a learning platform named ``NIGMS Sandbox for Cloud-based Learning'' https://github.com/NIGMS/NIGMS-Sandbox. The overall genesis of the Sandbox is described in the editorial NIGMS Sandbox [1] at the beginning of this Supplement. This module delivers learning materials on the analysis of bulk and single-cell ATAC-seq data in an interactive format that uses appropriate cloud resources for data access and analyses. Yujia Qin, Angela Maggio, Dale Hawkins, Laura Beaudry, Allen Kim, Daniel Pan, Youping Deng |
Briefings Bioinform. | 1 |
| 2024 | Exploring Universal Intrinsic Task Subspace for Few-Shot Learning via Prompt TuningabstractWhy can pre-trained language models (PLMs) learn universal representations and effectively adapt to broad NLP tasks differing a lot superficially? In this work, we empirically find evidence indicating that the adaptations of PLMs to various few-shot tasks can be reparameterized as optimizing only a few free parameters in a unified low-dimensionalintrinsic task subspace, which may help us understand why PLMs could easily adapt to various NLP tasks with small-scale data. To find such a subspace and examine its universality, we propose an analysis pipeline calledintrinsic prompt tuning(IPT). Specifically, we resort to the recent success of prompt tuning and decompose the soft prompts of multiple NLP tasks into the same low-dimensional nonlinear subspace, then we learn to adapt the PLM to unseen data or tasks by only tuning parameters in this subspace. In the experiments, we study diverse few-shot NLP tasks and surprisingly find that in a 250-dimensional subspace found with 100 tasks, by only tuning 250 free parameters, we can recover 97% and 83% of the full prompt tuning performance for 100 seen tasks (using different training data) and 20 unseen tasks, respectively, showing great generalization ability of the found intrinsic task subspace. Besides being an analysis tool, IPTcould further help us improve the prompt tuning stability. Yujia Qin, Xiaozhi Wang, Yusheng Su, Yankai Lin 0001, Ning Ding 0002, Jing Yi, Weize Chen, Zhiyuan Liu 0001, Juan-Zi Li, Lei Hou 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | WebCPM: Interactive Web Search for Chinese Long-form Question AnsweringabstractYujia Qin, Zihan Cai, Dian Jin, Lan Yan, Shihao Liang, Kunlun Zhu, Yankai Lin, Xu Han, Ning Ding, Huadong Wang, Ruobing Xie, Fanchao Qi, Zhiyuan Liu, Maosong Sun, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yujia Qin, Zihan Cai, Lan Yan, Shihao Liang, Kunlun Zhu, Yankai Lin 0001, Xu Han 0007, Ning Ding 0002, Ruobing Xie, Fanchao Qi, Zhiyuan Liu 0001, Maosong Sun 0001, Jie Zhou 0024 |
ACL (1) | 1 |
| 2023 | Enhancing Chat Language Models by Scaling High-quality Instructional ConversationsabstractFine-tuning on instruction data has been widely validated as an effective practice for implementing chat language models like ChatGPT.Scaling the diversity and quality of such data, although straightforward, stands a great chance of leading to improved performance.This paper aims to push the upper bound of opensource models further.We first provide a systematically designed, diverse, informative, large-scale dataset of instructional conversations, UltraChat, which does not involve human queries.Our objective is to capture the breadth of interactions between a human user and an AI assistant and employs a comprehensive framework to generate multi-turn conversation iteratively.UltraChat contains 1.5 million high-quality multi-turn dialogues and covers a wide range of topics and instructions.Our statistical analysis of UltraChat reveals its superiority in various key metrics, including scale, average length, diversity, coherence, etc., solidifying its position as a leading opensource dataset.Building upon UltraChat, we fine-tune a LLaMA model to create a powerful conversational model, UltraLM.Our evaluations indicate that UltraLM consistently outperforms other open-source models, including WizardLM and Vicuna, the previously recognized state-of-the-art open-source models. Ning Ding 0002, Yulin Chen 0001, Bokai Xu, Yujia Qin, Shengding Hu, Zhiyuan Liu 0001, Maosong Sun 0001, Bowen Zhou 0002 |
EMNLP | 4 |
| 2023 | Exploring the Impact of Model Scaling on Parameter-Efficient TuningabstractYusheng Su, Chi-Min Chan, Jiali Cheng, Yujia Qin, Yankai Lin, Shengding Hu, Zonghan Yang, Ning Ding, Xingzhi Sun, Guotong Xie, Zhiyuan Liu, Maosong Sun. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Yusheng Su, Chi-Min Chan, Jiali Cheng, Yujia Qin, Yankai Lin 0001, Shengding Hu, Zonghan Yang, Ning Ding 0002, Xingzhi Sun 0002, Guo Tong Xie, Zhiyuan Liu 0001, Maosong Sun 0001 |
EMNLP | 4 |
| 2023 | Removing Backdoors in Pre-trained Models by Regularized Continual Pre-trainingabstractAbstract Recent research has revealed that pre-trained models (PTMs) are vulnerable to backdoor attacks before the fine-tuning stage. The attackers can implant transferable task-agnostic backdoors in PTMs, and control model outputs on any downstream task, which poses severe security threats to all downstream applications. Existing backdoor-removal defenses focus on task-specific classification models and they are not suitable for defending PTMs against task-agnostic backdoor attacks. To this end, we propose the first task-agnostic backdoor removal method for PTMs. Based on the selective activation phenomenon in backdoored PTMs, we design a simple and effective backdoor eraser, which continually pre-trains the backdoored PTMs with a regularization term in an end-to-end approach. The regularization term removes backdoor functionalities from PTMs while the continual pre-training maintains the normal functionalities of PTMs. We conduct extensive experiments on pre-trained models across different modalities and architectures. The experimental results show that our method can effectively remove backdoors inside PTMs and preserve benign functionalities of PTMs with a few downstream-task-irrelevant auxiliary data, e.g., unlabeled plain texts. The average attack success rate on three downstream datasets is reduced from 99.88% to 8.10% after our defense on the backdoored BERT. The codes are publicly available at https://github.com/thunlp/RECIPE. Biru Zhu, Ganqu Cui, Yangyi Chen, Yujia Qin, Lifan Yuan, Yangdong Deng, Zhiyuan Liu 0001, Maosong Sun 0001, Ming Gu 0001 |
Trans. Assoc. Comput. Linguistics | 4 |
| 2022 | bert2BERT: Towards Reusable Pretrained Language ModelsabstractCheng Chen, Yichun Yin, Lifeng Shang, Xin Jiang, Yujia Qin, Fengyu Wang, Zhi Wang, Xiao Chen, Zhiyuan Liu, Qun Liu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yichun Yin, Lifeng Shang, Xin Jiang 0002, Yujia Qin, Zhi Wang 0001, Xiao Chen 0012, Zhiyuan Liu 0001, Qun Liu 0001 |
ACL (1) | 5 |
| 2022 | Pass off Fish Eyes for Pearls: Attacking Model Selection of Pre-trained ModelsabstractBiru Zhu, Yujia Qin, Fanchao Qi, Yangdong Deng, Zhiyuan Liu, Maosong Sun, Ming Gu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Biru Zhu, Yujia Qin, Fanchao Qi, Yangdong Deng, Zhiyuan Liu 0001, Maosong Sun 0001, Ming Gu 0001 |
ACL (1) | 2 |
| 2022 | Exploring Mode Connectivity for Pre-trained Language ModelsabstractRecent years have witnessed the prevalent application of pre-trained language models (PLMs) in NLP.From the perspective of parameter space, PLMs provide generic initialization, starting from which high-performance minima could be found.Although plenty of works have studied how to effectively and efficiently adapt PLMs to high-performance minima, little is known about the connection of various minima reached under different adaptation configurations.In this paper, we investigate the geometric connections of different minima through the lens of mode connectivity, which measures whether two minima can be connected with a low-loss path.We conduct empirical analyses to investigate three questions: (1) how could hyperparameters, specific tuning methods, and training data affect PLM's mode connectivity?(2) How does mode connectivity change during pretraining?(3) How does the PLM's task knowledge change along the path connecting two minima?In general, exploring the mode connectivity of PLMs conduces to understanding the geometric connection of different minima, which may help us fathom the inner workings of PLM downstream adaptation.The codes are publicly available at https://github.com/ thunlp/Mode-Connectivity-PLM. Yujia Qin, Cheng Qian 0008, Jing Yi, Weize Chen, Yankai Lin 0001, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001, Jie Zhou 0016 |
EMNLP | 1 |
| 2022 | Knowledge Inheritance for Pre-trained Language ModelsabstractYujia Qin, Yankai Lin, Jing Yi, Jiajie Zhang, Xu Han, Zhengyan Zhang, Yusheng Su, Zhiyuan Liu, Peng Li, Maosong Sun, Jie Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yujia Qin, Yankai Lin 0001, Jing Yi, Xu Han 0007, Zhengyan Zhang, Yusheng Su, Zhiyuan Liu 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016 |
NAACL-HLT | 1 |
| 2022 | On Transferability of Prompt Tuning for Natural Language ProcessingabstractYusheng Su, Xiaozhi Wang, Yujia Qin, Chi-Min Chan, Yankai Lin, Huadong Wang, Kaiyue Wen, Zhiyuan Liu, Peng Li, Juanzi Li, Lei Hou, Maosong Sun, Jie Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yusheng Su, Xiaozhi Wang, Yujia Qin, Chi-Min Chan, Yankai Lin 0001, Kaiyue Wen, Zhiyuan Liu 0001, Peng Li 0030, Juan-Zi Li, Lei Hou 0001, Maosong Sun 0001, Jie Zhou 0016 |
NAACL-HLT | 3 |
| 2022 | ProQA: Structural Prompt-based Pre-training for Unified Question AnsweringabstractWanjun Zhong, Yifan Gao, Ning Ding, Yujia Qin, Zhiyuan Liu, Ming Zhou, Jiahai Wang, Jian Yin, Nan Duan. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Wanjun Zhong, Yifan Gao 0001, Ning Ding 0002, Yujia Qin, Zhiyuan Liu 0001, Ming Zhou 0001, Jiahai Wang, Jian Yin 0001, Nan Duan 0001 |
NAACL-HLT | 4 |
| 2022 | Moderate-fitting as a Natural Backdoor Defender for Pre-trained Language ModelsabstractDespite the great success of pre-trained language models (PLMs) in a large set of natural language processing (NLP) tasks, there has been a growing concern about their security in real-world applications. Backdoor attack, which poisons a small number of training samples by inserting backdoor triggers, is a typical threat to security. Trained on the poisoned dataset, a victim model would perform normally on benign samples but predict the attacker-chosen label on samples containing pre-defined triggers. The vulnerability of PLMs under backdoor attacks has been proved with increasing evidence in the literature. In this paper, we present several simple yet effective training strategies that could effectively defend against such attacks. To the best of our knowledge, this is the first work to explore the possibility of backdoor-free adaptation for PLMs. Our motivation is based on the observation that, when trained on the poisoned dataset, the PLM's adaptation follows a strict order of two stages: (1) a moderate-fitting stage, where the model mainly learns the major features corresponding to the original task instead of subsidiary features of backdoor triggers, and (2) an overfitting stage, where both features are learned adequately. Therefore, if we could properly restrict the PLM's adaptation to the moderate-fitting stage, the model would neglect the backdoor triggers but still achieve satisfying performance on the original task. To this end, we design three methods to defend against backdoor attacks by reducing the model capacity, training epochs, and learning rate, respectively. Experimental results demonstrate the effectiveness of our methods in defending against several representative NLP backdoor attacks. We also perform visualization-based analysis to attain a deeper understanding of how the model learns different features, and explore the effect of the poisoning ratio. Finally, we explore whether our methods could defend against backdoor attacks for the pre-trained CV model. The codes are publicly available at https://github.com/thunlp/Moderate-fitting. Biru Zhu, Yujia Qin, Ganqu Cui, Yangyi Chen, Weilin Zhao, Yangdong Deng, Zhiyuan Liu 0001, Jingang Wang, Wei Wu 0014, Maosong Sun 0001, Ming Gu 0001 |
NeurIPS | 2 |
| 2021 | ERICA: Improving Entity and Relation Understanding for Pre-trained Language Models via Contrastive LearningabstractYujia Qin, Yankai Lin, Ryuichi Takanobu, Zhiyuan Liu, Peng Li, Heng Ji, Minlie Huang, Maosong Sun, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yujia Qin, Yankai Lin 0001, Ryuichi Takanobu, Zhiyuan Liu 0001, Peng Li 0030, Heng Ji 0001, Minlie Huang, Maosong Sun 0001, Jie Zhou 0016 |
ACL/IJCNLP (1) | 1 |
| 2020 | Learning from Explanations with Neural Execution Tree
Ziqi Wang 0003, Yujia Qin, Wenxuan Zhou 0002, Jun Yan 0012, Qinyuan Ye, Leonardo Neves, Zhiyuan Liu 0001, Xiang Ren 0001 |
ICLR | 2 |
| 2020 | Improving Sequence Modeling Ability of Recurrent Neural Networks via SememesabstractSememes, the minimum semantic units of human languages, have been successfully utilized in various natural language processing applications. However, most existing studies exploit sememes in specific tasks and few efforts are made to utilize sememes more fundamentally. In this paper, we propose to incorporate sememes into recurrent neural networks (RNNs) to improve their sequence modeling ability, which is beneficial to all kinds of downstream tasks. We design three different sememe incorporation methods and employ them in typical RNNs including LSTM, GRU and their bidirectional variants. In evaluation, we use several benchmark datasets involving PTB and WikiText-2 for language modeling, SNLI for natural language inference and another two datasets for sentiment analysis and paraphrase detection. Experimental results show evident and consistent improvement of our sememe-incorporated models compared with vanilla RNNs, which proves the effectiveness of our sememe incorporation methods. Moreover, we find the sememe-incorporated models have higher robustness and outperform adversarial training in defending adversarial attack. All the code and data of this work can be obtained at https://github.com/thunlp/SememeRNN. Yujia Qin, Fanchao Qi, Sicong Ouyang, Zhiyuan Liu 0001, Cheng Yang 0002, Yasheng Wang, Qun Liu 0001, Maosong Sun 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |