VLDB 2026 Research / reviewers in the wild / expert
Yaojie Lu 0001
dblp:15/3214-1
· DBLP profile ↗
54ranked-venue papers
5as first author
41since 2021 · last 2026
0000-0002-5842-7715ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 4 first-author · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AI-Salesman: Towards Reliable Large Language Model Driven TelemarketingabstractGoal-driven persuasive dialogue, exemplified by applications like telemarketing, requires sophisticated multi-turn planning and strict factual faithfulness, which remains a significant challenge for even state-of-the-art Large Language Models (LLMs). A lack of task-specific data often limits previous works, and direct LLM application suffers from strategic brittleness and factual hallucination. In this paper, we first construct and release TeleSalesCorpus, the first real-world-grounded dialogue dataset for this domain. We then propose AI-Salesman, a novel framework featuring a dual-stage architecture. For the training stage, we design a Bayesian-supervised reinforcement learning algorithm that learns robust sales strategies from noisy dialogues. For the inference stage, we introduce the Dynamic Outline-Guided Agent (DOGA), which leverages a pre-built script library to provide dynamic, turn-by-turn strategic guidance. Moreover, we design a comprehensive evaluation framework that combines fine-grained metrics for key sales skills with the LLM-as-a-Judge paradigm. Experimental results demonstrate that our proposed AI-Salesman significantly outperforms baseline models in both automatic metrics and comprehensive human evaluations, showcasing its effectiveness in complex persuasive scenarios. Chunlei Xin, Xuanang Chen, Yaojie Lu 0001, Xianpei Han, Le Sun 0001, Qianlong Xie |
AAAI | 4 |
| 2026 | PaperRegister: Boosting Flexible-grained Paper Search via Hierarchical Register IndexingabstractZhuoqun Li, Xuanang Chen, Hongyu Lin, Yaojie Lu, Xianpei Han, Shanshan Jiang, Bin Dong, Le Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xuanang Chen, Yaojie Lu 0001, Xianpei Han, Shanshan Jiang 0001, Bin Dong 0003, Le Sun 0001 |
ACL (1) | 4 |
| 2026 | All Languages Matter: Understanding and Mitigating Language Bias in Multilingual RAGabstractDan Wang, Guozhao Mo, Yafei Shi, Cheng Zhang, Bo Zheng, Boxi Cao, Xuanang Chen, Yaojie Lu, Hongyu Lin, Ben He, Xianpei Han, Le Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Guozhao Mo, Bo Zheng 0007, Boxi Cao, Xuanang Chen, Yaojie Lu 0001, Ben He 0001, Xianpei Han, Le Sun 0001 |
ACL (1) | 8 |
| 2026 | Answer First, Evidence Second? Uncovering Hidden Risks in Well-Structured AI Search SummariesabstractAs search engines increasingly adopt answer-centric interfaces, AI-generated summaries are often consumed as final answers. These summaries are typically well-structured and citation-rich, creating a strong appearance of reliability that can encourage user trust, even when evidential grounding is uncertain. Motivated by this gap between appearance and evidence, we analyze 14,175 real-world queries from MS MARCO by examining Google Search AI summaries and their cited sources. We find that reliable-looking summaries mask substantial evidential failures. Despite their credible appearance, 32.31% of summaries are incorrect, and 56.16% of these errors arise even when supporting evidence exists in the cited sources. Moreover, 31.08% of summaries exhibit citation inconsistencies, with conflict risk increasing as more sources are cited and as citations appear in more prominent positions. Our findings reveal systematic, user-facing risks in answer-centric search, highlighting the need for evaluation and design practices beyond surface-level reliability signals. https://github.com/icip-cas/AISummary. Jinman Li, Xuanang Chen, Ruoxi Xu, Yaojie Lu 0001, Zecheng Fan, Xianpei Han, Le Sun 0001 |
SIGIR | 5 |
| 2025 | DOMAINEVAL: An Auto-Constructed Benchmark for Multi-Domain Code GenerationabstractCode benchmarks such as HumanEval are widely adopted to evaluate the capabilities of Large Language Models (LLMs), providing insights into their strengths and weaknesses. However, current benchmarks primarily exercise LLMs' capability on common coding tasks (e.g., bubble sort, greatest common divisor), leaving domain-specific coding tasks (e.g., computation, system, cryptography) unexplored. To fill this gap, we propose a multi-domain code benchmark, DOMAINEVAL, designed to evaluate LLMs' coding capabilities thoroughly. Our pipeline works in a fully automated manner, enabling a push-button construction from code repositories into formatted subjects under study. Interesting findings are observed by evaluating 12 representative LLMs against DOMAINEVAL. We notice that LLMs are generally good at computation tasks while falling short on cryptography and system coding tasks. The performance gap can be as much as 68.94% (80.94% - 12.0%) in some LLMs. We also observe that generating more samples can increase the overall performance of LLMs, while the domain bias may even increase. The contributions of this study include a code generation benchmark dataset DOMAINEVAL, encompassing six popular domains, a fully automated pipeline for constructing code benchmarks, and an identification of the limitations of LLMs in code generation tasks based on their performance on DOMAINEVAL, providing directions for future research improvements. Qiming Zhu, Jialun Cao, Yaojie Lu 0001, Xianpei Han, Le Sun 0001, Shing-Chi Cheung |
AAAI | 3 |
| 2025 | From Informal to Formal - Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal ProofsabstractJialun Cao, Yaojie Lu, Meiziniu Li, Haoyang Ma, Haokun Li, Mengda He, Cheng Wen, Le Sun, Hongyu Zhang, Shengchao Qin, Shing-Chi Cheung, Cong Tian. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jialun Cao, Yaojie Lu 0001, Meiziniu Li, Haokun Li, Mengda He, Cheng Wen 0002, Le Sun 0001, Hongyu Zhang 0002, Shengchao Qin, Shing-Chi Cheung, Cong Tian 0001 |
ACL (1) | 2 |
| 2025 | DeepSolution: Boosting Complex Engineering Solution Design via Tree-based Exploration and Bi-point ThinkingabstractZhuoqun Li, Haiyang Yu, Xuanang Chen, Hongyu Lin, Yaojie Lu, Fei Huang, Xianpei Han, Yongbin Li, Le Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Haiyang Yu 0003, Xuanang Chen, Yaojie Lu 0001, Fei Huang 0002, Xianpei Han, Yongbin Li 0001, Le Sun 0001 |
ACL (1) | 5 |
| 2025 | Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from ScratchabstractXueru Wen, Jie Lou, Zichao Li, Yaojie Lu, XingYu XingYu, Yuqiu Ji, Guohai Xu, Hongyu Lin, Ben He, Xianpei Han, Le Sun, Debing Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xueru Wen, Jie Lou, Yaojie Lu 0001, XingYu, Yuqiu Ji, Guohai Xu, Ben He 0001, Xianpei Han, Le Sun 0001, Debing Zhang |
ACL (1) | 4 |
| 2025 | Sparse Latents Steer Retrieval-Augmented GenerationabstractChunlei Xin, Shuheng Zhou, Huijia Zhu, Weiqiang Wang, Xuanang Chen, Xinyan Guan, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Chunlei Xin, Shuheng Zhou 0001, Huijia Zhu, Weiqiang Wang 0002, Xuanang Chen, Xinyan Guan, Yaojie Lu 0001, Xianpei Han, Le Sun 0001 |
ACL (1) | 7 |
| 2025 | CRUXEVAL-X: A Benchmark for Multilingual Code Reasoning, Understanding and ExecutionabstractCode benchmarks such as HumanEval are widely adopted to evaluate Large Language Models' (LLMs) coding capabilities. However, there is an unignorable programming language bias in existing code benchmarks - over 95% code generation benchmarks are dominated by Python, leaving the LLMs' capabilities in other programming languages such as Java and C/C++ unknown. Moreover, coding task bias is also crucial. Most benchmarks focus on code generation capability, while benchmarks for code reasoning (given input, reasoning output; and given output, reasoning input), an essential coding capability, are insufficient. Yet, constructing multi-lingual benchmarks can be expensive and labor-intensive, and codes in contest websites such as Leetcode suffer from data contamination during training. To fill this gap, we propose CRUXEVAL-X, a multi-lingual code reasoning benchmark that contains 19 programming languages. It comprises at least 600 subjects for each language, along with 19K content-consistent tests in total. In particular, the construction pipeline of CRUXEVAL-X works in a fully automated and test-guided manner, which iteratively generates and repairs based on execution feedback. Also, to cross language barriers (e.g., dynamic/static type systems in Python/C++), we formulated various transition rules between language pairs to facilitate translation. Our extensive evaluation of 24 representative LLMs reveals the correlation between language pairs. For example, TypeScript and JavaScript show a significant positive correlation, while Racket has less correlation with other languages. More interestingly, even a model trained solely on Python can achieve at most 34.4% Pass@1 in other languages, revealing the cross-language generalization of LLMs. Jialun Cao, Yaojie Lu 0001, Ming Wen 0001, Xianpei Han, Ben He 0001, Shing-Chi Cheung, Le Sun 0001 |
ACL (1) | 3 |
| 2025 | Memorizing is Not Enough: Deep Knowledge Injection Through ReasoningabstractRuoxi Xu, Yunjie Ji, Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Ben He, Yingfei Sun, Xiangang Li, Le Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ruoxi Xu, Yunjie Ji, Boxi Cao, Yaojie Lu 0001, Xianpei Han, Ben He 0001, Yingfei Sun, Xiangang Li, Le Sun 0001 |
ACL (1) | 4 |
| 2025 | Improved Sparse Upcycling for Instruction TuningabstractThe Mixture-of-Experts (MoE) architecture has demonstrated significant potential in both large-scale pre-training and instruction tuning by offering increased parameter capacity without additional inference costs. However, developing MoE models faces challenges including training instability and the need for substantial high-quality training data. While efficient methodologies like sparse upcycling exist, they often lead to performance degradation in instruction tuning scenarios. We introduce representation-based sparse upcycling, a straightforward yet effective technique for converting dense language models into sparsely activated ones while maintaining similar computational costs. Unlike conventional sparse upcycling, our approach leverages intermediate representations from language models to initialize router weights. This strategy addresses the mismatch between randomly initialized and well-trained parameters while providing prior knowledge to guide expert specialization during training. Extensive experiments across diverse benchmarks demonstrate significant improvements in both model capabilities and routing consistency compared to existing approaches. Wangyi Jiang, Yaojie Lu 0001, Xianpei Han, Le Sun 0001 |
COLING | 2 |
| 2025 | Aligning Retrieval with Reader Needs: Reader-Centered Passage Selection for Open-Domain Question AnsweringabstractOpen-Domain Question Answering (ODQA) systems often struggle with the quality of retrieved passages, which may contain conflicting information and be misaligned with the reader’s needs. Existing retrieval methods aim to gather relevant passages but often fail to prioritize consistent and useful information for the reader. In this paper, we introduce a novel Reader-Centered Passage Selection (R-CPS) method, which enhances the performance of the retrieve-then-read pipeline by re-ranking and clustering passages from the reader’s perspective. Our method re-ranks passages based on the reader’s prediction probability distribution and clusters passages according to the predicted answers, prioritizing more useful and relevant passages to the top and reducing inconsistent information. Experiments on ODQA datasets demonstrate the effectiveness of our approach in improving the quality of evidence passages under zero-shot settings. Chunlei Xin, Shuheng Zhou 0001, Xuanang Chen, Yaojie Lu 0001, Huijia Zhu, Weiqiang Wang 0002, Xianpei Han, Le Sun 0001 |
COLING | 4 |
| 2025 | ConsistentChat: Building Skeleton-Guided Consistent Multi-Turn Dialogues for Large Language Models from ScratchabstractJiawei Chen, Xinyan Guan, Qianhao Yuan, Mo Guozhao, Weixiang Zhou, Yaojie Lu, Hongyu Lin, Ben He, Le Sun, Xianpei Han. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jiawei Chen 0011, Xinyan Guan, Qianhao Yuan, Guozhao Mo, Weixiang Zhou, Yaojie Lu 0001, Ben He 0001, Le Sun 0001, Xianpei Han |
EMNLP | 6 |
| 2025 | Teach Small Models to Reason by Curriculum DistillationabstractLarge Reasoning Models (LRMs) show strong System-2-style reasoning, but at the cost of significant computational overhead.In contrast, efficient System-1-style Large Language Models (LLMs) often struggle on complex tasks.We identify a critical asymmetry between these two paradigms: LRMs can implicitly self-distill their own reasoning, solving hard problems with near System-1-style efficiency while retaining superior performance.LLMs, however, lack such deep internal modes and collapse when forced to rely on their own reasoning rather than imitating external traces.This asymmetry explains why direct distillation from strong LRMs to weaker LLMs often fails: student models struggle to learn from LRMs' overly complex explicit reasoning and gain little from their overly compact implicit solutions.To address this, we introduce a two-stage curriculum distillation framework, which first builds a robust internal problem-solving student model and then teaches the student model to externalize this latent knowledge as explicit reasoning.On challenging mathematical benchmarks, our method significantly outperforms single-stage baselines, creating compact models with strong reasoning ability. Wangyi Jiang, Yaojie Lu 0001, Xianpei Han, Le Sun 0001 |
EMNLP | 2 |
| 2025 | PPTAgent: Generating and Evaluating Presentations Beyond Text-to-SlidesabstractHao Zheng, Xinyan Guan, Hao Kong, Wenkai Zhang, Jia Zheng, Weixiang Zhou, Hongyu Lin, Yaojie Lu, Xianpei Han, Le Sun. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Xinyan Guan, Jia Zheng 0009, Weixiang Zhou, Yaojie Lu 0001, Xianpei Han, Le Sun 0001 |
EMNLP | 8 |
| 2025 | ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective LayersabstractMultimodal Large Language Models (MLLMs) suffer from high computational costs due to their massive size and the large number of visual tokens. In this paper, we investigate layer-wise redundancy in MLLMs by introducing a novel metric, Layer Contribution (LC), which quantifies the impact of a layer's transformations on visual and text tokens, respectively. The calculation of LC involves measuring the divergence in model output that results from removing the layer's transformations on the specified tokens. Our pilot experiment reveals that many layers of MLLMs exhibit minimal contribution during the processing of visual tokens. Motivated by this observation, we propose ShortV, a training-free method that leverages LC to identify ineffective layers, and freezes visual token updates in these layers. Experiments show that ShortV can freeze visual token in approximately 60\% of the MLLM layers, thereby dramatically reducing computational costs related to updating visual tokens. For example, it achieves a 50\% reduction in FLOPs on LLaVA-NeXT-13B while maintaining superior performance. The code will be publicly available at https://github.com/icip-cas/ShortV Qianhao Yuan, Yanjiang Liu, Jiawei Chen 0011, Yaojie Lu 0001, Jia Zheng 0009, Xianpei Han, Le Sun 0001 |
ICCV | 5 |
| 2025 | The Rise and Down of Babel Tower: Investigating the Evolution Process of Multilingual Code Large Language ModelabstractLarge language models (LLMs) have shown significant multilingual capabilities. However, the mechanisms underlying the development of these capabilities during pre-training are not well understood. In this paper, we use code LLMs as an experimental platform to explore the evolution of multilingual capabilities in LLMs during the pre-training process. Based on our observations, we propose the Babel Tower Hypothesis, which describes the entire process of LLMs acquiring new language capabilities. During the learning process, multiple languages initially share a single knowledge system dominated by the primary language and gradually develop language-specific knowledge systems. We then validate the above hypothesis by tracking the internal states of the LLM using specific methods. Experimental results show that the internal state changes of the LLM are consistent with our Babel Tower Hypothesis. Building on these insights, we propose a novel method to construct an optimized pre-training corpus for multilingual code LLMs, which significantly outperforms LLMs trained on the original corpus. The proposed Babel Tower Hypothesis provides new insights into designing pre-training data distributions to achieve optimal multilingual capabilities in LLMs. Jiawei Chen 0011, Mengjie Ren, Yaojie Lu 0001, Xianpei Han, Le Sun 0001 |
ICLR | 7 |
| 2025 | StructRAG: Boosting Knowledge Intensive Reasoning of LLMs via Inference-time Hybrid Information StructurizationabstractRetrieval-augmented generation (RAG) is a key means to effectively enhance large language models (LLMs) in many knowledge-based tasks.
However, existing RAG methods struggle with knowledge-intensive reasoning tasks, because useful information required to these tasks are badly scattered.
This characteristic makes it difficult for existing RAG methods to accurately identify key information and perform global reasoning with such noisy augmentation.
In this paper, motivated by the cognitive theories that humans convert raw information into various structured knowledge when tackling knowledge-intensive reasoning, we proposes a new framework, StructRAG, which can identify the optimal structure type for the task at hand, reconstruct original documents into this structured format, and infer answers based on the resulting structure.
Extensive experiments across various knowledge-intensive tasks show that StructRAG achieves state-of-the-art performance, particularly excelling in challenging scenarios, demonstrating its potential as an effective solution for enhancing LLMs in complex real-world applications. Xuanang Chen, Haiyang Yu 0003, Yaojie Lu 0001, Qiaoyu Tang, Fei Huang 0002, Xianpei Han, Le Sun 0001, Yongbin Li 0001 |
ICLR | 5 |
| 2025 | Rethinking Reward Model Evaluation: Are We Barking up the Wrong Tree?abstractReward Models (RMs) are crucial for aligning language models with human preferences.
Currently, the evaluation of RMs depends on measuring accuracy against a validation set of manually annotated preference data.
Although this method is straightforward and widely adopted, the relationship between RM accuracy and downstream policy performance remains under-explored.
In this work, we conduct experiments in a synthetic setting to investigate how differences in RM measured by accuracy translate into gaps in optimized policy performance.
Our findings reveal that while there is a weak positive correlation between accuracy and downstream performance, policies optimized towards RMs with similar accuracy can exhibit quite different performance.
Moreover, we discover that the way of measuring accuracy significantly impacts its ability to predict the final policy performance.
Through the lens of the Regressional Goodhart effect, we recognize that accuracy, when used for measuring RM quality, can fail to fully capture the potential RM overoptimization.
This underscores the inadequacy of relying solely on accuracy to reflect their impact on policy optimization. Xueru Wen, Jie Lou, Yaojie Lu 0001, XingYu, Ben He 0001, Xianpei Han, Debing Zhang, Le Sun 0001 |
ICLR | 3 |
| 2025 | The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward ModelsabstractMultimodal Reward Models (MM-RMs) are crucial for aligning Large Language Models (LLMs) with human preferences, particularly as LLMs increasingly interact with multimodal data. However, we find that MM-RMs trained on existing datasets often struggle to generalize to out-of-distribution data due to their reliance on unimodal spurious correlations, primarily text-only shortcuts within the training distribution, which prevents them from leveraging true multimodal reward functions. To address this, we introduce a Shortcut-aware MM-RM learning algorithm that mitigates this issue by dynamically reweighting training samples, shifting the distribution toward better multimodal understanding, and reducing dependence on unimodal spurious correlations. Our experiments demonstrate significant improvements in generalization, downstream task performance, and scalability, establishing a more robust framework for multimodal reward modeling. Our source code is provided on https://github.com/alignrm/Generalizable-MM-RM. Xueru Wen, Jie Lou, Yuqiu Ji, Yaojie Lu 0001, Xianpei Han, Debing Zhang, Le Sun 0001 |
ICML | 5 |
| 2025 | Transferable Post-training via Inverse Value LearningabstractXinyu Lu, Xueru Wen, Yaojie Lu, Bowen Yu, Hongyu Lin, Haiyang Yu, Le Sun, Xianpei Han, Yongbin Li. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Xueru Wen, Yaojie Lu 0001, Bowen Yu 0002, Haiyang Yu 0003, Le Sun 0001, Xianpei Han, Yongbin Li 0001 |
NAACL (Long Papers) | 3 |
| 2025 | Influence of External Information on Large Language Models Mirrors Social Cognitive PatternsabstractSocial cognitive theory explains how people learn and acquire knowledge through observing others. Recent years have witnessed the rapid development of large language models (LLMs), which suggests their potential significance as agents in the society. LLMs, as AI agents, can observe external information, which shapes their cognition and behaviors. However, the extent to which external information influences LLMs’ cognition and behaviors remains unclear. This study investigates how external statements and opinions influence LLMs’ thoughts and behaviors from a social cognitive perspective. Three experiments were conducted to explore the effects of external information on LLMs’ memories, opinions, and social media behavioral decisions. Sociocognitive factors, including source authority, social identity, and social role, were analyzed to investigate their moderating effects. Results showed that external information can significantly shape LLMs’ memories, opinions, and behaviors, with these changes mirroring human social cognitive patterns such as authority bias, in-group bias, emotional positivity, and emotion contagion. This underscores the challenges in developing safe and unbiased LLMs, and emphasizes the importance of understanding the susceptibility of LLMs to external influences. Ning Bian, Yaojie Lu 0001, Chunkang Zhang, Ben He 0001, Xianpei Han, Le Sun 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | Mitigating Large Language Model Hallucinations via Autonomous Knowledge Graph-Based RetrofittingabstractIncorporating factual knowledge in knowledge graph is regarded as a promising approach for mitigating the hallucination of large language models (LLMs). Existing methods usually only use the user's input to query the knowledge graph, thus failing to address the factual hallucination generated by LLMs during its reasoning process. To address this problem, this paper proposes Knowledge Graph-based Retrofitting (KGR), a new framework that incorporates LLMs with KGs to mitigate factual hallucination during the reasoning process by retrofitting the initial draft responses of LLMs based on the factual knowledge stored in KGs. Specifically, KGR leverages LLMs to extract, select, validate, and retrofit factual statements within the model-generated responses, which enables an autonomous knowledge verifying and refining procedure without any additional manual efforts. Experiments show that KGR can significantly improve the performance of LLMs on factual QA benchmarks especially when involving complex reasoning processes, which demonstrates the necessity and effectiveness of KGR in mitigating hallucination and enhancing the reliability of LLMs. Xinyan Guan, Yanjiang Liu, Yaojie Lu 0001, Ben He 0001, Xianpei Han, Le Sun 0001 |
AAAI | 4 |
| 2024 | Rule or Story, Which is a Better Commonsense Expression for Talking with Large Language Models?abstractBuilding machines with commonsense has been a longstanding challenge in NLP due to the reporting bias of commonsense rules and the exposure bias of rule-based commonsense reasoning.In contrast, humans convey and pass down commonsense implicitly through stories.This paper investigates the inherent commonsense ability of large language models (LLMs) expressed through storytelling.We systematically investigate and compare stories and rules for retrieving and leveraging commonsense in LLMs.Experimental results on 28 commonsense QA datasets show that stories outperform rules as the expression for retrieving commonsense from LLMs, exhibiting higher generation confidence and commonsense accuracy.Moreover, stories are the more effective commonsense expression for answering questions regarding daily events, while rules are more effective for scientific questions.This aligns with the reporting bias of commonsense in text corpora.We further show that the correctness and relevance of commonsense stories can be further improved via iterative self-supervised fine-tuning.These findings emphasize the importance of using appropriate language to express, retrieve, and leverage commonsense for LLMs, highlighting a promising direction for better exploiting their commonsense abilities. Ning Bian, Xianpei Han, Yaojie Lu 0001, Ben He 0001, Le Sun 0001 |
ACL (1) | 4 |
| 2024 | Open Grounded Planning: Challenges and Benchmark ConstructionabstractThe emergence of large language models (LLMs) has increasingly drawn attention to the use of LLMs for human-like planning.Existing work on LLM-based planning either focuses on leveraging the inherent language generation capabilities of LLMs to produce free-style plans or employs reinforcement learning approaches to learn decision-making for a limited set of actions within restricted environments.However, both approaches exhibit significant discrepancies between the open and executable requirements in real-world planning.In this paper, we propose a new planning task-open grounded planning.The primary objective of open grounded planning is to ask the model to generate an executable plan based on a variable action set, thereby ensuring the executability of the produced plan.To this end, we establish a benchmark for open grounded planning spanning a wide range of domains.Then we test current state-of-the-art LLMs along with five planning approaches, revealing that existing LLMs and methods still struggle to address the challenges posed by grounded planning in open domains.The outcomes of this paper define and establish a foundational dataset for open grounded planning, and shed light on the potential challenges and future directions of LLM-based planning.Our code and datasets are at https://github.com/Shiguang-Guo/ Shiguang Guo, Ziliang Deng, Yaojie Lu 0001, Xianpei Han, Le Sun 0001 |
ACL (1) | 4 |
| 2024 | Few-shot Named Entity Recognition via Superposition Concept DiscriminationabstractFew-shot NER aims to identify entities of target types with only limited number of illustrative instances. Unfortunately, few-shot NER is severely challenged by the intrinsic precise generalization problem, i.e., it is hard to accurately determine the desired target type due to the ambiguity stemming from information deficiency. In this paper, we propose Superposition Concept Discriminator (SuperCD), which resolves the above challenge via an active learning paradigm. Specifically, a concept extractor is first introduced to identify superposition concepts from illustrative instances, with each concept corresponding to a possible generalization boundary. Then a superposition instance retriever is applied to retrieve corresponding instances of these superposition concepts from large-scale text corpus. Finally, annotators are asked to annotate the retrieved instances and these annotated instances together with original illustrative instances are used to learn FS-NER models. To this end, we learn a universal concept extractor and superposition instance retriever using a large-scale openly available knowledge bases. Experiments show that SuperCD can effectively identify superposition concepts from illustrative instances, retrieve superposition instances from large-scale corpus, and significantly improve the few-shot NER performance with minimal additional efforts. Jiawei Chen 0011, Xianpei Han, Yaojie Lu 0001, Shanshan Jiang 0001, Bin Dong 0003, Le Sun 0001 |
LREC/COLING | 4 |
| 2024 | ChatGPT Is a Knowledgeable but Inexperienced Solver: An Investigation of Commonsense Problem in Large Language ModelsabstractLarge language models (LLMs) have made significant progress in NLP. However, their ability to memorize, represent, and leverage commonsense knowledge has been a well-known pain point. In this paper, we specifically focus on ChatGPT, a widely used and easily accessible LLM, and ask the following questions: (1) Can ChatGPT effectively answer commonsense questions? (2) Is ChatGPT aware of the underlying commonsense knowledge for answering a specific question? (3) Is ChatGPT knowledgeable in commonsense? (4) Can ChatGPT effectively leverage commonsense for answering questions? We conduct a series of experiments on 11 datasets to evaluate ChatGPT’s commonsense abilities, including answering commonsense questions, identifying necessary knowledge, generating knowledge descriptions, and using knowledge descriptions to answer questions again. Experimental results show that: (1) ChatGPT can achieve good QA accuracies in commonsense tasks, while still struggling with certain domains of datasets. (2) ChatGPT is knowledgeable, and can accurately generate most of the commonsense knowledge using knowledge prompts. (3) Despite its knowledge, ChatGPT is an inexperienced commonsense problem solver, which cannot precisely identify the needed commonsense for answering a specific question. These findings raise the need to explore improved mechanisms for effectively incorporating commonsense into LLMs like ChatGPT, such as better instruction following and commonsense guidance. Ning Bian, Xianpei Han, Le Sun 0001, Yaojie Lu 0001, Ben He 0001, Shanshan Jiang 0001, Bin Dong 0003 |
LREC/COLING | 5 |
| 2024 | Meta-Cognitive Analysis: Evaluating Declarative and Procedural Knowledge in Datasets and Large Language ModelsabstractDeclarative knowledge and procedural knowledge are two key parts in meta-cognitive theory, and these two hold significant importance in pre-training and inference of LLMs. However, a comprehensive analysis comparing these two types of knowledge is lacking, primarily due to challenges in definition, probing and quantitative assessment. In this paper, we explore from a new perspective by providing ground-truth knowledge for LLMs and evaluating the effective score. Through extensive experiments with widely-used datasets and models, we get conclusions: (1) In most tasks, benefits from declarative knowledge are greater than those from procedural knowledge. (2) Profits of procedural knowledge are larger than declarative knowledge only in reasoning tasks with simple logic. (3) As pre-training progresses and size increases, model ability to utilize both kinds of knowledge significantly improves, but in different speed. We do detailed analysis for the findings and this can provide primary guidance for evaluation and enhancement of large language models. Yaojie Lu 0001, Xianpei Han, Le Sun 0001 |
LREC/COLING | 3 |
| 2024 | Beyond Full Fine-tuning: Harnessing the Power of LoRA for Multi-Task Instruction TuningabstractLow-Rank Adaptation (LoRA) is a widespread parameter-efficient fine-tuning algorithm for large-scale language models. It has been commonly accepted that LoRA mostly achieves promising results in single-task, low-resource settings, and struggles to handle multi-task instruction tuning scenarios. In this paper, we conduct a systematic study of LoRA on diverse tasks and rich resources with different learning capacities, examining its performance on seen tasks during training and its cross-task generalization on unseen tasks. Our findings challenge the prevalent assumption that the limited learning capacity will inevitably result in performance decline. In fact, our study reveals that when configured with an appropriate rank, LoRA can achieve remarkable performance in high-resource and multi-task scenarios, even comparable to that achieved through full fine-tuning. It turns out that the constrained learning capacity encourages LoRA to prioritize conforming to instruction requirements rather than memorizing specialized features of particular tasks or instances. This study reveals the underlying connection between learning capacity and generalization capabilities for robust parameter-efficient fine-tuning, highlighting a promising direction for the broader application of LoRA across various tasks and settings. Chunlei Xin, Yaojie Lu 0001, Shuheng Zhou 0001, Huijia Zhu, Weiqiang Wang 0002, Xianpei Han, Le Sun 0001 |
LREC/COLING | 2 |
| 2024 | Executing Natural Language-Described Algorithms with Large Language Models: An InvestigationabstractExecuting computer programs described in natural language has long been a pursuit of computer science. With the advent of enhanced natural language understanding capabilities exhibited by large language models (LLMs), the path toward this goal has been illuminated. In this paper, we seek to examine the capacity of present-day LLMs to comprehend and execute algorithms outlined in natural language. We established an algorithm test set sourced from Introduction to Algorithm, a well-known textbook that contains many representative widely-used algorithms. To systematically assess LLMs’ code execution abilities, we selected 30 algorithms, generated 300 random-sampled instances in total, and evaluated whether popular LLMs can understand and execute these algorithms. Our findings reveal that LLMs, notably GPT-4, can effectively execute programs described in natural language, as long as no heavy numeric computation is involved. We believe our findings contribute to evaluating LLMs’ code execution abilities and would encourage further investigation and application for the computation power of LLMs. Qiming Zhu, Yaojie Lu 0001, Xianpei Han, Le Sun 0001 |
LREC/COLING | 4 |
| 2024 | Seg2Act: Global Context-aware Action Generation for Document Logical StructuringabstractDocument logical structuring aims to extract the underlying hierarchical structure of documents, which is crucial for document intelligence.Traditional approaches often fall short in handling the complexity and the variability of lengthy documents.To address these issues, we introduce SEG2ACT, an end-to-end, generation-based method for document logical structuring, revisiting logical structure extraction as an action generation task.Specifically, given the text segments of a document, SEG2ACT iteratively generates the action sequence via a global context-aware generative model, and simultaneously updates its global context and current logical structure based on the generated actions.Experiments on ChCa-tExt and HierDoc datasets demonstrate the superior performance of SEG2ACT in both supervised and transfer learning settings 1 . Shaojie He, Meng Liao, Xuanang Chen, Yaojie Lu 0001, Yanxiong Lu, Xianpei Han, Le Sun 0001 |
EMNLP | 5 |
| 2024 | Self-Retrieval: End-to-End Information Retrieval with One Large Language ModelabstractThe rise of large language models (LLMs) has significantly transformed both the construction and application of information retrieval (IR) systems.
However, current interactions between IR systems and LLMs remain limited, with LLMs merely serving as part of components within IR systems, and IR systems being constructed independently of LLMs. This separated architecture restricts knowledge sharing and deep collaboration between them.
In this paper, we introduce Self-Retrieval, a novel end-to-end LLM-driven information retrieval architecture.
Self-Retrieval unifies all essential IR functions within a single LLM, leveraging the inherent capabilities of LLMs throughout the IR process.
Specifically, Self-Retrieval internalizes the retrieval corpus through self-supervised learning, transforms the retrieval process into sequential passage generation, and performs relevance assessment for reranking.
Experimental results demonstrate that Self-Retrieval not only outperforms existing retrieval approaches by a significant margin, but also substantially enhances the performance of LLM-driven downstream applications like retrieval-augmented generation. Qiaoyu Tang, Jiawei Chen 0011, Bowen Yu 0002, Yaojie Lu 0001, Cheng Fu 0003, Haiyang Yu 0003, Fei Huang 0002, Ben He 0001, Xianpei Han, Le Sun 0001, Yongbin Li 0001 |
NeurIPS | 5 |
| 2023 | Universal Information Extraction as Unified Semantic MatchingabstractThe challenge of information extraction (IE) lies in the diversity of label schemas and the heterogeneity of structures. Traditional methods require task-specific model design and rely heavily on expensive supervision, making them difficult to generalize to new schemas. In this paper, we decouple IE into two basic abilities, structuring and conceptualizing, which are shared by different tasks and schemas. Based on this paradigm, we propose to universally model various IE tasks with Unified Semantic Matching (USM) framework, which introduces three unified token linking operations to model the abilities of structuring and conceptualizing. In this way, USM can jointly encode schema and input text, uniformly extract substructures in parallel, and controllably decode target structures on demand. Empirical evaluation on 4 IE tasks shows that the proposed method achieves state-of-the-art performance under the supervised experiments and shows strong generalization ability in zero/few-shot transfer settings. Jie Lou, Yaojie Lu 0001, Dai Dai, Xianpei Han, Le Sun 0001, Hua Wu 0003 |
AAAI | 2 |
| 2023 | Learning In-context Learning for Named Entity RecognitionabstractJiawei Chen, Yaojie Lu, Hongyu Lin, Jie Lou, Wei Jia, Dai Dai, Hua Wu, Boxi Cao, Xianpei Han, Le Sun. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jiawei Chen 0011, Yaojie Lu 0001, Jie Lou, Dai Dai, Hua Wu 0003, Boxi Cao, Xianpei Han, Le Sun 0001 |
ACL (1) | 2 |
| 2023 | Testing Coreference Resolution Systems without Labeled Test SetsabstractCoreference resolution (CR) is a task to resolve different expressions (e.g., named entities, pronouns) that refer to the same real-world en- tity/event. It is a core natural language processing (NLP) component that underlies and empowers major downstream NLP applications such as machine translation, chatbots, and question-answering. De- spite its broad impact, the problem of testing CR systems has rarely been studied. A major difficulty is the shortage of a labeled dataset for testing. While it is possible to feed arbitrary sentences as test inputs to a CR system, a test oracle that captures their expected test outputs (coreference relations) is hard to define automatically. To address the challenge, we propose Crest, an automated testing methodology for CR systems. Crest uses constituency and depen- dency relations to construct pairs of test inputs subject to the same coreference. These relations can be leveraged to define the meta- morphic relation for metamorphic testing. We compare Crest with five state-of-the-art test generation baselines on two popular CR systems, and apply them to generate tests from 1,000 sentences randomly sampled from CoNLL-2012, a popular dataset for corefer- ence resolution. Experimental results show that Crest outperforms baselines significantly. The issues reported by Crest are all true positives (i.e., 100% precision), compared with 63% to 75% achieved by the baselines. Jialun Cao, Yaojie Lu 0001, Ming Wen 0001, Shing-Chi Cheung |
ESEC/SIGSOFT FSE | 2 |
| 2022 | Procedural Text Understanding via Scene-Wise EvolutionabstractProcedural text understanding requires machines to reason about entity states within the dynamical narratives. Current procedural text understanding approaches are commonly entity-wise, which separately track each entity and independently predict different states of each entity. Such an entity-wise paradigm does not consider the interaction between entities and their states. In this paper, we propose a new scene-wise paradigm for procedural text understanding, which jointly tracks states of all entities in a scene-by-scene manner. Based on this paradigm, we propose Scene Graph Reasoner (SGR), which introduces a series of dynamically evolving scene graphs to jointly formulate the evolution of entities, states and their associations throughout the narrative. In this way, the deep interactions between all entities and states can be jointly captured and simultaneously derived from scene graphs. Experiments show that SGR not only achieves the new state-of-the-art performance but also significantly accelerates the speed of reasoning. Jialong Tang, Meng Liao, Yaojie Lu 0001, Xianpei Han, Le Sun 0001, Weijian Xie, Jin Xu 0014 |
AAAI | 4 |
| 2022 | Unified Structure Generation for Universal Information ExtractionabstractYaojie Lu, Qing Liu, Dai Dai, Xinyan Xiao, Hongyu Lin, Xianpei Han, Le Sun, Hua Wu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yaojie Lu 0001, Dai Dai, Xinyan Xiao, Xianpei Han, Le Sun 0001, Hua Wu 0003 |
ACL (1) | 1 |
| 2022 | End-to-end neural event coreference resolution
Yaojie Lu 0001, Jialong Tang, Xianpei Han, Le Sun 0001 |
Artif. Intell. | 1 |
| 2021 | Text2Event: Controllable Sequence-to-Structure Generation for End-to-end Event ExtractionabstractYaojie Lu, Hongyu Lin, Jin Xu, Xianpei Han, Jialong Tang, Annan Li, Le Sun, Meng Liao, Shaoyi Chen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yaojie Lu 0001, Jin Xu 0014, Xianpei Han, Jialong Tang, Annan Li, Le Sun 0001, Meng Liao, Shaoyi Chen |
ACL/IJCNLP (1) | 1 |
| 2021 | From Discourse to Narrative: Knowledge Projection for Event Relation ExtractionabstractJialong Tang, Hongyu Lin, Meng Liao, Yaojie Lu, Xianpei Han, Le Sun, Weijian Xie, Jin Xu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jialong Tang, Meng Liao, Yaojie Lu 0001, Xianpei Han, Le Sun 0001, Weijian Xie, Jin Xu 0014 |
ACL/IJCNLP (1) | 4 |
| 2020 | A Rigorous Study on Named Entity Recognition: Can Fine-tuning Pretrained Model Lead to the Promised Land?abstractFine-tuning pretrained model has achieved promising performance on standard NER benchmarks.Generally, these benchmarks are blessed with strong name regularity, high mention coverage and sufficient context diversity.Unfortunately, when scaling NER to open situations, these advantages may no longer exist.And therefore it raises a critical question of whether previous creditable approaches can still work well when facing these challenges.As there is no currently available dataset to investigate this problem, this paper proposes to conduct randomization test on standard benchmarks.Specifically, we erase name regularity, mention coverage and context diversity respectively from the benchmarks, in order to explore their impact on the generalization ability of models.To further verify our conclusions, we also construct a new open NER dataset that focuses on entity types with weaker name regularity and lower mention coverage to verify our conclusion.From both randomization test and empirical experiments, we draw the conclusions that 1) name regularity is critical for the models to generalize to unseen mentions; 2) high mention coverage may undermine the model generalization ability and 3) context patterns may not require enormous data to capture when using pretrained encoders. Yaojie Lu 0001, Jialong Tang, Xianpei Han, Le Sun 0001, Zhicheng Wei, Nicholas Jing Yuan |
EMNLP (1) | 2 |
| 2020 | Semantically Smooth Bilingual Phrase Embeddings Based on Recursive Autoencoders
Xiangwen Zhang, Yaojie Lu 0001, Jinsong Su |
Neural Process. Lett. | 5 |
| 2019 | Sequence-to-Nuggets: Nested Entity Mention Detection via Anchor-Region NetworksabstractSequential labeling-based NER approaches restrict each word belonging to at most one entity mention, which will face a serious problem when recognizing nested entity mentions.In this paper, we propose to resolve this problem by modeling and leveraging the head-driven phrase structures of entity mentions, i.e., although a mention can nest other mentions, they will not share the same head word.Specifically, we propose Anchor-Region Networks (ARNs), a sequence-to-nuggets architecture for nested mention detection.ARNs first identify anchor words (i.e., possible head words) of all mentions, and then recognize the mention boundaries for each anchor word by exploiting regular phrase structures.Furthermore, we also design Bag Loss, an objective function which can train ARNs in an end-toend manner without using any anchor word annotation.Experiments show that ARNs achieve the state-of-the-art performance on three standard nested entity mention detection benchmarks. Yaojie Lu 0001, Xianpei Han, Le Sun 0001 |
ACL (1) | 2 |
| 2019 | Cost-sensitive Regularization for Label Confusion-aware Event DetectionabstractIn supervised event detection, most of the mislabeling occurs between a small number of confusing type pairs, including trigger-NIL pairs and sibling sub-types of the same coarse type.To address this label confusion problem, this paper proposes cost-sensitive regularization, which can force the training procedure to concentrate more on optimizing confusing type pairs.Specifically, we introduce a costweighted term into the training loss, which penalizes more on mislabeling between confusing label pairs.Furthermore, we also propose two estimators which can effectively measure such label confusion based on instance-level or population-level statistics.Experiments on TAC-KBP 2017 datasets demonstrate that the proposed method can significantly improve the performances of different models in both English and Chinese event detection. Yaojie Lu 0001, Xianpei Han, Le Sun 0001 |
ACL (1) | 2 |
| 2019 | Distilling Discrimination and Generalization Knowledge for Event Detection via Delta-Representation LearningabstractEvent detection systems rely on discrimination knowledge to distinguish ambiguous trigger words and generalization knowledge to detect unseen/sparse trigger words.Current neural event detection approaches focus on trigger-centric representations, which work well on distilling discrimination knowledge, but poorly on learning generalization knowledge.To address this problem, this paper proposes a ∆-learning approach to distill discrimination and generalization knowledge by effectively decoupling, incrementally learning and adaptively fusing event representation.Experiments show that our method significantly outperforms previous approaches on unseen/sparse trigger words, and achieves state-of-the-art performance on both ACE2005 and KBP2017 datasets. Yaojie Lu 0001, Xianpei Han, Le Sun 0001 |
ACL (1) | 1 |
| 2019 | Gazetteer-Enhanced Attentive Neural Networks for Named Entity RecognitionabstractHongyu Lin, Yaojie Lu, Xianpei Han, Le Sun, Bin Dong, Shanshan Jiang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yaojie Lu 0001, Xianpei Han, Le Sun 0001, Bin Dong 0003, Shanshan Jiang 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Iterative Dual Domain Adaptation for Neural Machine TranslationabstractJiali Zeng, Yang Liu, Jinsong Su, Yubing Ge, Yaojie Lu, Yongjing Yin, Jiebo Luo. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jiali Zeng, Yang Liu 0005, Jinsong Su, Yubin Ge, Yaojie Lu 0001, Yongjing Yin, Jiebo Luo 0001 |
EMNLP/IJCNLP (1) | 5 |
| 2018 | Variational Recurrent Neural Machine TranslationabstractPartially inspired by successful applications of variational recurrent neural networks, we propose a novel variational recurrent neural machine translation (VRNMT) model in this paper. Different from the variational NMT, VRNMT introduces a series of latent random variables to model the translation procedure of a sentence in a generative way, instead of a single latent variable. Specifically, the latent random variables are included into the hidden states of the NMT decoder with elements from the variational autoencoder. In this way, these variables are recurrently generated, which enables them to further capture strong and complex dependencies among the output translations at different timesteps. In order to deal with the challenges in performing efficient posterior inference and large-scale training during the incorporation of latent variables, we build a neural posterior approximator, and equip it with a reparameterization technique to estimate the variational lower bound. Experiments on Chinese-English and English-German translation tasks demonstrate that the proposed model achieves significant improvements over both the conventional and variational NMT models. Jinsong Su, Deyi Xiong, Yaojie Lu 0001, Xianpei Han, Biao Zhang 0002 |
AAAI | 4 |
| 2018 | Adaptive Scaling for Sparse Detection in Information ExtractionabstractThis paper focuses on detection tasks in information extraction, where positive instances are sparsely distributed and models are usually evaluated using F-measure on positive classes.These characteristics often result in deficient performance of neural network based detection models.In this paper, we propose adaptive scaling, an algorithm which can handle the positive sparsity problem and directly optimize over F-measure via dynamic costsensitive learning.To this end, we borrow the idea of marginal utility from economics and propose a theoretical framework for instance importance measuring without introducing any additional hyperparameters.Experiments show that our algorithm leads to a more effective and stable training of neural network based detection models. Yaojie Lu 0001, Xianpei Han, Le Sun 0001 |
ACL (1) | 2 |
| 2018 | Nugget Proposal Networks for Chinese Event DetectionabstractNeural network based models commonly regard event detection as a word-wise classification task, which suffer from the mismatch problem between words and event triggers, especially in languages without natural word delimiters such as Chinese.In this paper, we propose Nugget Proposal Networks (NPNs), which can solve the word-trigger mismatch problem by directly proposing entire trigger nuggets centered at each character regardless of word boundaries.Specifically, NPNs perform event detection in a character-wise paradigm, where a hybrid representation for each character is first learned to capture both structural and semantic information from both characters and words.Then based on learned representations, trigger nuggets are proposed and categorized by exploiting character compositional structures of Chinese event triggers.Experiments on both ACE2005 and TAC KBP 2017 datasets show that NPNs significantly outperform the state-of-the-art methods. Yaojie Lu 0001, Xianpei Han, Le Sun 0001 |
ACL (1) | 2 |
| 2018 | Cross-lingual implicit discourse relation recognition with co-trainingabstractA lack of labeled corpora obstructs the research progress on implicit discourse relation recognition (DRR) for Chinese, while there are some available discourse corpora in other languages, such as English. In this paper, we propose a cross-lingual implicit DRR framework that exploits an available English corpus for the Chinese DRR task. We use machine translation to generate Chinese instances from a labeled English discourse corpus. In this way, each instance has two independent views: Chinese and English views. Then we train two classifiers in Chinese and English in a co-training way, which exploits unlabeled Chinese data to implement better implicit DRR for Chinese. Experimental results demonstrate the effectiveness of our method. Yaojie Lu 0001, Mu Xu, Changxing Wu, Deyi Xiong, Jinsong Su |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2018 | Exploring Implicit Semantic Constraints for Bilingual Word Embeddings
Jinsong Su, Zhenqiao Song, Yaojie Lu 0001, Mu Xu, Changxing Wu, Yidong Chen 0001 |
Neural Process. Lett. | 3 |
| 2015 | Shallow Convolutional Neural Network for Implicit Discourse Relation RecognitionabstractImplicit discourse relation recognition remains a serious challenge due to the absence of discourse connectives.In this paper, we propose a Shallow Convolutional Neural Network (SCNN) for implicit discourse relation recognition, which contains only one hidden layer but is effective in relation recognition.The shallow structure alleviates the overfitting problem, while the convolution and nonlinear operations help preserve the recognition and generalization ability of our model.Experiments on the benchmark data set show that our model achieves comparable and even better performance when comparing against current state-of-the-art systems. Biao Zhang 0002, Jinsong Su, Deyi Xiong, Yaojie Lu 0001, Hong Duan, Junfeng Yao |
EMNLP | 4 |