VLDB 2026 Research / reviewers in the wild / expert
Jiangjie Chen
dblp:236/6076
· DBLP profile ↗
35ranked-venue papers
6as first author
33since 2021 · last 2026
0000-0002-2613-0386ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 6 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can LLMs Learn to Map the World from Local Descriptions?abstractSirui Xia, Aili Chen, Xintao Wang, Tinghui Zhu, Yikai Zhang, Jiangjie Chen, Yanghua Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Sirui Xia, Aili Chen, Xintao Wang 0001, Tinghui Zhu, Yikai Zhang 0004, Jiangjie Chen, Yanghua Xiao |
ACL (1) | 6 |
| 2025 | DEEPER Insight into Your User: Directed Persona Refinement for Dynamic Persona ModelingabstractAili Chen, Chengyu Du, Jiangjie Chen, Jinghan Xu, Yikai Zhang, Siyu Yuan, Zulong Chen, Liangyue Li, Yanghua Xiao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Aili Chen, Chengyu Du, Jiangjie Chen, Yikai Zhang 0004, Zulong Chen, Liangyue Li, Yanghua Xiao |
ACL (1) | 3 |
| 2025 | Past Meets Present: Creating Historical Analogy with Large Language ModelsabstractNianqi Li, Siyu Yuan, Jiangjie Chen, Jiaqing Liang, Feng Wei, Zujie Liang, Deqing Yang, Yanghua Xiao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Nianqi Li, Jiangjie Chen, Jiaqing Liang, Zujie Liang, Deqing Yang, Yanghua Xiao |
ACL (1) | 3 |
| 2025 | CoSER: Coordinating LLM-Based Persona Simulation of Established RolesabstractRole-playing language agents (RPLAs) have emerged as promising applications of large language models (LLMs). However, simulating established characters presents a challenging task for RPLAs, due to the lack of authentic character datasets and nuanced evaluation methods using such data. In this paper, we present CoSER, a collection of a high-quality dataset, open models, and an evaluation protocol towards effective RPLAs of established characters. The CoSER dataset covers 17,966 characters from 771 renowned books. It provides authentic dialogues with real-world intricacies, as well as diverse data types such as character experiences and internal thoughts. Drawing from acting methodology, we introduce given-circumstance acting for training and evaluating role-playing LLMs, where LLMs sequentially portray multiple characters in book scenes. Using our dataset, we develop CoSER 8B and CoSER 70B, i.e., advanced open role-playing LLMs built on LLaMA-3.1 models. Extensive experiments demonstrate the value of the CoSER dataset for RPLA training, evaluation and retrieval. Moreover, CoSER 70B exhibits state-of-the-art performance surpassing or matching GPT-4o on our evaluation and three existing benchmarks, i.e., achieving 75.80% and 93.47% accuracy on the InCharacter and LifeChoice benchmarks respectively. Our code, dataset and models are available at: https://github.com/Neph0s/CoSER. Xintao Wang 0001, Xinfeng Yuan, Rui Xu 0026, Jen-tse Huang 0001, Haoran Guo, Jiangjie Chen, Shuchang Zhou 0003, Wei Wang 0009, Yanghua Xiao |
ICML | 9 |
| 2025 | Revealing the Barriers of Language Agents in PlanningabstractJian Xie, Kexun Zhang, Jiangjie Chen, Siyu Yuan, Kai Zhang, Yikai Zhang, Lei Li, Yanghua Xiao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kexun Zhang, Jiangjie Chen, Kai Zhang 0033, Yikai Zhang 0004, Lei Li 0005, Yanghua Xiao |
NAACL (Long Papers) | 3 |
| 2025 | SELFGOAL: Your Language Agents Already Know How to Achieve High-level GoalsabstractRuihan Yang, Jiangjie Chen, Yikai Zhang, Siyu Yuan, Aili Chen, Kyle Richardson, Yanghua Xiao, Deqing Yang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Ruihan Yang, Jiangjie Chen, Yikai Zhang 0004, Aili Chen, Kyle Richardson 0001, Yanghua Xiao, Deqing Yang |
NAACL (Long Papers) | 2 |
| 2025 | EvoAgent: Towards Automatic Multi-Agent Generation via Evolutionary AlgorithmsabstractSiyu Yuan, Kaitao Song, Jiangjie Chen, Xu Tan, Dongsheng Li, Deqing Yang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kaitao Song, Jiangjie Chen, Xu Tan 0003, Dongsheng Li 0002, Deqing Yang |
NAACL (Long Papers) | 3 |
| 2025 | EASYTOOL: Enhancing LLM-based Agents with Concise Tool InstructionabstractSiyu Yuan, Kaitao Song, Jiangjie Chen, Xu Tan, Yongliang Shen, Kan Ren, Dongsheng Li, Deqing Yang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kaitao Song, Jiangjie Chen, Xu Tan 0003, Yongliang Shen 0001, Kan Ren, Dongsheng Li 0002, Deqing Yang |
NAACL (Long Papers) | 3 |
| 2025 | Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable PuzzlesabstractLarge Language Models (LLMs), such as OpenAI’s o1 and DeepSeek’s R1, excel at advanced reasoning tasks like math and coding via Reinforcement Learning with Verifiable Rewards (RLVR), but still struggle with puzzles solvable by humans without domain knowledge. We introduce ENIGMATA, the first comprehensive suite tailored for improving LLMs with puzzle reasoning skills. It includes 36 tasks across 7 categories, each with: 1) a generator that produces unlimited examples with controllable difficulty, and 2) a rule-based verifier for automatic evaluation. This generator-verifier design supports scalable, multi-task RL training, fine-grained analysis, and seamless RLVR integration. We further propose ENIGMATA-Eval, a rigorous benchmark, and develop optimized multi-task RLVR strategies. Our trained model, Qwen2.5-32B-ENIGMATA, consistently surpasses o3-mini-high and o1 on the puzzle reasoning benchmarks like ENIGMATA-Eval, ARC-AGI (32.8%), and ARC-AGI 2 (0.6%). It also generalizes well to out-of-domain puzzle benchmarks and mathematical reasoning, with little multi-tasking trade-off. When trained on larger models like Seed1.5-Thinking (20B activated parameters and 200B total parameters), puzzle data from ENIGMATA further boosts SoTA performance on advanced math and STEM reasoning tasks such as AIME (2024-2025), BeyondAIME and GPQA (Diamond), showing nice generalization benefits of ENIGMATA. This work offers a unified, controllable framework for advancing logical reasoning in LLMs. Project page: https://seed-enigmata.github.io. Jiangjie Chen, Qianyu He, Aili Chen, Zhicheng Cai, Weinan Dai, Hongli Yu, Jiaze Chen, Qiying Yu, Hao Zhou 0012, Mingxuan Wang |
NeurIPS | 1 |
| 2025 | KORGym: A Dynamic Game Platform for LLM Reasoning EvaluationabstractRecent advancements in large language models (LLMs) underscore the need for more comprehensive evaluation methods to accurately assess their reasoning capabilities. Existing benchmarks are often domain-specific and thus cannot fully capture an LLM’s general reasoning potential. To address this limitation, we introduce the **Knowledge Orthogonal Reasoning Gymnasium (KORGym)**, a dynamic evaluation platform inspired by KOR-Bench and Gymnasium. KORGym offers over fifty games in either textual or visual formats and supports interactive, multi-turn assessments with reinforcement learning scenarios. Using KORGym, we conduct extensive experiments on 19 LLMs and 8 VLMs, revealing consistent reasoning patterns within model families and demonstrating the superior performance of closed-source models. Further analysis examines the effects of modality, reasoning strategies, reinforcement learning techniques, and response length on model performance. We expect KORGym to become a valuable resource for advancing LLM reasoning research and developing evaluation methodologies suited to complex, interactive environments. Jiajun Shi, Jian Yang 0037, Xingyuan Bu, Jiangjie Chen, Junting Zhou, Kaijing Ma, Zhoufutu Wen, Bingli Wang, Yancheng He, Hualei Zhu, Wei Zhang 0021, Ruibin Yuan, Yunli Wang, Siyuan Fang, Qianyu He, Robert Tang, Yingshui Tan, Wangchunshu Zhou, Zhaoxiang Zhang 0001, Zhoujun Li 0001, Wenhao Huang 0001, Ge Zhang 0009 |
NeurIPS | 5 |
| 2025 | ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical ConstraintsabstractSpatial reasoning is a key capability in the field of artificial intelligence, especially crucial in areas such as robotics, computer vision, and natural language understanding. However, evaluating the ability of multimodal large language models (MLLMs) in complex spatial reasoning still faces challenges, particularly in scenarios requiring multi-step reasoning and precise mathematical constraints. This paper introduces ORIGAMISPACE, a new dataset and benchmark designed to evaluate the multi-step spatial reasoning ability and the capacity to handle mathematical constraints of MLLMs through origami tasks. The dataset contains 350 data instances, each comprising a strictly formatted crease pattern (CP diagram), the Compiled Flat Pattern, the complete Folding Process, and the final Folded Shape Image. We propose four evaluation tasks: Pattern Prediction, Multi-step Spatial Reasoning, Spatial Relationship Prediction, and End-to-End CP Code Generation. For the CP code generation task, we design an interactive environment and explore the possibility of using reinforcement learning methods to train MLLMs. Through experiments on existing MLLMs, we initially reveal the strengths and weaknesses of these models in handling complex spatial reasoning tasks. Rui Xu 0026, Dakuan Lu, Zicheng Zhao, Xiaoyu Tan, Xintao Wang 0001, Jiangjie Chen, Yinghui Xu 0001 |
NeurIPS | 7 |
| 2025 | ARIA: Training Language Agents with Intention-driven Reward AggregationabstractLarge language models (LLMs) have enabled agents to perform complex reasoning and decision-making through free-form language interactions. However, in open-ended language action environments (e.g., negotiation or question-asking games), the action space can be formulated as a joint distribution over tokens, resulting in an extremely large and combinatorial action space. Sampling actions in such a space can lead to extreme reward sparsity, which brings large reward variance, hindering effective reinforcement learning (RL). To address this, we propose **ARIA**, a method that **A**ggregates **R**ewards in **I**ntention space to enable efficient and effective language **A**gents training. ARIA aims to project natural language actions from the high-dimensional joint token distribution space into a low-dimensional intention space, where semantically similar actions are clustered and assigned shared rewards. This intention-aware reward aggregation reduces reward variance by densifying reward signals, fostering efficient and effective policy optimization. Extensive experiments demonstrate that ARIA not only significantly reduces gradient variance, but also delivers substantial performance gains of average 9.95% across four downstream tasks (e.g., negotiation and text-based games), consistently outperforming strong offline and online RL baselines. Ruihan Yang, Yikai Zhang 0004, Aili Chen, Xintao Wang 0001, Jiangjie Chen, Deqing Yang, Yanghua Xiao |
NeurIPS | 5 |
| 2025 | DAPO: An Open-Source LLM Reinforcement Learning System at ScaleabstractInference scaling empowers LLMs with unprecedented reasoning ability, with reinforcement learning as the core technique to elicit complex reasoning. However, key technical details of state-of-the-art reasoning LLMs are concealed (such as in OpenAI o1 blog and DeepSeek R1 technical report), thus the community still struggles to reproduce their RL training results. We propose the **D**ecoupled Clip and **D**ynamic s**A**mpling **P**olicy **O**ptimization (**DAPO**) algorithm, and fully open-source a state-of-the-art large-scale RL system that achieves 50 points on AIME 2024 using Qwen2.5-32B base model. Unlike previous works that withhold training details, we introduce four key techniques of our algorithm that make large-scale LLM RL a success. In addition, we open-source our training code, which is built on the verl framework, along with a carefully curated and processed dataset. These components of our open-source system enhance reproducibility and support future research in large-scale LLM RL. Qiying Yu, Zheng Zhang 0001, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Juncai Liu, Lingjun Liu, Xin Liu 0039, Haibin Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang 0022, Mofan Zhang, Ru Zhang 0006, Wang Zhang 0017, Jiaze Chen, Jiangjie Chen, Hongli Yu, Yuxuan Song 0002, Xiangpeng Wei, Hao Zhou 0012, Wei-Ying Ma, Ya-Qin Zhang, Mingxuan Wang |
NeurIPS | 25 |
| 2024 | Translate Meanings, Not Just Words: IdiomKB's Role in Optimizing Idiomatic Translation with Language ModelsabstractTo translate well, machine translation (MT) systems and general-purposed language models (LMs) need a deep understanding of both source and target languages and cultures. Therefore, idioms, with their non-compositional nature, pose particular challenges for Transformer-based systems, as literal translations often miss the intended meaning. Traditional methods, which replace idioms using existing knowledge bases (KBs), often lack scale and context-awareness. Addressing these challenges, our approach prioritizes context-awareness and scalability, allowing for offline storage of idioms in a manageable KB size. This ensures efficient serving with smaller models and provides a more comprehensive understanding of idiomatic expressions. We introduce a multilingual idiom KB (IdiomKB) developed using large LMs to address this. This KB facilitates better translation by smaller models, such as BLOOMZ (7.1B), Alpaca (7B), and InstructGPT (6.7B), by retrieving idioms' figurative meanings. We present a novel, GPT-4-powered metric for human-aligned evaluation, demonstrating that IdiomKB considerably boosts model performance. Human evaluations further validate our KB's quality. Jiangjie Chen, Hao Yang 0006, Shimin Tao, Yanghua Xiao |
AAAI | 2 |
| 2024 | GumbelSoft: Diversified Language Model Watermarking via the GumbelMax-trickabstractLarge language models (LLMs) excellently generate human-like text, but also raise concerns about misuse in fake news and academic dishonesty.Decoding-based watermark, particularly the GumbelMax-trick-based watermark (GM watermark), is a standout solution for safeguarding machine-generated texts due to its notable detectability.However, GM watermark encounters a major challenge with generation diversity, always yielding identical outputs for the same prompt, negatively impacting generation diversity and user experience.To overcome this limitation, we propose a new type of GM watermark, the Logits-Addition watermark, and its three variants, specifically designed to enhance diversity.Among these, the GumbelSoft watermark (a softmax variant of the Logits-Addition watermark) demonstrates superior performance in high diversity settings, with its AUROC score outperforming those of the two alternative variants by 0.1 to 0.3 and surpassing other decoding-based watermarking methods by a minimum of 0.1.1 Jiayi Fu, Xuandong Zhao, Ruihan Yang, Yuansen Zhang, Jiangjie Chen, Yanghua Xiao |
ACL (1) | 5 |
| 2024 | InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological InterviewsabstractXintao Wang, Yunze Xiao, Jen-tse Huang, Siyu Yuan, Rui Xu, Haoran Guo, Quan Tu, Yaying Fei, Ziang Leng, Wei Wang, Jiangjie Chen, Cheng Li, Yanghua Xiao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Xintao Wang 0001, Yunze Xiao, Jen-tse Huang 0001, Rui Xu 0026, Haoran Guo, Quan Tu, Yaying Fei, Ziang Leng, Wei Wang 0009, Jiangjie Chen, Yanghua Xiao |
ACL (1) | 11 |
| 2024 | ANALOGYKB: Unlocking Analogical Reasoning of Language Models with A Million-scale Knowledge BaseabstractAnalogical reasoning is a fundamental cognitive ability of humans.However, current language models (LMs) still struggle to achieve human-like performance in analogical reasoning tasks due to a lack of resources for model training.In this work, we address this gap by proposing ANALOGYKB, a million-scale analogy knowledge base (KB) derived from existing knowledge graphs (KGs).ANALOGYKB identifies two types of analogies from the KGs: 1) analogies of the same relations, which can be directly extracted from the KGs, and 2) analogies of analogous relations, which are identified with a selection and filtering pipeline enabled by large language models (LLMs), followed by minor human efforts for data quality control.Evaluations on a series of datasets of two analogical reasoning tasks (analogy recognition and generation) demonstrate that ANAL-OGYKB successfully enables both smaller LMs and LLMs to gain better analogical reasoning capabilities.Resources of this paper can be found at https://github.com/siyuyuan/ analogykb. Jiangjie Chen, Changzhi Sun, Jiaqing Liang, Yanghua Xiao, Deqing Yang |
ACL (1) | 2 |
| 2024 | TimeArena: Shaping Efficient Multitasking Language Agents in a Time-Aware SimulationabstractDespite remarkable advancements in emulating human-like behavior through Large Language Models (LLMs), current textual simulations do not adequately address the notion of time.To this end, we introduce TIMEARENA, a novel textual simulated environment that incorporates complex temporal dynamics and constraints that better reflect real-life planning scenarios.In TIMEARENA, agents are asked to complete multiple tasks as soon as possible, allowing for parallel processing to save time.We implement the dependency between actions, the time duration for each action, and the occupancy of the agent and the objects in the environment.TIMEARENA grounds to 30 real-world tasks in cooking, household activity, and laboratory work.We conduct extensive experiments with various LLMs using TIMEARENA.Our findings reveal that even the most powerful models, e.g., GPT-4, still lag behind humans in effective multitasking, underscoring the need for enhanced temporal awareness in the development of language agents. 1 Yikai Zhang 0004, Caiyu Hu, Kyle Richardson 0001, Yanghua Xiao, Jiangjie Chen |
ACL (1) | 6 |
| 2024 | SEGMENT+: Long Text Processing with Short-Context Language ModelsabstractWei Shi, Shuang Li, Kerun Yu, Jinglei Chen, Zujie Liang, Xinhui Wu, Yuxi Qian, Feng Wei, Bo Zheng, Jiaqing Liang, Jiangjie Chen, Yanghua Xiao. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Kerun Yu, Jinglei Chen, Zujie Liang, Xinhui Wu, Yuxi Qian, Jiaqing Liang, Jiangjie Chen, Yanghua Xiao |
EMNLP | 11 |
| 2024 | Evaluating Character Understanding of Large Language Models via Character Profiling from Fictional WorksabstractLarge language models (LLMs) have demonstrated impressive performance and spurred numerous AI applications, in which role-playing agents (RPAs) are particularly popular, especially for fictional characters.The prerequisite for these RPAs lies in the capability of LLMs to understand characters from fictional works.Previous efforts have evaluated this capability via basic classification tasks or characteristic imitation, failing to capture the nuanced character understanding with LLMs.In this paper, we propose evaluating LLMs' character understanding capability via the character profiling task, i.e., summarizing character profiles from corresponding materials, a widely adopted yet understudied practice for RPA development.Specifically, we construct the CROSS dataset from literature experts and assess the generated profiles by comparing them with ground truth references and evaluating their applicability in downstream tasks.Our experiments, which cover various summarization methods and LLMs, have yielded promising results.These results strongly validate the character understanding capability of LLMs.Resources are available at https://github. com/Joanna0123/character_profiling. Xinfeng Yuan, Yuhan Cui, Tianhe Lin, Xintao Wang 0001, Rui Xu 0026, Jiangjie Chen, Deqing Yang |
EMNLP | 7 |
| 2024 | Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge ConflictsabstractBy providing external information to large language models (LLMs), tool augmentation (including retrieval augmentation) has emerged as a promising solution for addressing the limitations of LLMs' static parametric memory.
However, how receptive are LLMs to such external evidence, especially when the evidence conflicts with their parametric memory?
We present the first comprehensive and controlled investigation into the behavior of LLMs when encountering knowledge conflicts.
We propose a systematic framework to elicit high-quality parametric memory from LLMs and construct the corresponding counter-memory, which enables us to conduct a series of controlled experiments.
Our investigation reveals seemingly contradicting behaviors of LLMs.
On the one hand, different from prior wisdom, we find that LLMs can be highly receptive to external evidence even when that conflicts with their parametric memory, given that the external evidence is coherent and convincing.
On the other hand, LLMs also demonstrate a strong confirmation bias when the external evidence contains some information that is consistent with their parametric memory, despite being presented with conflicting evidence at the same time.
These results pose important implications that are worth careful consideration for the further development and deployment of tool- and retrieval-augmented LLMs.
Resources are available at https://github.com/OSU-NLP-Group/LLM-Knowledge-Conflict. Kai Zhang 0033, Jiangjie Chen, Renze Lou, Yu Su 0001 |
ICLR | 3 |
| 2024 | TravelPlanner: A Benchmark for Real-World Planning with Language AgentsabstractPlanning has been part of the core pursuit for artificial intelligence since its conception, but earlier AI agents mostly focused on constrained settings because many of the cognitive substrates necessary for human-level planning have been lacking. Recently, language agents powered by large language models (LLMs) have shown interesting capabilities such as tool use and reasoning. Are these language agents capable of planning in more complex settings that are out of the reach of prior AI agents? To advance this investigation, we propose TravelPlanner, a new planning benchmark that focuses on travel planning, a common real-world planning scenario. It provides a rich sandbox environment, various tools for accessing nearly four million data records, and 1,225 meticulously curated planning intents and reference plans. Comprehensive evaluations show that the current language agents are not yet capable of handling such complex planning tasks—even GPT-4 only achieves a success rate of 0.6%. Language agents struggle to stay on task, use the right tools to collect information, or keep track of multiple constraints. However, we note that the mere possibility for language agents to tackle such a complex problem is in itself non-trivial progress. TravelPlanner provides a challenging yet meaningful testbed for future language agents. Kai Zhang 0033, Jiangjie Chen, Tinghui Zhu, Renze Lou, Yuandong Tian, Yanghua Xiao, Yu Su 0001 |
ICML | 3 |
| 2023 | Converge to the Truth: Factual Error Correction via Iterative Constrained EditingabstractGiven a possibly false claim sentence, how can we automatically correct it with minimal editing? Existing methods either require a large number of pairs of false and corrected claims for supervised training or do not handle well errors spanning over multiple tokens within an utterance. In this paper, we propose VENCE, a novel method for factual error correction (FEC) with minimal edits. VENCE formulates the FEC problem as iterative sampling editing actions with respect to a target density function. We carefully design the target function with predicted truthfulness scores from an offline trained fact verification model. VENCE samples the most probable editing positions based on back-calculated gradients of the truthfulness score concerning input tokens and the editing actions using a distantly-supervised language model (T5). Experiments on a public dataset show that VENCE improves the well-adopted SARI metric by 5.3 (or a relative improvement of 11.8%) over the previous best distantly-supervised methods. Jiangjie Chen, Rui Xu 0026, Wenxuan Zeng, Changzhi Sun, Lei Li 0005, Yanghua Xiao |
AAAI | 1 |
| 2023 | Unsupervised Explanation Generation via Correct InstantiationsabstractWhile large pre-trained language models (PLM) have shown their great skills at solving discriminative tasks, a significant gap remains when compared with humans for explanation-related tasks. Among them, explaining the reason why a statement is wrong (e.g., against commonsense) is incredibly challenging. The major difficulty is finding the conflict point, where the statement contradicts our real world. This paper proposes Neon, a two-phrase, unsupervised explanation generation framework. Neon first generates corrected instantiations of the statement (phase I), then uses them to prompt large PLMs to find the conflict point and complete the explanation (phase II). We conduct extensive experiments on two standard explanation benchmarks, i.e., ComVE and e-SNLI. According to both automatic and human evaluations, Neon outperforms baselines, even for those with human-annotated instantiations. In addition to explaining a negative prediction, we further demonstrate that Neon remains effective when generalizing to different scenarios. The resources of Neon are available at: https://github.com/Shark-NLP/Neon. Sijie Cheng, Zhiyong Wu 0003, Jiangjie Chen, Lingpeng Kong |
AAAI | 3 |
| 2023 | Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense KnowledgeabstractLarge language models (LLMs) have been widely studied for their ability to store and utilize positive knowledge.However, negative knowledge, such as "lions don't live in the ocean", is also ubiquitous in the world but rarely mentioned explicitly in the text.What do LLMs know about negative knowledge?This work examines the ability of LLMs to negative commonsense knowledge.We design a constrained keywords-to-sentence generation task (CG) and a Boolean question-answering task (QA) to probe LLMs.Our experiments reveal that LLMs frequently fail to generate valid sentences grounded in negative commonsense knowledge, yet they can correctly answer polar yes-or-no questions.We term this phenomenon the belief conflict of LLMs.Our further analysis shows that statistical shortcuts and negation reporting bias from language modeling pre-training cause this conflict.1 * Work done while at Brain Technologies, Inc. Jiangjie Chen, Ziquan Fu, Sijie Cheng, Lei Li 0005, Yanghua Xiao |
ACL (1) | 1 |
| 2023 | Distilling Script Knowledge from Large Language Models for Constrained Language PlanningabstractSiyu Yuan, Jiangjie Chen, Ziquan Fu, Xuyang Ge, Soham Shah, Charles Jankowski, Yanghua Xiao, Deqing Yang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jiangjie Chen, Ziquan Fu, Xuyang Ge, Soham Shah, Charles Robert Jankowski, Yanghua Xiao, Deqing Yang |
ACL (1) | 2 |
| 2022 | LOREN: Logic-Regularized Reasoning for Interpretable Fact VerificationabstractGiven a natural language statement, how to verify its veracity against a large-scale textual knowledge source like Wikipedia? Most existing neural models make predictions without giving clues about which part of a false claim goes wrong. In this paper, we propose LOREN, an approach for interpretable fact verification. We decompose the verification of the whole claim at phrase-level, where the veracity of the phrases serves as explanations and can be aggregated into the final verdict according to logical rules. The key insight of LOREN is to represent claim phrase veracity as three-valued latent variables, which are regularized by aggregation logical rules. The final claim verification is based on all latent variables. Thus, LOREN enjoys the additional benefit of interpretability --- it is easy to explain how it reaches certain results with claim phrase veracity. Experiments on a public fact verification benchmark show that LOREN is competitive against previous approaches while enjoying the merit of faithful and accurate interpretability. The resources of LOREN are available at: https://github.com/jiangjiechen/LOREN. Jiangjie Chen, Qiaoben Bao, Changzhi Sun, Xinbo Zhang, Jiaze Chen, Hao Zhou 0012, Yanghua Xiao, Lei Li 0005 |
AAAI | 1 |
| 2022 | Unsupervised Editing for Counterfactual StoriesabstractCreating what-if stories requires reasoning about prior statements and possible outcomes of the changed conditions. One can easily generate coherent endings under new conditions, but it would be challenging for current systems to do it with minimal changes to the original story. Therefore, one major challenge is the trade-off between generating a logical story and rewriting with minimal-edits. In this paper, we propose EDUCAT, an editing-based unsupervised approach for counterfactual story rewriting. EDUCAT includes a target position detection strategy based on estimating causal effects of the what-if conditions, which keeps the causal invariant parts of the story. EDUCAT then generates the stories under fluency, coherence and minimal-edits constraints. We also propose a new metric to alleviate the shortcomings of current automatic metrics and better evaluate the trade-off. We evaluate EDUCAT on a public counterfactual story rewriting benchmark. Experiments show that EDUCAT achieves the best trade-off over unsupervised SOTA methods according to both automatic and human evaluation. The resources of EDUCAT are available at: https://github.com/jiangjiechen/EDUCAT. Jiangjie Chen, Chun Gan, Sijie Cheng, Hao Zhou 0012, Yanghua Xiao, Lei Li 0005 |
AAAI | 1 |
| 2022 | FalCon: A Faithful Contrastive Framework for Response Generation in TableQA Systems
Shineng Fang, Jiangjie Chen, Xinyao Shen, Yunwen Chen, Yanghua Xiao |
DASFAA (3) | 2 |
| 2022 | Neighbors Are Not Strangers: Improving Non-Autoregressive Translation under Low-Frequency Lexical ConstraintsabstractChun Zeng, Jiangjie Chen, Tianyi Zhuang, Rui Xu, Hao Yang, Qin Ying, Shimin Tao, Yanghua Xiao. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Chun Zeng, Jiangjie Chen, Tianyi Zhuang, Rui Xu 0026, Hao Yang 0006, Shimin Tao, Yanghua Xiao |
NAACL-HLT | 2 |
| 2022 | Harvesting More Answer Spans from Paragraph beyond AnnotationabstractAutomaticA nswer spanE xtraction (AE) focuses on identifying key information from paragraphs that can be asked. It has been used to facilitate downstream question generation tasks or data augmentation for question answering. Current work of AE heavily relies on the annotated answer spans fromM achineR eadingC omprehension (MRC) datasets. However, these methods suffer from the partial annotation problem due to the annotation protocols of MRC tasks. To tackle this problem, we propose \mymethod, a S tructured Co ntext graph network with P ositive -unlabeled learning. \mymethod first represents the paragraph by constructing a graph with both syntactic and semantic edges, then adopts a unified pointer network for answer span identification. \mymethod narrows the discrenpency between AE and MRC by formulating AE as aP ositive-\textitu nlabeled (PU) learning problem, thus recovering more answer spans from paragraphs. To evaluate newly extracted spans without annotation, we also present an automatic metric from the perspective of question answering and text summarization, which correlates well with human judgments. Comprehensive experiments on both AE and downstream tasks demonstrate the effectiveness of our proposed framework. Our code is available at \urlhttps://github.com/iambabao/SCOPE. Qiaoben Bao, Jiangjie Chen, Linfang Liu, Jiaqing Liang, Yanghua Xiao |
WSDM | 2 |
| 2022 | Diversified Query Generation Guided by Knowledge GraphabstractRelevant articles recommendation plays an important role in online news platforms. Directly displaying recalled articles by a search engine lacks a deep understanding of the article contents. Generating clickable queries, on the other hand, summarizes an article in various aspects, which can be henceforth utilized to better connect relevant articles. Most existing approaches for generating article queries, however, do not consider the diversity of queries or whether they are appealing enough, which are essential for boosting user experience and platform drainage. To this end, we propose a Knowledge-Enhanced Diversified QuerY Generator (KEDY), which leverages an external knowledge graph (KG) as guidance. We diversify the query generation with the information of semantic neighbors of the entities in articles. We further constrain the diversification process with entity popularity knowledge to build appealing queries that users may be more interested in. The information within KG is propagated towards more popular entities with popularity-guided graph attention. We collect a news-query dataset from the search logs of a real-world search engine. Extensive experiments demonstrate our proposed KEDY can generate more diversified and insightful related queries than several strong baselines. Xinyao Shen, Jiangjie Chen, Jiaze Chen, Chun Zeng, Yanghua Xiao |
WSDM | 2 |
| 2021 | Diversified Paraphrase Generation with Commonsense Knowledge Graph
Xinyao Shen, Jiangjie Chen, Yanghua Xiao |
NLPCC (1) | 2 |
| 2019 | Ensuring Readability and Data-fidelity using Head-modifier Templates in Deep Type Description GenerationabstractA type description is a succinct noun compound which helps human and machines to quickly grasp the informative and distinctive information of an entity.Entities in most knowledge graphs (KGs) still lack such descriptions, thus calling for automatic methods to supplement such information.However, existing generative methods either overlook the grammatical structure or make factual mistakes in generated texts.To solve these problems, we propose a head-modifier template-based method to ensure the readability and data fidelity of generated type descriptions.We also propose a new dataset and two automatic metrics for this task.Experiments show that our method improves substantially compared with baselines and achieves stateof-the-art performance on both datasets. Jiangjie Chen, Haiyun Jiang, Suo Feng, Yanghua Xiao |
ACL (1) | 1 |
| 2019 | CN-Probase: A Data-Driven Approach for Large-Scale Chinese Taxonomy ConstructionabstractTaxonomies play an important role in machine intelligence. However, most well-known taxonomies are in English, and non-English taxonomies, especially Chinese ones, are still very rare. In this paper, we focus on automatic Chinese taxonomy construction and propose an effective generation and verification framework to build a large-scale and high-quality Chinese taxonomy. In the generation module, we extract isA relations from multiple sources of Chinese encyclopedia, which ensures the coverage. To further improve the precision of taxonomy, we apply three heuristic approaches in verification module. As a result, we construct the largest Chinese taxonomy with high precision about 95% called CN-Probase. Our taxonomy has been deployed on Aliyun, with over 82 million API calls in six months. Jindong Chen, Jiangjie Chen, Yanghua Xiao, Zhendong Chu, Jiaqing Liang, Wei Wang 0009 |
ICDE | 3 |