Pengjun Xie

dblp:212/1755 · DBLP profile ↗
← Back
63ranked-venue papers
0as first author
59since 2021 · last 2026
0009-0004-8412-359XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 57 · 53 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021
YearPublicationVenuePosition
2026 ERank: Fusing Supervised Fine-Tuning and Reinforcement Learning for Effective and Efficient Text Reranking
abstract
Text reranking models are a crucial component in modern systems like Retrieval-Augmented Generation, tasked with selecting the most relevant documents prior to generation. However, current Large Language Models (LLMs) powered rerankers often face a fundamental trade-off. On one hand, Supervised Fine-Tuning based pointwise methods that frame relevance as a binary classification task lack the necessary scoring discrimination, particularly for those built on reasoning LLMs. On the other hand, approaches designed for complex reasoning often employ powerful yet inefficient listwise formulations, rendering them impractical for low latency applications. To resolve this dilemma, we introduce ERank, a highly Effective and Efficient pointwise reranker built from a reasoning LLM that excels across diverse relevance scenarios. We propose a novel two-stage training pipeline that begins with Supervised Fine-Tuning (SFT). In this stage, we move beyond binary labels and train the model generatively to output fine grained integer scores, which significantly enhances relevance discrimination. The model is then further refined using Reinforcement Learning (RL) with a novel, listwise derived reward. This technique instills global ranking awareness into the efficient pointwise architecture. We evaluate the ERank reranker on the BRIGHT, FollowIR, TREC DL, and BEIR benchmarks, demonstrating superior effectiveness and robustness compared to existing approaches. On the reasoning-intensive BRIGHT benchmark, our ERank-4B achieves an nDCG@10 of 38.7, while a larger 32B variant reaches a state of the art nDCG@10 of 40.2.
Yuzheng Cai, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Weiguo Zheng
AAAI5
2026 Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning
abstract
While Reinforcement Learning (RL) has advanced LLM reasoning, applying it to longcontext scenarios is hindered by sparsity of outcome rewards.This limitation fails to penalize ungrounded "lucky guesses," leaving the critical process of needle-in-a-haystack evidence retrieval largely unsupervised.To address this, we propose EAPO (Evidence-Augmented Policy Optimization).We first establish the Evidence-Augmented Reasoning paradigm, validating via Tree-Structured Evidence Sampling that precise evidence extraction is the decisive bottleneck for long-context reasoning.Guided by this insight, EAPO introduces a specialized RL algorithm where a reward model computes a Group-Relative Evidence Reward, providing dense process supervision to explicitly improve evidence quality.To sustain accurate supervision throughout training, we further incorporate an Adaptive Reward-Policy Co-Evolution mechanism.This mechanism iteratively refines the reward model using outcome-consistent rollouts, sharpening its discriminative capability to ensure precise process guidance.Comprehensive evaluations across eight benchmarks demonstrate that EAPO significantly enhances long-context reasoning performance compared to SOTA baselines.
Shen Huang, Pengjun Xie, Jingren Zhou 0001, Jiuxin Cao
ACL (1)4
2026 Nested Browser-Use Learning for Agentic Information Seeking
abstract
Baixuan Li, Jialong Wu, Wenbiao Yin, Kuan Li, Zhongwang Zhang, Huifeng Yin, Zhengwei Tao, Liwen Zhang, Pengjun Xie, Jingren Zhou, Yong Jiang, Wentao Zhang, Zhiqiang Gao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Baixuan Li, Jialong Wu 0007, Wenbiao Yin, Kuan Li, Zhongwang Zhang, Huifeng Yin, Zhengwei Tao, Pengjun Xie, Jingren Zhou 0001, Yong Jiang 0005, Wentao Zhang 0001
ACL (1)9
2026 Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing
abstract
Tingyu Song, Yanzhao Zhang, Mingxin Li, Zhuoning Guo, Dingkun Long, Pengjun Xie, Siyue Zhang, Yilun Zhao, Shu Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tingyu Song, Yanzhao Zhang, Zhuoning Guo, Dingkun Long, Pengjun Xie, Siyue Zhang, Yilun Zhao 0001
ACL (1)6
2026 Internalizing Explicit Reasoning into Latent Space for Dense Retrieval
abstract
Large Language Models (LLMs) have fundamentally transformed dense retrieval, upgrading backbones from discriminative encoders to generative architectures. However, a critical disconnect remains: while LLMs possess strong reasoning capabilities, current retrievers predominantly utilize them as static encoders, leaving their potential for complex reasoning unexplored. To address this, existing approaches typically adopt ''rewrite-then-retrieve'' pipelines to generate explicit Chain-of-Thought (CoT) rationales before retrieval. However, this incurs prohibitive latency. Conversely, implicit reasoning methods utilizing latent tokens offer efficiency but often suffer from semantic degeneration due to the lack of explicit supervision. In this paper, we propose LaSER, a novel self-distillation framework that internalizes explicit reasoning into the latent space of dense retrievers. Operating on a shared LLM backbone, LaSER introduces a dual-view training mechanism: an Explicit view that explicitly encodes ground-truth reasoning paths, and a Latent view that performs implicit latent thinking. To bridge the gap between these views, we design a multi-grained alignment strategy. Beyond standard output alignment, we introduce a trajectory alignment mechanism that synchronizes the intermediate latent states of the latent path with the semantic progression of the explicit reasoning segments. This allows the retriever to ''think'' silently and effectively without autoregressive text generation. Extensive experiments on both in-domain and out-of-domain reasoning-intensive benchmarks demonstrate that LaSER significantly outperforms state-of-the-art baselines. Furthermore, analyses across diverse backbones and model scales validate the robustness of our approach, confirming that our unified learning framework is essential for eliciting effective latent thinking. Our method successfully combines the reasoning depth of explicit CoT pipelines with the inference efficiency of standard dense retrievers. The code, model, and training data are available at https://github.com/RUC-NLPIR/LaSER.
Jiajie Jin, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Yutao Zhu 0001, Zhicheng Dou
SIGIR5
2026 Progressive Adaptation of Large Language Models for Multilingual Text Ranking
abstract
Despite increasing research attention to text ranking, most studies focus on monolingual scenarios, with a particular emphasis on English-language contexts. This narrow focus limits the applicability of ranking models in cross-lingual contexts, such as ranking Chinese documents based on English queries. Recent advances in large language models (LLMs) have significantly reduced inter-language barriers through pre-training on extensive multilingual corpora, thus facilitating the study of multilingual text ranking (MTR). In this work, we explore the potential of LLMs in MTR tasks. Specifically, we first introduce an MTR benchmark encompassing both monolingual and cross-lingual scenarios. Then, we propose a two-stage training pipeline to alleviate the misalignment between LLMs and text ranking. Lastly, we adapt this training pipeline to multilingual scenarios from the perspective of training data and methods. Our experiments on the MTR benchmark demonstrate that the proposed multilingual two-stage training pipeline significantly improves LLM ranking performance in both monolingual and cross-lingual scenarios, particularly in out-domain settings. We complement these findings with a thorough analysis to deepen the understanding of our approach.
Longhui Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, Jing Li 0034, Min Zhang 0005
ACM Trans. Inf. Syst.4
2025 WebWalker: Benchmarking LLMs in Web Traversal
abstract
Retrieval-augmented generation (RAG) demonstrates remarkable performance across tasks in open-domain question-answering. However, traditional search engines may retrieve shallow content, limiting the ability of LLMs to handle complex, multi-layered information. To address this, we introduce WebWalkerQA, a benchmark designed to assess the ability of LLMs to perform web traversal. It evaluates the capacity of LLMs to traverse a website’s subpages to extract high-quality data systematically. We propose WebWalker, which is a multi-agent framework that mimics human-like web navigation through an explore-critic paradigm. Extensive experimental results show that WebWalkerQA is challenging and demonstrates the effectiveness of RAG combined with WebWalker, through this horizontal and vertical integration in real-world scenarios.
Jialong Wu 0007, Wenbiao Yin, Yong Jiang 0005, Zhenglin Wang, Zekun Xi, Runnan Fang, Linhai Zhang, Yulan He 0001, Pengjun Xie, Fei Huang 0002
ACL (1)10
2025 Agentic Knowledgeable Self-awareness
abstract
Shuofei Qiao, Zhisong Qiu, Baochang Ren, Xiaobin Wang, Xiangyuan Ru, Ningyu Zhang, Xiang Chen, Yong Jiang, Pengjun Xie, Fei Huang, Huajun Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shuofei Qiao, Zhisong Qiu, Baochang Ren, Xiaobin Wang, Xiangyuan Ru, Ningyu Zhang 0001, Xiang Chen 0016, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Huajun Chen
ACL (1)9
2025 Towards Text-Image Interleaved Retrieval
abstract
Xin Zhang, Ziqi Dai, Yongqi Li, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, Jun Yu, Wenjie Li, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xin Zhang 0097, Ziqi Dai, Yongqi Li 0001, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, Jun Yu 0002, Wenjie Li 0002, Min Zhang 0005
ACL (1)6
2025 Let LLMs Take on the Latest Challenges! A Chinese Dynamic Question Answering Benchmark
abstract
How to better evaluate the capabilities of Large Language Models (LLMs) is the focal point and hot topic in current LLMs research. Previous work has noted that due to the extremely high cost of iterative updates of LLMs, they are often unable to answer the latest dynamic questions well. To promote the improvement of Chinese LLMs’ ability to answer dynamic questions, in this paper, we introduce CDQA, a Chinese Dynamic QA benchmark containing question-answer pairs related to the latest news on the Chinese Internet. We obtain high-quality data through a pipeline that combines humans and models, and carefully classify the samples according to the frequency of answer changes to facilitate a more fine-grained observation of LLMs’ capabilities. We have also evaluated and analyzed mainstream and advanced Chinese LLMs on CDQA. Extensive experiments and valuable insights suggest that our proposed CDQA is challenging and worthy of more further study. We believe that the benchmark we provide will become one of the key data resources for improving LLMs’ Chinese question-answering ability in the future.
Zhikun Xu, Ruixue Ding, Xinyu Wang 0013, Boli Chen, Yong Jiang 0005, Hai-Tao Zheng 0002, Wenlian Lu, Pengjun Xie, Fei Huang 0002
COLING9
2025 Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language Models
abstract
Universal Multimodal Retrieval (UMR) aims to enable search across various modalities using a unified model, where queries and candidates can consist of pure text, images, or a combination of both. Previous work has attempted to adopt multimodal large language models (MLLMs) to realize UMR using only text data. However, our preliminary experiments demonstrate that more diverse multimodal training data can further unlock the potential of MLLMs. Despite its effectiveness, the existing multimodal training data is highly imbalanced in terms of modality, which motivates us to develop a training data synthesis pipeline and construct a large-scale, high-quality fused-modal training dataset. Based on the synthetic training data, we develop the General Multimodal Embedder (GME), an MLLM-based dense retriever designed for UMR. Furthermore, we construct a comprehensive UMR Benchmark (UMRB) to evaluate the effectiveness of our approach. Experimental results show that our method achieves state-of-the-art performance among existing UMR methods. Last, we provide in-depth analyses of model scaling and training strategies, and perform ablation studies on both the model and synthetic data.
Xin Zhang 0097, Yanzhao Zhang, Wen Xie 0006, Ziqi Dai, Dingkun Long, Pengjun Xie, Meishan Zhang, Wenjie Li 0002, Min Zhang 0005
CVPR7
2025 Detecting Knowledge Boundary of Vision Large Language Models by Sampling-Based Inference
abstract
Despite the advancements made in Vision Large Language Models (VLLMs), like text Large Language Models (LLMs), they have limitations in addressing questions that require real-time information or are knowledgeintensive.Indiscriminately adopting Retrieval Augmented Generation (RAG) techniques is an effective yet expensive way to enable models to answer queries beyond their knowledge scopes.To mitigate the dependence on retrieval and simultaneously maintain, or even improve, the performance benefits provided by retrieval, we propose a method to detect the knowledge boundary of VLLMs, allowing for more efficient use of techniques like RAG.Specifically, we propose a method with two variants that finetune a VLLM on an automatically constructed dataset for boundary identification.Experimental results on various types of Visual Question Answering datasets show that our method successfully depicts a VLLM's knowledge boundary, based on which we are able to reduce indiscriminate retrieval while maintaining or improving the performance.In addition, we show that the knowledge boundary identified by our method for one VLLM can be used as a surrogate boundary for other VLLMs.Code will be released at https://github.com/Chord-Che n-30/VLLM-KnowledgeBoundary
Xinyu Wang 0013, Yong Jiang 0005, Zhen Zhang 0008, Xinyu Geng, Pengjun Xie, Fei Huang 0002, Kewei Tu
EMNLP6
2025 DecoupleSearch: Decouple Planning and Search via Hierarchical Reward Modeling
abstract
Retrieval-Augmented Generation (RAG) systems have emerged as a pivotal methodology for enhancing Large Language Models (LLMs) through the dynamic integration of external knowledge.To further improve RAG's flexibility, Agentic RAG introduces autonomous agents into the workflow.However, Agentic RAG faces several challenges: (1) the success of each step depends on both high-quality planning and accurate search, (2) the lack of supervision for intermediate reasoning steps, and (3) the exponentially large candidate space for planning and searching.To address these challenges, we propose DecoupleSearch, a novel framework that decouples planning and search processes using dual value models, enabling independent optimization of plan reasoning and search grounding.Our approach constructs a reasoning tree, where each node represents planning and search steps.We leverage Monte Carlo Tree Search to assess the quality of each step.During inference, Hierarchical Beam Search iteratively refines planning and search candidates with dual value models.Extensive experiments across policy models of varying parameter sizes, demonstrate the effectiveness of our method.
Hao Sun 0015, Zile Qiao, Bo Wang 0134, Guoxin Chen, Yingyan Hou, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Yan Zhang 0117
EMNLP7
2025 ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents
abstract
Understanding information from visually rich documents remains a significant challenge for traditional Retrieval-Augmented Generation (RAG) methods.Existing benchmarks predominantly focus on image-based question answering (QA), overlooking the fundamental challenges of efficient retrieval, comprehension, and reasoning within dense visual documents.To bridge this gap, we introduce ViDoSeek, a novel dataset designed to evaluate RAG performance on visually rich documents requiring complex reasoning.Based on it, we identify key limitations in current RAG approaches: (i) purely visual retrieval methods struggle to effectively integrate both textual and visual features, and (ii) previous approaches often allocate insufficient reasoning tokens, limiting their effectiveness.To address these challenges, we propose ViDoRAG, a novel multi-agent RAG framework tailored for complex reasoning across visual documents.ViDoRAG employs a Gaussian Mixture Model (GMM)-based hybrid strategy to effectively handle multimodal retrieval.To further elicit the model's reasoning capabilities, we introduce an iterative agent workflow incorporating exploration, summarization, and reflection, providing a framework for investigating test-time scaling in RAG domains.Extensive experiments on ViDoSeek validate the effectiveness and generalization of our approach.Notably, ViDoRAG outperforms existing methods by over 10% on the competitive benchmark.The code is available at https: //github.com/Alibaba-NLP/ViDoRAG.
Qiuchen Wang, Ruixue Ding, Weiqi Wu, Pengjun Xie, Feng Zhao 0004
EMNLP6
2025 OmniThink: Expanding Knowledge Boundaries in Machine Writing through Thinking
abstract
Zekun Xi, Wenbiao Yin, Jizhan Fang, Jialong Wu, Runnan Fang, Yong Jiang, Pengjun Xie, Fei Huang, Huajun Chen, Ningyu Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zekun Xi, Wenbiao Yin, Jizhan Fang, Jialong Wu 0007, Runnan Fang, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Huajun Chen, Ningyu Zhang 0001
EMNLP7
2025 EvolveSearch: An Iterative Self-Evolving Search Agent
abstract
Ding-Chu Zhang, Yida Zhao, Jialong Wu, Liwen Zhang, Baixuan Li, Wenbiao Yin, Yong Jiang, Yu-Feng Li, Kewei Tu, Pengjun Xie, Fei Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Dingchu Zhang, Yida Zhao, Jialong Wu 0007, Baixuan Li, Wenbiao Yin, Yong Jiang 0005, Kewei Tu, Pengjun Xie, Fei Huang 0002
EMNLP10
2025 Benchmarking Agentic Workflow Generation
abstract
Large Language Models (LLMs), with their exceptional ability to handle a wide range of tasks, have driven significant advancements in tackling reasoning and planning tasks, wherein decomposing complex problems into executable workflows is a crucial step in this process. Existing workflow evaluation frameworks either focus solely on holistic performance or suffer from limitations such as restricted scenario coverage, simplistic workflow structures, and lax evaluation standards. To this end, we introduce WorfBench, a unified workflow generation benchmark with multi-faceted scenarios and intricate graph workflow structures. Additionally, we present WorfEval, a systemic evaluation protocol utilizing subsequence and subgraph matching algorithms to accurately quantify the LLM agent's workflow generation capabilities. Through comprehensive evaluations across different types of LLMs, we discover distinct gaps between the sequence planning capabilities and graph planning capabilities of LLM agents, with even GPT-4 exhibiting a gap of around 15%. We also train two open-source models and evaluate their generalization abilities on held-out tasks. Furthermore, we observe that the generated workflows can enhance downstream tasks, enabling them to achieve superior performance with less time during inference. Code and dataset are available at https://github.com/zjunlp/WorfBench.
Shuofei Qiao, Runnan Fang, Zhisong Qiu, Xiaobin Wang, Ningyu Zhang 0001, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Huajun Chen
ICLR7
2025 An End-to-End Model for Photo-Sharing Multi-Modal Dialogue Generation
abstract
Photo-Sharing Multi-modal dialogue generation requires a dialogue agent not only to generate text responses but also to share photos at the proper moment. Using image text caption as the bridge, a pipeline model integrates an image caption model, a text generation model, and an image generation model to handle this complex multi-modal task. However, representing the images with text captions may lose important visual details and information and cause error propagation in the complex dialogue system. Besides, the pipeline model isolates the three models separately because discrete image text captions hinder end-to-end gradient propagation. We propose the first end-to-end model for photo-sharing multi-modal dialogue generation, which integrates an image perceptron and an image generator with a large language model. The large language model employs the vision encoder to perceive visual images in the input end. For image generation in the output end, we propose a dynamic vocabulary transformation matrix and use straight-through and gumbel-softmax techniques to align the large language model and stable diffusion model and achieve end-to-end gradient propagation. We perform experiments on PhotoChat and DialogCC datasets to evaluate our end-to-end model. Compared with pipeline models, the end-to-end model gains state-of-the-art performances on various metrics of text and image generation.
Peiming Guo, Sinuo Liu, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, Min Zhang 0005
ICME5
2025 LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs - No Silver Bullet for LC or RAG Routing
abstract
As Large Language Model (LLM) context windows expand, the necessity of Retrieval-Augmented Generation (RAG) for integrating external knowledge is debated. Existing RAG vs. long-context (LC) LLM comparisons are often inconclusive due to benchmark limitations. We introduce LaRA, a novel benchmark with 2326 test cases across four QA tasks and three long context types, for rigorous evaluation. Our analysis of eleven LLMs reveals the optimal choice between RAG and LC depends on a complex interplay of model capabilities, context length, task type, and retrieval characteristics, offering actionable guidelines for practitioners. Our code and dataset is provided at:https://github.com/Alibaba-NLP/LaRA
Kuan Li, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Shuai Wang 0028, Minhao Cheng
ICML4
2025 VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning
abstract
Effectively retrieving, reasoning and understanding visually rich information remains a challenge for traditional Retrieval-Augmented Generation (RAG) methods. On the one hand, traditional text-based methods cannot handle visual-related information. On the other hand, current vision-based RAG approaches are often limited by fixed pipelines and frequently struggle to reason effectively due to the insufficient activation of the fundamental capabilities of models. As reinforcement learning (RL) has been proven to be beneficial for model reasoning, we introduce VRAG-RL, a novel RL framework tailored for complex reasoning across visually rich information. With this framework, VLMs interact with search engines, autonomously sampling single-turn or multi-turn reasoning trajectories with the help of visual perception tokens and undergoing continual optimization based on these samples. Our approach highlights key limitations of RL in RAG domains: (i) Prior Multi-modal RAG approaches tend to merely incorporate images into the context, leading to insufficient reasoning token allocation and neglecting visual-specific perception; and (ii) When models interact with search engines, their queries often fail to retrieve relevant information due to the inability to articulate requirements, thereby leading to suboptimal performance. To address these challenges, we define an action space tailored for visually rich inputs, with actions including cropping and scaling, allowing the model to gather information from a coarse-to-fine perspective. Furthermore, to bridge the gap between users' original inquiries and the retriever, we employ a simple yet effective reward that integrates query rewriting and retrieval performance with a model-based reward. Our VRAG-RL optimizes VLMs for RAG tasks using specially designed RL strategies, aligning the model with real-world applications. Extensive experiments on diverse and challenging benchmarks show that our VRAG-RL outperforms existing methods by 20\% (Qwen2.5-VL-7B) and 30\% (Qwen2.5-VL-3B), demonstrating the effectiveness of our approach. The code is available at https://github.com/Alibaba-NLP/VRAG.
Qiuchen Wang, Ruixue Ding, Pengjun Xie, Fei Huang 0002, Feng Zhao 0004
NeurIPS7
2025 WebDancer: Towards Autonomous Information Seeking Agency
abstract
Addressing intricate real-world problems necessitates in-depth information seeking and multi-step reasoning. Recent progress in agentic systems, exemplified by Deep Research, underscores the potential for autonomous multi-step research. In this work, we present a cohesive paradigm for building end-to-end agentic information seeking agents from a data-centric and training-stage perspective. Our approach consists of four key stages: (1) browsing data construction, (2) trajectories sampling, (3) supervised fine-tuning for effective cold start, and (4) reinforcement learning for enhanced generalisation. We instantiate this framework in a web agent based on the ReAct format, WebDancer. Empirical evaluations on the challenging GAIA and WebWalkerQA benchmarks demonstrate the strong performance of WebDancer, achieving considerable results and highlighting the efficacy of our training paradigm. Further analysis of agent training provides valuable insights and actionable, systematic pathways for developing more capable agentic models.
Jialong Wu 0007, Baixuan Li, Runnan Fang, Wenbiao Yin, Zhenglin Wang, Zhengwei Tao, Dingchu Zhang, Zekun Xi, Robert Tang, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Jingren Zhou 0001
NeurIPS12
2025 SSRB: Direct Natural Language Querying to Massive Heterogeneous Semi-Structured Data
abstract
Searching over semi-structured data with natural language (NL) queries has attracted sustained attention, enabling broader audiences to access information easily. As more applications, such as LLM agents and RAG systems, emerge to search and interact with semi-structured data, two major challenges have become evident: (1) the increasing diversity of domains and schema variations, making domain-customized solutions prohibitively costly; (2) the growing complexity of NL queries, which combine both exact field matching conditions and fuzzy semantic requirements, often involving multiple fields and implicit reasoning. These challenges make formal language querying or keyword-based search insufficient. In this work, we explore neural retrievers as a unified non-formal querying solution by directly index semi-structured collections and understand NL queries. We employ LLM-based automatic evaluation and build a large-scale semi-structured retrieval benchmark (SSRB) using LLM generation and filtering, containing 14M semi-structured objects from 99 different schemas across 6 domains, along with 8,485 test queries that combine both exact and fuzzy matching conditions. Our systematic evaluation of popular retrievers shows that current state-of-the-art models could achieve acceptable performance, yet they still lack precise understanding of matching constraints. While by in-domain training of dense retrievers, the performance can be significantly improved. We believe that our SSRB could serve as a valuable resource for future research in this area, and we hope to inspire further exploration of semi-structured retrieval with complex queries.
Xin Zhang 0097, Yanzhao Zhang, Dingkun Long, Yongqi Li 0001, Pengjun Xie, Meishan Zhang, Wenjie Li 0002, Min Zhang 0005, Philip S. Yu
NeurIPS7
2024 Three Heads Are Better than One: Improving Cross-Domain NER with Progressive Decomposed Network
abstract
Cross-domain named entity recognition (NER) tasks encourage NER models to transfer knowledge from data-rich source domains to sparsely labeled target domains. Previous works adopt the paradigms of pre-training on the source domain followed by fine-tuning on the target domain. However, these works ignore that general labeled NER source domain data can be easily retrieved in the real world, and soliciting more source domains could bring more benefits. Unfortunately, previous paradigms cannot efficiently transfer knowledge from multiple source domains. In this work, to transfer multiple source domains' knowledge, we decouple the NER task into the pipeline tasks of mention detection and entity typing, where the mention detection unifies the training object across domains, thus providing the entity typing with higher-quality entity mentions. Additionally, we request multiple general source domain models to suggest the potential named entities for sentences in the target domain explicitly, and transfer their knowledge to the target domain models through the knowledge progressive networks implicitly. Furthermore, we propose two methods to analyze in which source domain knowledge transfer occurs, thus helping us judge which source domain brings the greatest benefit. In our experiment, we develop a Chinese cross-domain NER dataset. Our model improved the F1 score by an average of 12.50% across 8 Chinese and English datasets compared to models without source domain data.
Xuming Hu, Zhaochen Hong, Yong Jiang 0005, Zhichao Lin, Xiaobin Wang, Pengjun Xie, Philip S. Yu
AAAI6
2024 EcomGPT: Instruction-Tuning Large Language Models with Chain-of-Task Tasks for E-commerce
abstract
Recently, instruction-following Large Language Models (LLMs) , represented by ChatGPT, have exhibited exceptional performance in general Natural Language Processing (NLP) tasks. However, the unique characteristics of E-commerce data pose significant challenges to general LLMs. An LLM tailored specifically for E-commerce scenarios, possessing robust cross-dataset/task generalization capabilities, is a pressing necessity. To solve this issue, in this work, we proposed the first E-commerce instruction dataset EcomInstruct, with a total of 2.5 million instruction data. EcomInstruct scales up the data size and task diversity by constructing atomic tasks with E-commerce basic data types, such as product information, user reviews. Atomic tasks are defined as intermediate tasks implicitly involved in solving a final task, which we also call Chain-of-Task tasks. We developed EcomGPT with different parameter scales by training the backbone model BLOOMZ with the EcomInstruct. Benefiting from the fundamental semantic understanding capabilities acquired from the Chain-of-Task tasks, EcomGPT exhibits excellent zero-shot generalization capabilities. Extensive experiments and human evaluations demonstrate that EcomGPT outperforms ChatGPT in term of cross-dataset/task generalization on E-commerce tasks. The EcomGPT will be public at https://github.com/Alibaba-NLP/EcomGPT.
Yangning Li, Shirong Ma, Xiaobin Wang, Shen Huang, Chengyue Jiang, Hai-Tao Zheng 0002, Pengjun Xie, Fei Huang 0002, Yong Jiang 0005
AAAI7
2024 SeqGPT: An Out-of-the-Box Large Language Model for Open Domain Sequence Understanding
abstract
Large language models (LLMs) have shown impressive abilities for open-domain NLP tasks. However, LLMs are sometimes too footloose for natural language understanding (NLU) tasks which always have restricted output and input format. Their performances on NLU tasks are highly related to prompts or demonstrations and are shown to be poor at performing several representative NLU tasks, such as event extraction and entity typing. To this end, we present SeqGPT, a bilingual (i.e., English and Chinese) open-source autoregressive model specially enhanced for open-domain natural language understanding. We express all NLU tasks with two atomic tasks, which define fixed instructions to restrict the input and output format but still ``open'' for arbitrarily varied label sets. The model is first instruction-tuned with extremely fine-grained labeled data synthesized by ChatGPT and then further fine-tuned by 233 different atomic tasks from 152 datasets across various domains. The experimental results show that SeqGPT has decent classification and extraction ability, and is capable of performing language understanding tasks on unseen domains. We also conduct empirical studies on the scaling of data and model size as well as on the transfer across tasks. Our models are accessible at https://github.com/Alibaba-NLP/SeqGPT.
Tianyu Yu 0002, Chengyue Jiang, Chao Lou, Shen Huang, Xiaobin Wang, Wei Liu 0131, Jiong Cai, Yangning Li, Kewei Tu, Hai-Tao Zheng 0002, Ningyu Zhang 0001, Pengjun Xie, Fei Huang 0002, Yong Jiang 0005
AAAI13
2024 GeoGLUE: A Chinese GeoGraphic Language Understanding Evaluation Benchmark
abstract
With the rapid growth of geographic applications, automatable and intelligent models are essential to be designed to handle the large volume of information. However, few researchers focus on geographic natural language processing, and there has never been a benchmark to build a unified standard. In this work, we propose a GeoGraphic Language Understanding Evaluation benchmark, named GeoGLUE. We collect data from open-released geographic resources and introduce six natural language understanding tasks, including geographic textual similarity on recall, geographic textual similarity on rerank, geographic elements tagging, geographic composition analysis, geographic where what cut, and geographic entity alignment. We also provide evaluation experiments and analysis of general baselines, indicating the effectiveness and significance of the GeoGLUE benchmark ( https://modelscope.cn/datasets/iic/GeoGLUE/summary ).
Ruixue Ding, Qiang Zhang 0051, Boli Chen, Pengjun Xie, Xin Li 0144, Fei Huang 0002
ADMA (5)6
2024 Chinese Sequence Labeling with Semi-Supervised Boundary-Aware Language Model Pre-training
abstract
Chinese sequence labeling tasks are sensitive to word boundaries. Although pretrained language models (PLM) have achieved considerable success in these tasks, current PLMs rarely consider boundary information explicitly. An exception to this is BABERT, which incorporates unsupervised statistical boundary information into Chinese BERT’s pre-training objectives. Building upon this approach, we input supervised high-quality boundary information to enhance BABERT’s learning, developing a semi-supervised boundary-aware PLM. To assess PLMs’ ability to encode boundaries, we introduce a novel “Boundary Information Metric” that is both simple and effective. This metric allows comparison of different PLMs without task-specific fine-tuning. Experimental results on Chinese sequence labeling datasets demonstrate that the improved BABERT version outperforms the vanilla version, not only in these tasks but also in broader Chinese natural language understanding tasks. Additionally, our proposed metric offers a convenient and accurate means of evaluating PLMs’ boundary awareness.
Longhui Zhang, Dingkun Long, Meishan Zhang, Yanzhao Zhang, Pengjun Xie, Min Zhang 0005
LREC/COLING5
2024 Geo-Encoder: A Chunk-Argument Bi-Encoder Framework for Chinese Geographic Re-Ranking
abstract
Yong Cao, Ruixue Ding, Boli Chen, Xianzhi Li, Min Chen, Daniel Hershcovich, Pengjun Xie, Fei Huang. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yong Cao 0001, Ruixue Ding, Boli Chen, Xianzhi Li 0001, Min Chen 0003, Daniel Hershcovich, Pengjun Xie, Fei Huang 0002
EACL (1)7
2024 Effective Demonstration Annotation for In-Context Learning via Language Model-Based Determinantal Point Process
abstract
In-context learning (ICL) is a few-shot learning paradigm that involves learning mappings through input-output pairs and appropriately applying them to new instances.Despite the remarkable ICL capabilities demonstrated by Large Language Models (LLMs), existing works are highly dependent on large-scale labeled support sets, not always feasible in practical scenarios.To refine this approach, we focus primarily on an innovative selective annotation mechanism, which precedes the standard demonstration retrieval.We introduce the Language Model-based Determinant Point Process (LM-DPP) that simultaneously considers the uncertainty and diversity of unlabeled instances for optimal selection.Consequently, this yields a subset for annotation that strikes a trade-off between the two factors.We apply LM-DPP to various language models, including GPT-J, LlaMA, and GPT-3.Experimental results on 9 NLU and 2 Generation datasets demonstrate that LM-DPP can effectively select canonical examples.Further analysis reveals that LLMs benefit most significantly from subsets that are both low uncertainty and high diversity.
Peng Wang 0104, Xiaobin Wang, Chao Lou, Shengyu Mao, Pengjun Xie, Yong Jiang 0005
EMNLP5
2024 Retrieved In-Context Principles from Previous Mistakes
abstract
In-context learning (ICL) has been instrumental in adapting large language models (LLMs) to downstream tasks using correct input-output examples.Recent advances have attempted to improve model performance through principles derived from mistakes, yet these approaches suffer from lack of customization and inadequate error coverage.To address these limitations, we propose Retrieved In-Context Principles (RICP), a novel teacherstudent framework.In RICP, the teacher model analyzes mistakes from the student model to generate reasons and insights for preventing similar mistakes.These mistakes are clustered based on their underlying reasons for developing task-level principles, enhancing the error coverage of principles.During inference, the most relevant mistakes for each question are retrieved to create question-level principles, improving the customization of the provided guidance.RICP is orthogonal to existing prompting methods and does not require intervention from the teacher model during inference.Experimental results across seven reasoning benchmarks reveal that RICP effectively enhances performance when applied to various prompting strategies.
Hao Sun 0015, Yong Jiang 0005, Bo Wang 0134, Yingyan Hou, Yan Zhang 0117, Pengjun Xie, Fei Huang 0002
EMNLP6
2024 Exploring Key Point Analysis with Pairwise Generation and Graph Partitioning
abstract
Xiao Li, Yong Jiang, Shen Huang, Pengjun Xie, Gong Cheng, Fei Huang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Xiao Li 0043, Yong Jiang 0005, Shen Huang, Pengjun Xie, Gong Cheng 0001, Fei Huang 0002
NAACL-HLT4
2024 WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language Models
abstract
Large language models (LLMs) need knowledge updates to meet the ever-growing world facts and correct the hallucinated responses, facilitating the methods of lifelong model editing. Where the updated knowledge resides in memories is a fundamental question for model editing. In this paper, we find that editing either long-term memory (direct model parameters) or working memory (non-parametric knowledge of neural network activations/representations by retrieval) will result in an impossible triangle---reliability, generalization, and locality can not be realized together in the lifelong editing settings. For long-term memory, directly editing the parameters will cause conflicts with irrelevant pretrained knowledge or previous edits (poor reliability and locality). For working memory, retrieval-based activations can hardly make the model understand the edits and generalize (poor generalization). Therefore, we propose WISE to bridge the gap between memories. In WISE, we design a dual parametric memory scheme, which consists of the main memory for the pretrained knowledge and a side memory for the edited knowledge. We only edit the knowledge in the side memory and train a router to decide which memory to go through when given a query. For continual editing, we devise a knowledge-sharding mechanism where different sets of edits reside in distinct subspaces of parameters, and are subsequently merged into a shared memory without conflicts. Extensive experiments show that WISE can outperform previous model editing methods and overcome the impossible triangle under lifelong model editing of question answering, hallucination, and out-of-distribution settings across trending LLM architectures, e.g., GPT, LLaMA, and Mistral.
Peng Wang 0104, Ningyu Zhang 0001, Ziwen Xu, Yunzhi Yao, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Huajun Chen
NeurIPS7
2024 Agent Planning with World Knowledge Model
abstract
Recent endeavors towards directly using large language models (LLMs) as agent models to execute interactive planning tasks have shown commendable results. Despite their achievements, however, they still struggle with brainless trial-and-error in global planning and generating hallucinatory actions in local planning due to their poor understanding of the "real" physical world. Imitating humans' mental world knowledge model which provides global prior knowledge before the task and maintains local dynamic knowledge during the task, in this paper, we introduce parametric World Knowledge Model (WKM) to facilitate agent planning. Concretely, we steer the agent model to self-synthesize knowledge from both expert and sampled trajectories. Then we develop WKM, providing prior task knowledge to guide the global planning and dynamic state knowledge to assist the local planning. Experimental results on three real-world simulated datasets with Mistral-7B, Gemma-7B, and Llama-3-8B demonstrate that our method can achieve superior performance compared to various strong baselines. Besides, we analyze to illustrate that our WKM can effectively alleviate the blind trial-and-error and hallucinatory action issues, providing strong support for the agent's understanding of the world. Other interesting findings include: 1) our instance-level task knowledge can generalize better to unseen tasks, 2) weak WKM can guide strong agent model planning, and 3) unified WKM training has promising potential for further development.
Shuofei Qiao, Runnan Fang, Ningyu Zhang 0001, Xiang Chen 0016, Shumin Deng, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Huajun Chen
NeurIPS8
2024 MCFC: A Momentum-Driven Clicked Feature Compressed Pre-trained Language Model for Information Retrieval
Ruixue Ding, Pengjun Xie
NLPCC (1)3
2024 Editing Personality For Large Language Models
Shengyu Mao, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Ningyu Zhang 0001
NLPCC (2)5
2023 Adversarial Self-Attention for Language Understanding
abstract
Deep neural models (e.g. Transformer) naturally learn spurious features, which create a ``shortcut'' between the labels and inputs, thus impairing the generalization and robustness. This paper advances self-attention mechanism to its robust variant for Transformer-based pre-trained language models (e.g. BERT). We propose Adversarial Self-Attention mechanism (ASA), which adversarially biases the attentions to effectively suppress the model reliance on features (e.g. specific keywords) and encourage its exploration of broader semantics. We conduct comprehensive evaluation across a wide range of tasks for both pre-training and fine-tuning stages. For pre-training, ASA unfolds remarkable performance gain compared to naive training for longer steps. For fine-tuning, ASA-empowered models outweigh naive models by a large margin considering both generalization and robustness.
Hongqiu Wu, Ruixue Ding, Hai Zhao 0001, Pengjun Xie, Fei Huang 0002, Min Zhang 0005
AAAI4
2023 Exploring Lottery Prompts for Pre-trained Language Models
abstract
Yulin Chen, Ning Ding, Xiaobin Wang, Shengding Hu, Haitao Zheng, Zhiyuan Liu, Pengjun Xie. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yulin Chen 0001, Ning Ding 0002, Xiaobin Wang, Shengding Hu, Hai-Tao Zheng 0002, Zhiyuan Liu 0001, Pengjun Xie
ACL (1)7
2023 MANNER: A Variational Memory-Augmented Model for Cross Domain Few-Shot Named Entity Recognition
abstract
This paper focuses on the task of cross domain few-shot named entity recognition (NER), which aims to adapt the knowledge learned from source domain to recognize named entities in target domain with only a few labeled examples.To address this challenging task, we propose MANNER, a variational memoryaugmented few-shot NER model.Specifically, MANNER uses a memory module to store information from the source domain and then retrieve relevant information from the memory to augment few-shot tasks in the target domain.In order to effectively utilize the information from memory, MANNER uses optimal transport to retrieve and process information from memory, which can explicitly adapt the retrieved information from source domain to target domain and improve the performance in the cross domain few-shot setting.We conduct experiments on both English and Chinese cross domain fewshot NER datasets, and the experimental results demonstrate that MANNER can achieve superior performance 1 .
Jinyuan Fang, Xiaobin Wang, Zaiqiao Meng, Pengjun Xie, Fei Huang 0002, Yong Jiang 0005
ACL (1)4
2023 Recall, Expand, and Multi-Candidate Cross-Encode: Fast and Accurate Ultra-Fine Entity Typing
abstract
Ultra-fine entity typing (UFET) predicts extremely free-formed types (e.g., president, politician) of a given entity mention (e.g., Joe Biden) in context.State-of-the-art (SOTA) methods use the cross-encoder (CE) based architecture.CE concatenates a mention (and its context) with each type and feeds the pair into a pretrained language model (PLM) to score their relevance.It brings deeper interaction between the mention and the type to reach better performance but has to perform N (the type set size) forward passes to infer all the types of a single mention.CE is therefore very slow in inference when the type set is large (e.g., N = 10k for UFET).To this end, we propose to perform entity typing in a recall-expand-filter manner.The recall and expansion stages prune the large type set and generate K (typically much smaller than N ) most relevant type candidates for each mention.At the filter stage, we use a novel model called MCCE to concurrently encode and score all these K candidates in only one forward pass to obtain the final type prediction.We investigate different model options for each stage and conduct extensive experiments to compare each option, experiments show that our method reaches SOTA performance on UFET and is thousands of times faster than the CE-based architecture.We also found our method is very effective in fine-grained (130 types) and coarse-grained (9 types) entity typing.
Chengyue Jiang, Wenyang Hui, Yong Jiang 0005, Xiaobin Wang, Pengjun Xie, Kewei Tu
ACL (1)5
2023 Do PLMs Know and Understand Ontological Knowledge?
abstract
Ontological knowledge, which comprises classes and properties and their relationships, is integral to world knowledge.It is significant to explore whether Pretrained Language Models (PLMs) know and understand such knowledge.However, existing PLM-probing studies focus mainly on factual knowledge, lacking a systematic probing of ontological knowledge.In this paper, we focus on probing whether PLMs store ontological knowledge and have a semantic understanding of the knowledge rather than rote memorization of the surface form.To probe whether PLMs know ontological knowledge, we investigate how well PLMs memorize: (1) types of entities; (2) hierarchical relationships among classes and properties, e.g., Person is a subclass of Animal and Member of Sports Team is a subproperty of Member of ; (3) domain and range constraints of properties, e.g., the subject of Member of Sports Team should be a Person and the object should be a Sports Team.To further probe whether PLMs truly understand ontological knowledge beyond memorization, we comprehensively study whether they can reliably perform logical reasoning with given knowledge according to ontological entailment rules.Our probing results show that PLMs can memorize certain ontological knowledge and utilize implicit knowledge in reasoning.However, both the memorizing and reasoning performances are less than perfect, indicating incomplete knowledge and understanding.
Weiqi Wu, Chengyue Jiang, Yong Jiang 0005, Pengjun Xie, Kewei Tu
ACL (1)4
2023 COMBO: A Complete Benchmark for Open KG Canonicalization
abstract
Open knowledge graph (KG) consists of (subject, relation, object) triples extracted from millions of raw text.The subject and object noun phrases and the relation in open KG have severe redundancy and ambiguity and need to be canonicalized.Existing datasets for open KG canonicalization only provide gold entitylevel canonicalization for noun phrases.In this paper, we present COMBO, a Complete Benchmark for Open KG canonicalization.Compared with existing datasets, we additionally provide gold canonicalization for relation phrases, gold ontology-level canonicalization for noun phrases, as well as source sentences from which triples are extracted.We also propose metrics for evaluating each type of canonicalization.On the COMBO dataset, we empirically compare previously proposed canonicalization methods as well as a few simple baseline methods based on pretrained language models.We find that properly encoding the phrases in a triple using pretrained language models results in better relation canonicalization and ontology-level canonicalization of the noun phrase.We release our dataset, baselines, and evaluation scripts at
Chengyue Jiang, Yong Jiang 0005, Weiqi Wu, Pengjun Xie, Kewei Tu
EACL5
2023 Text Representation Distillation via Information Bottleneck Principle
abstract
Pre-trained language models (PLMs) have recently shown great success in text representation field.However, the high computational cost and high-dimensional representation of PLMs pose significant challenges for practical applications.To make models more accessible, an effective method is to distill large models into smaller representation models.In order to relieve the issue of performance degradation after distillation, we propose a novel Knowledge Distillation method called IBKD.This approach is motivated by the Information Bottleneck principle and aims to maximize the mutual information between the final representation of the teacher and student model, while simultaneously reducing the mutual information between the student model's representation and the input data.This enables the student model to preserve important learned information while avoiding unnecessary information, thus reducing the risk of over-fitting.Empirical studies on two main downstream applications of text representation (Semantic Textual Similarity and Dense Retrieval tasks) demonstrate the effectiveness of our proposed approach 1 .
Yanzhao Zhang, Dingkun Long, Zehan Li, Pengjun Xie
EMNLP4
2023 MGeo: Multi-Modal Geographic Language Model Pre-Training
abstract
Query and point of interest (POI) matching is a core task in location-based services~(LBS), e.g., navigation maps. It connects users' intent with real-world geographic information. Lately, pre-trained language models (PLMs) have made notable advancements in many natural language processing (NLP) tasks. To overcome the limitation that generic PLMs lack geographic knowledge for query-POI matching, related literature attempts to employ continued pre-training based on domain-specific corpus. However, a query generally describes the geographic context (GC) about its destination and contains mentions of multiple geographic objects like nearby roads and regions of interest (ROIs). These diverse geographic objects and their correlations are pivotal to retrieving the most relevant POI. Text-based single-modal PLMs can barely make use of the important GC and are therefore limited. In this work, we propose a novel method for query-POI matching, namely Multi-modal Geographic language model (MGeo), which comprises a geographic encoder and a multi-modal interaction module. Representing GC as a new modality, MGeo is able to fully extract multi-modal correlations to perform accurate query-POI matching. Moreover, there exists no publicly available query-POI matching benchmark. Intending to facilitate further research, we build a new open-source large-scale benchmark for this topic, i.e., Geographic TExtual Similarity (GeoTES). The POIs come from an open-source geographic information system (GIS) and the queries are manually generated by annotators to prevent privacy issues. Compared with several strong baselines, the extensive experiment results and detailed ablation analyses demonstrate that our proposed multi-modal geographic pre-training method can significantly improve the query-POI matching capability of PLMs with or without users' locations. Our code and benchmark are publicly available at https://github.com/PhantomGrapes/MGeo.
Ruixue Ding, Boli Chen, Pengjun Xie, Fei Huang 0002, Xin Li 0144, Qiang Zhang 0051
SIGIR3
2023 Fine-Grained Domain Adaptation for Chinese Syntactic Processing
abstract
Syntactic processing is fundamental to natural language processing. It provides rich and comprehensive syntax information in sentences that could be potentially beneficial for downstream tasks. Recently, pretrained language models have shown great success in Chinese syntactic processing, which typically involves word segmentation, POS tagging, and dependency parsing. However, the on-going research never ends since performance would be degraded drastically when tested on a highly-discrepant domain. This problem is widely accepted as domain adaptation, where the test domain differs from the training domain in supervised learning. Self-training is one promising solution for it, and straightforward source-to-target adaptation has already shown remarkable effectiveness in previous work. While this strategy ignores the fact that sentences of the target domain sentences may have very different gaps from the source training domain. More specifically, sentences with large gaps might fail by direct self-training adaptation. To this end, we propose fine-grained domain adaptation for Chinese syntactic processing in this work, aiming to model the gaps between the source and the target domains accurately and progressively. The key idea is to divide the target domain into fine-grained subdomains by using a specified domain distance metric, and then perform gradual self-training on the subdomains. We further offer an intuitive theoretical illustration based on the theory of Kumar et al. (2020) approximately. In addition, a novel representation learning framework is proposed to encode fine-grained subdomains effectively, aiming to utilize the above idea fully. Experimental results on benchmark datasets show that our method can achieve significant improvements over a variety of baselines.
Meishan Zhang, Peiming Guo, Peijie Jiang, Dingkun Long, Yueheng Sun, Pengjun Xie, Min Zhang 0005
ACM Trans. Asian Low Resour. Lang. Inf. Process.7
2022 Parallel Instance Query Network for Named Entity Recognition
abstract
Yongliang Shen, Xiaobin Wang, Zeqi Tan, Guangwei Xu, Pengjun Xie, Fei Huang, Weiming Lu, Yueting Zhuang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yongliang Shen 0001, Xiaobin Wang, Zeqi Tan, Pengjun Xie, Fei Huang 0002, Weiming Lu 0001, Yueting Zhuang
ACL (1)5
2022 Domain-Specific NER via Retrieving Correlated Samples
abstract
Successful Machine Learning based Named Entity Recognition models could fail on texts from some special domains, for instance, Chinese addresses and e-commerce titles, where requires adequate background knowledge. Such texts are also difficult for human annotators. In fact, we can obtain some potentially helpful information from correlated texts, which have some common entities, to help the text understanding. Then, one can easily reason out the correct answer by referencing correlated samples. In this paper, we suggest enhancing NER models with correlated samples. We draw correlated samples by the sparse BM25 retriever from large-scale in-domain unlabeled data. To explicitly simulate the human reasoning process, we perform a training-free entity type calibrating by majority voting. To capture correlation features in the training stage, we suggest to model correlated samples by the transformer-based multi-instance cross-encoder. Empirical results on datasets of the above two domains show the efficacy of our methods.
Xin Zhang 0097, Yong Jiang 0005, Xiaobin Wang, Xuming Hu, Yueheng Sun, Pengjun Xie, Meishan Zhang
COLING6
2022 Modeling Label Correlations for Ultra-Fine Entity Typing with Neural Pairwise Conditional Random Field
abstract
Ultra-fine entity typing (UFET) aims to predict a wide range of type phrases that correctly describe the categories of a given entity mention in a sentence.Most recent works infer each entity type independently, ignoring the correlations between types, e.g., when an entity is inferred as a president, it should also be a politician and a leader.To this end, we use an undirected graphical model called pairwise conditional random field (PCRF) to formulate the UFET problem, in which the type variables are not only unarily influenced by the input but also pairwisely relate to all the other type variables.We use various modern backbones for entity typing to compute unary potentials, and derive pairwise potentials from type phrase representations that both capture prior semantic information and facilitate accelerated inference.We use mean-field variational inference for efficient type inference on very large type sets and unfold it as a neural network module to enable end-to-end training.Experiments on UFET show that the Neural-PCRF consistently outperforms its backbones with little cost and results in a competitive performance against crossencoder based SOTA while being thousands of times faster.We also find Neural-PCRF effective on a widely used fine-grained entity typing dataset with a smaller type set.We pack Neural-PCRF as a network module that can be plugged onto multi-label type classifiers with ease and release it in github.com/modelscope/ adaseq/examples/NPCRF.
Chengyue Jiang, Yong Jiang 0005, Weiqi Wu, Pengjun Xie, Kewei Tu
EMNLP4
2022 Unsupervised Boundary-Aware Language Model Pretraining for Chinese Sequence Labeling
abstract
Boundary information is critical for various Chinese language processing tasks, such as word segmentation, part-of-speech tagging, and named entity recognition.Previous studies usually resorted to the use of a high-quality external lexicon, where lexicon items can offer explicit boundary information.However, to ensure the quality of the lexicon, great human effort is always necessary, which has been generally ignored.In this work, we suggest unsupervised statistical boundary information instead, and propose an architecture to encode the information directly into pre-trained language models, resulting in Boundary-Aware BERT (BABERT).We apply BABERT for feature induction of Chinese sequence labeling tasks.Experimental results on ten benchmarks of Chinese sequence labeling demonstrate that BABERT can provide consistent improvements on all datasets.In addition, our method can complement previous supervised lexicon exploration, where further improvements can be achieved when integrated with external lexicon information.
Peijie Jiang, Dingkun Long, Yanzhao Zhang, Pengjun Xie, Meishan Zhang, Min Zhang 0005
EMNLP4
2022 AISHELL-NER: Named Entity Recognition from Chinese Speech
abstract
Named Entity Recognition (NER) from speech is among Spoken Language Understanding (SLU) tasks, aiming to extract semantic information from the speech signal. NER from speech is usually made through a two-step pipeline that consists of (1) processing the audio using an Automatic Speech Recognition (ASR) system and (2) applying an NER tagger to the ASR outputs. Recent works have shown the capability of the End-to-End (E2E) approach for NER from English and French speech, which is essentially entity-aware ASR. However, due to the many homophones and polyphones that exist in Chinese, NER from Chinese speech is effectively a more challenging task. In this paper, we introduce a new dataset AISEHLL-NER for NER from Chinese speech. Extensive experiments are conducted to explore the performance of several state-of-the-art methods. The results demonstrate that the performance could be improved by combining entity-aware ASR and pretrained NER tagger, which can be easily applied to the modern SLU pipeline. The dataset is publicly available at github.com/Alibaba-NLP/AISHELL-NER.
Boli Chen, Xiaobin Wang, Pengjun Xie, Meishan Zhang, Fei Huang 0002
ICASSP4
2022 Robust Self-Augmentation for Named Entity Recognition with Meta Reweighting
abstract
Linzhi Wu, Pengjun Xie, Jie Zhou, Meishan Zhang, Ma Chunping, Guangwei Xu, Min Zhang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Linzhi Wu, Pengjun Xie, Jie Zhou 0013, Meishan Zhang, Chunping Ma, Min Zhang 0005
NAACL-HLT2
2022 DAMO-NLP at NLPCC-2022 Task 2: Knowledge Enhanced Robust NER for Speech Entity Linking
Shen Huang, Yuchen Zhai, Xinwei Long, Yong Jiang 0005, Xiaobin Wang, Yin Zhang 0006, Pengjun Xie
NLPCC (2)7
2022 Multi-CPR: A Multi Domain Chinese Dataset for Passage Retrieval
abstract
Passage retrieval is a fundamental task in information retrieval (IR) research, which has drawn much attention recently. In the English field, the availability of large-scale annotated dataset (e.g, MS MARCO) and the emergence of deep pre-trained language models (e.g, BERT) has resulted in a substantial improvement of existing passage retrieval systems. However, in the Chinese field, especially for specific domains, passage retrieval systems are still immature due to quality-annotated dataset being limited by scale. Therefore, in this paper, we present a novel multi-domain Chinese dataset for passage retrieval (Multi-CPR). The dataset is collected from three different domains, including E-commerce, Entertainment video and Medical. Each dataset contains millions of passages and a certain amount of human annotated query-passage related pairs. We implement various representative passage retrieval methods as baselines. We find that the performance of retrieval models trained on dataset from general domain will inevitably decrease on specific domain. Nevertheless, a passage retrieval system built on in-domain annotated dataset can achieve significant improvement, which indeed demonstrates the necessity of domain labeled data for further optimization. We hope the release of the Multi-CPR dataset could benchmark Chinese passage retrieval task in specific domain and also make advances for future studies.
Dingkun Long, Qiong Gao, Kuan Zou, Pengjun Xie, Ruijie Guo, Guanjun Jiang, Luxi Xing
SIGIR5
2021 Knowledge-aware Named Entity Recognition with Alleviating Heterogeneity
abstract
Named Entity Recognition (NER) is a fundamental and important research topic for many downstream NLP tasks, aiming at detecting and classifying named entities (NEs) mentioned in unstructured text into pre-defined categories. Learning from labeled data only is far from enough when it comes to domain-specific or temporally-evolving entities (medical terminologies or restaurant names). Luckily, open-source Knowledge Bases (KBs) (Wikidata and Freebase) contain NEs that are manually labeled with predefined types in different domains, which is potentially beneficial to identify entity boundaries and recognize entity types more accurately. However, the type system of a domain-specific NER task is typically independent of that of current KBs and thus exhibits heterogeneity issue inevitably, which makes matching between the original NER and KB types (Person in NER potentially matches President in KBs) less likely, or introduces unintended noises without considering domain-specific knowledge (Band in NER should be mapped to Out_of_Entity_Types in the restaurant-related task). To better incorporate and denoise the abundant knowledge in KBs, we propose a new KB-aware NER framework (KaNa), which utilizes type-heterogeneous knowledge to improve NER. Specifically, for an entity mention along with a set of candidate entities that are linked from KBs, KaNa first uses a type projection mechanism that maps the mention type and entity types into a shared space to homogenize the heterogeneous entity types. Then, based on projected types, a noise detector filters out certain less-confident candidate entities in an unsupervised manner. Finally, the filtered mention-entity pairs are injected into a NER model as a graph to predict answers. The experimental results demonstrate KaNa's state-of-the-art performance on five public benchmark datasets from different domains.
Binling Nie, Ruixue Ding, Pengjun Xie, Fei Huang 0002, Chen Qian 0003, Luo Si
AAAI3
2021 Counterfactual Inference for Text Classification Debiasing
abstract
Chen Qian, Fuli Feng, Lijie Wen, Chunping Ma, Pengjun Xie. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Chen Qian 0003, Fuli Feng, Lijie Wen 0001, Chunping Ma, Pengjun Xie
ACL/IJCNLP (1)5
2021 Few-NERD: A Few-shot Named Entity Recognition Dataset
abstract
Ning Ding, Guangwei Xu, Yulin Chen, Xiaobin Wang, Xu Han, Pengjun Xie, Haitao Zheng, Zhiyuan Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Ning Ding 0002, Yulin Chen 0001, Xiaobin Wang, Xu Han 0007, Pengjun Xie, Hai-Tao Zheng 0002, Zhiyuan Liu 0001
ACL/IJCNLP (1)6
2021 Crowdsourcing Learning as Domain Adaptation: A Case Study on Named Entity Recognition
abstract
Xin Zhang, Guangwei Xu, Yueheng Sun, Meishan Zhang, Pengjun Xie. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Xin Zhang 0097, Yueheng Sun, Meishan Zhang, Pengjun Xie
ACL/IJCNLP (1)5
2021 A Fine-Grained Domain Adaption Model for Joint Word Segmentation and POS Tagging
abstract
Domain adaption for word segmentation and POS tagging is a challenging problem for Chinese lexical processing.Self-training is one promising solution for it, which struggles to construct a set of high-quality pseudo training instances for the target domain.Previous work usually assumes a universal sourceto-target adaption to collect such pseudo corpus, ignoring the different gaps from the target sentences to the source domain.In this work, we start from joint word segmentation and POS tagging, presenting a fine-grained domain adaption method to model the gaps accurately.We measure the gaps by one simple and intuitive metric, and adopt it to develop a pseudo target domain corpus based on finegrained subdomains incrementally.A novel domain-mixed representation learning model is proposed accordingly to encode the multiple subdomains effectively.The whole process is performed progressively for both corpus construction and model training.Experimental results on a benchmark dataset show that our method can gain significant improvements over a vary of baselines.Extensive analyses are performed to show the advantages of our final domain adaption model as well.
Peijie Jiang, Dingkun Long, Yueheng Sun, Meishan Zhang, Pengjun Xie
EMNLP (1)6
2021 Probing BERT in Hyperbolic Spaces
Boli Chen, Pengjun Xie, Chuanqi Tan, Mosha Chen, Liping Jing
ICLR4
2021 Prototypical Representation Learning for Relation Extraction
Ning Ding 0002, Xiaobin Wang, Rui Wang 0005, Pengjun Xie, Ying Shen 0001, Fei Huang 0002, Hai-Tao Zheng 0002, Rui Zhang 0003
ICLR6
2020 Coupling Distant Annotation and Adversarial Training for Cross-Domain Chinese Word Segmentation
abstract
Fully supervised neural approaches have achieved significant progress in the task of Chinese word segmentation (CWS).Nevertheless, the performance of supervised models tends to drop dramatically when they are applied to outof-domain data.Performance degradation is caused by the distribution gap across domains and the out of vocabulary (OOV) problem.In order to simultaneously alleviate these two issues, this paper proposes to couple distant annotation and adversarial training for crossdomain CWS.For distant annotation, we rethink the essence of "Chinese words" and design an automatic distant annotation mechanism that does not need any supervision or pre-defined dictionaries from the target domain.The approach could effectively explore domain-specific words and distantly annotate the raw texts for the target domain.For adversarial training, we develop a sentence-level training procedure to perform noise reduction and maximum utilization of the source domain information.Experiments on multiple realworld datasets across various domains show the superiority and robustness of our model, significantly outperforming previous state-ofthe-art cross-domain CWS methods.
Ning Ding 0002, Dingkun Long, Muhua Zhu, Pengjun Xie, Xiaobin Wang, Hai-Tao Zheng 0002
ACL5
2020 Hierarchy-Aware Global Model for Hierarchical Text Classification
abstract
Hierarchical text classification is an essential yet challenging subtask of multi-label text classification with a taxonomic hierarchy.Existing methods have difficulties in modeling the hierarchical label structure in a global view.Furthermore, they cannot make full use of the mutual interactions between the text feature space and the label space.In this paper, we formulate the hierarchy as a directed graph and introduce hierarchy-aware structure encoders for modeling label dependencies.Based on the hierarchy encoder, we propose a novel end-to-end hierarchy-aware global model (Hi-AGM) with two variants.A multi-label attention variant (HiAGM-LA) learns hierarchyaware label embeddings through the hierarchy encoder and conducts inductive fusion of labelaware text features.A text feature propagation model (HiAGM-TP) is proposed as the deductive variant that directly feeds text features into hierarchy encoders.Compared with previous works, both HiAGM-LA and HiAGM-TP achieve significant and consistent improvements on three benchmark datasets.
Jie Zhou 0013, Chunping Ma, Dingkun Long, Ning Ding 0002, Pengjun Xie, Gongshen Liu
ACL7
2020 Learning with Noise: Improving Distantly-Supervised Fine-grained Entity Typing via Automatic Relabeling
abstract
Fine-grained entity typing (FET) is a fundamental task for various entity-leveraging applications. Although great success has been made, existing systems still have challenges in handling noisy samples in training data introduced by distant supervision methods. To address these noise, previous studies either focus on processing the clean samples (i,e., have only one label) and noisy samples (i,e., have multiple labels) with different strategies or filtering the noisy labels based on the assumption that the distantly-supervised label set certainly contains the correct type label. In this paper, we propose a probabilistic automatic relabeling method which treats all training samples uniformly. Our method aims to estimate the pseudo-truth label distribution of each sample, and the pseudo-truth distribution will be treated as part of trainable parameters which are jointly updated during the training process. The proposed approach does not rely on any prerequisite or extra supervision, making it effective on real applications. Experiments on several benchmarks show that our method outperforms previous approaches and alleviates the noisy labeling problem.
Dingkun Long, Muhua Zhu, Pengjun Xie, Fei Huang 0002, Ji Wang 0001
IJCAI5
2019 A Neural Multi-digraph Model for Chinese NER with Gazetteers
abstract
Gazetteers were shown to be useful resources for named entity recognition (NER) (Ratinov and Roth, 2009).Many existing approaches to incorporating gazetteers into machine learning based NER systems rely on manually defined selection strategies or handcrafted templates, which may not always lead to optimal effectiveness, especially when multiple gazetteers are involved.This is especially the case for the task of Chinese NER, where the words are not naturally tokenized, leading to additional ambiguities.To automatically learn how to incorporate multiple gazetteers into an NER system, we propose a novel approach based on graph neural networks with a multidigraph structure that captures the information that the gazetteers offer.Experiments on various datasets show that our model is effective in incorporating rich gazetteer information while resolving ambiguities, outperforming previous approaches.
Ruixue Ding, Pengjun Xie, Wei Lu 0011, Linlin Li 0001, Luo Si
ACL (1)2