VLDB 2026 Research / reviewers in the wild / expert
Yong Jiang 0005
dblp:74/1552-5
· DBLP profile ↗
61ranked-venue papers
4as first author
42since 2021 · last 2026
0000-0003-4482-1559ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 60 · 4 first-author · 41 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Nested Browser-Use Learning for Agentic Information SeekingabstractBaixuan Li, Jialong Wu, Wenbiao Yin, Kuan Li, Zhongwang Zhang, Huifeng Yin, Zhengwei Tao, Liwen Zhang, Pengjun Xie, Jingren Zhou, Yong Jiang, Wentao Zhang, Zhiqiang Gao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Baixuan Li, Jialong Wu 0007, Wenbiao Yin, Kuan Li, Zhongwang Zhang, Huifeng Yin, Zhengwei Tao, Pengjun Xie, Jingren Zhou 0001, Yong Jiang 0005, Wentao Zhang 0001 |
ACL (1) | 11 |
| 2026 | STORM: A Spatio-Temporal Factor Model Based on Dual Vector Quantized Variational Autoencoders for Financial TradingabstractIn financial trading, factor models are widely used to price assets and capture excess returns from mispricing. Recently, we have witnessed the rise of variational autoencoder-based latent factor models, which learn latent factors self-adaptively. While these models focus on modeling overall market conditions, they often fail to effectively capture the temporal patterns of individual stocks. Additionally, representing multiple factors as single values simplifies the model but limits its ability to capture complex relationships and dependencies. As a result, the learned factors are of low quality and lack diversity, reducing their effectiveness and robustness across different trading periods. To address these issues, we propose a Spatio-Temporal factOR Model based on dual vector quantized variational autoencoders, named STORM, which extracts features of stocks from temporal and spatial perspectives, then fuses and aligns these features at the fine-grained and semantic level, and represents the factors as multi-dimensional embeddings. The discrete codebooks cluster similar factor embeddings, ensuring orthogonality and diversity, which helps distinguish between different factors and enables factor selection in financial trading. To show the performance of the proposed factor model, we apply it to two downstream experiments: portfolio management on two stock datasets and individual trading tasks on six specific stocks. The extensive experiments demonstrate STORM's flexibility in adapting to downstream tasks and superior performance over baseline models. Yilei Zhao 0001, Wentao Zhang 0007, Tingran Yang, Yong Jiang 0005, Fei Huang 0002, Wei Yang Bryan Lim |
WSDM | 4 |
| 2025 | WebWalker: Benchmarking LLMs in Web TraversalabstractRetrieval-augmented generation (RAG) demonstrates remarkable performance across tasks in open-domain question-answering. However, traditional search engines may retrieve shallow content, limiting the ability of LLMs to handle complex, multi-layered information. To address this, we introduce WebWalkerQA, a benchmark designed to assess the ability of LLMs to perform web traversal. It evaluates the capacity of LLMs to traverse a website’s subpages to extract high-quality data systematically. We propose WebWalker, which is a multi-agent framework that mimics human-like web navigation through an explore-critic paradigm. Extensive experimental results show that WebWalkerQA is challenging and demonstrates the effectiveness of RAG combined with WebWalker, through this horizontal and vertical integration in real-world scenarios. Jialong Wu 0007, Wenbiao Yin, Yong Jiang 0005, Zhenglin Wang, Zekun Xi, Runnan Fang, Linhai Zhang, Yulan He 0001, Pengjun Xie, Fei Huang 0002 |
ACL (1) | 3 |
| 2025 | Agentic Knowledgeable Self-awarenessabstractShuofei Qiao, Zhisong Qiu, Baochang Ren, Xiaobin Wang, Xiangyuan Ru, Ningyu Zhang, Xiang Chen, Yong Jiang, Pengjun Xie, Fei Huang, Huajun Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shuofei Qiao, Zhisong Qiu, Baochang Ren, Xiaobin Wang, Xiangyuan Ru, Ningyu Zhang 0001, Xiang Chen 0016, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Huajun Chen |
ACL (1) | 8 |
| 2025 | Let LLMs Take on the Latest Challenges! A Chinese Dynamic Question Answering BenchmarkabstractHow to better evaluate the capabilities of Large Language Models (LLMs) is the focal point and hot topic in current LLMs research. Previous work has noted that due to the extremely high cost of iterative updates of LLMs, they are often unable to answer the latest dynamic questions well. To promote the improvement of Chinese LLMs’ ability to answer dynamic questions, in this paper, we introduce CDQA, a Chinese Dynamic QA benchmark containing question-answer pairs related to the latest news on the Chinese Internet. We obtain high-quality data through a pipeline that combines humans and models, and carefully classify the samples according to the frequency of answer changes to facilitate a more fine-grained observation of LLMs’ capabilities. We have also evaluated and analyzed mainstream and advanced Chinese LLMs on CDQA. Extensive experiments and valuable insights suggest that our proposed CDQA is challenging and worthy of more further study. We believe that the benchmark we provide will become one of the key data resources for improving LLMs’ Chinese question-answering ability in the future. Zhikun Xu, Ruixue Ding, Xinyu Wang 0013, Boli Chen, Yong Jiang 0005, Hai-Tao Zheng 0002, Wenlian Lu, Pengjun Xie, Fei Huang 0002 |
COLING | 6 |
| 2025 | Detecting Knowledge Boundary of Vision Large Language Models by Sampling-Based InferenceabstractDespite the advancements made in Vision Large Language Models (VLLMs), like text Large Language Models (LLMs), they have limitations in addressing questions that require real-time information or are knowledgeintensive.Indiscriminately adopting Retrieval Augmented Generation (RAG) techniques is an effective yet expensive way to enable models to answer queries beyond their knowledge scopes.To mitigate the dependence on retrieval and simultaneously maintain, or even improve, the performance benefits provided by retrieval, we propose a method to detect the knowledge boundary of VLLMs, allowing for more efficient use of techniques like RAG.Specifically, we propose a method with two variants that finetune a VLLM on an automatically constructed dataset for boundary identification.Experimental results on various types of Visual Question Answering datasets show that our method successfully depicts a VLLM's knowledge boundary, based on which we are able to reduce indiscriminate retrieval while maintaining or improving the performance.In addition, we show that the knowledge boundary identified by our method for one VLLM can be used as a surrogate boundary for other VLLMs.Code will be released at https://github.com/Chord-Che n-30/VLLM-KnowledgeBoundary Xinyu Wang 0013, Yong Jiang 0005, Zhen Zhang 0008, Xinyu Geng, Pengjun Xie, Fei Huang 0002, Kewei Tu |
EMNLP | 3 |
| 2025 | DecoupleSearch: Decouple Planning and Search via Hierarchical Reward ModelingabstractRetrieval-Augmented Generation (RAG) systems have emerged as a pivotal methodology for enhancing Large Language Models (LLMs) through the dynamic integration of external knowledge.To further improve RAG's flexibility, Agentic RAG introduces autonomous agents into the workflow.However, Agentic RAG faces several challenges: (1) the success of each step depends on both high-quality planning and accurate search, (2) the lack of supervision for intermediate reasoning steps, and (3) the exponentially large candidate space for planning and searching.To address these challenges, we propose DecoupleSearch, a novel framework that decouples planning and search processes using dual value models, enabling independent optimization of plan reasoning and search grounding.Our approach constructs a reasoning tree, where each node represents planning and search steps.We leverage Monte Carlo Tree Search to assess the quality of each step.During inference, Hierarchical Beam Search iteratively refines planning and search candidates with dual value models.Extensive experiments across policy models of varying parameter sizes, demonstrate the effectiveness of our method. Hao Sun 0015, Zile Qiao, Bo Wang 0134, Guoxin Chen, Yingyan Hou, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Yan Zhang 0117 |
EMNLP | 6 |
| 2025 | OmniThink: Expanding Knowledge Boundaries in Machine Writing through ThinkingabstractZekun Xi, Wenbiao Yin, Jizhan Fang, Jialong Wu, Runnan Fang, Yong Jiang, Pengjun Xie, Fei Huang, Huajun Chen, Ningyu Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zekun Xi, Wenbiao Yin, Jizhan Fang, Jialong Wu 0007, Runnan Fang, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Huajun Chen, Ningyu Zhang 0001 |
EMNLP | 6 |
| 2025 | EvolveSearch: An Iterative Self-Evolving Search AgentabstractDing-Chu Zhang, Yida Zhao, Jialong Wu, Liwen Zhang, Baixuan Li, Wenbiao Yin, Yong Jiang, Yu-Feng Li, Kewei Tu, Pengjun Xie, Fei Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Dingchu Zhang, Yida Zhao, Jialong Wu 0007, Baixuan Li, Wenbiao Yin, Yong Jiang 0005, Kewei Tu, Pengjun Xie, Fei Huang 0002 |
EMNLP | 7 |
| 2025 | Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning AgentabstractMultimodal Retrieval Augmented Generation (mRAG) plays an important role in mitigating the “hallucination” issue inherent in multimodal large language models (MLLMs). Although promising, existing heuristic mRAGs typically predefined fixed retrieval processes, which causes two issues: (1) Non-adaptive Retrieval Queries. (2) Overloaded Retrieval Queries. However, these flaws cannot be adequately reflected by current knowledge-seeking visual question answering (VQA) datasets, since the most required knowledge can be readily obtained with a standard two-step retrieval. To bridge the dataset gap, we first construct Dyn-VQA dataset, consisting of three types of ``dynamic'' questions, which require complex knowledge retrieval strategies variable in query, tool, and time: (1) Questions with rapidly changing answers. (2) Questions requiring multi-modal knowledge. (3) Multi-hop questions. Experiments on Dyn-VQA reveal that existing heuristic mRAGs struggle to provide sufficient and precisely relevant knowledge for dynamic questions due to their rigid retrieval processes. Hence, we further propose the first self-adaptive planning agent for multimodal retrieval, **OmniSearch**. The underlying idea is to emulate the human behavior in question solution which dynamically decomposes complex multimodal questions into sub-question chains with retrieval action. Extensive experiments prove the effectiveness of our OmniSearch, also provide direction for advancing mRAG. Code and dataset will be open-sourced. Yangning Li, Xinyu Wang 0013, Yong Jiang 0005, Zhen Zhang 0008, Xinran Zheng, Hui Wang 0030, Hai-Tao Zheng 0002, Fei Huang 0002, Jingren Zhou 0001, Philip S. Yu |
ICLR | 4 |
| 2025 | Benchmarking Agentic Workflow GenerationabstractLarge Language Models (LLMs), with their exceptional ability to handle a wide range of tasks, have driven significant advancements in tackling reasoning and planning tasks, wherein decomposing complex problems into executable workflows is a crucial step in this process. Existing workflow evaluation frameworks either focus solely on holistic performance or suffer from limitations such as restricted scenario coverage, simplistic workflow structures, and lax evaluation standards. To this end, we introduce WorfBench, a unified workflow generation benchmark with multi-faceted scenarios and intricate graph workflow structures. Additionally, we present WorfEval, a systemic evaluation protocol utilizing subsequence and subgraph matching algorithms to accurately quantify the LLM agent's workflow generation capabilities. Through comprehensive evaluations across different types of LLMs, we discover distinct gaps between the sequence planning capabilities and graph planning capabilities of LLM agents, with even GPT-4 exhibiting a gap of around 15%. We also train two open-source models and evaluate their generalization abilities on held-out tasks. Furthermore, we observe that the generated workflows can enhance downstream tasks, enabling them to achieve superior performance with less time during inference. Code and dataset are available at https://github.com/zjunlp/WorfBench. Shuofei Qiao, Runnan Fang, Zhisong Qiu, Xiaobin Wang, Ningyu Zhang 0001, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Huajun Chen |
ICLR | 6 |
| 2025 | LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs - No Silver Bullet for LC or RAG RoutingabstractAs Large Language Model (LLM) context windows expand, the necessity of Retrieval-Augmented Generation (RAG) for integrating external knowledge is debated. Existing RAG vs. long-context (LC) LLM comparisons are often inconclusive due to benchmark limitations. We introduce LaRA, a novel benchmark with 2326 test cases across four QA tasks and three long context types, for rigorous evaluation. Our analysis of eleven LLMs reveals the optimal choice between RAG and LC depends on a complex interplay of model capabilities, context length, task type, and retrieval characteristics, offering actionable guidelines for practitioners. Our code and dataset is provided at:https://github.com/Alibaba-NLP/LaRA Kuan Li, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Shuai Wang 0028, Minhao Cheng |
ICML | 3 |
| 2025 | WebDancer: Towards Autonomous Information Seeking AgencyabstractAddressing intricate real-world problems necessitates in-depth information seeking and multi-step reasoning.
Recent progress in agentic systems, exemplified by Deep Research, underscores the potential for autonomous multi-step research.
In this work, we present a cohesive paradigm for building end-to-end agentic information seeking agents from a data-centric and training-stage perspective.
Our approach consists of four key stages: (1) browsing data construction, (2) trajectories sampling, (3) supervised fine-tuning for effective cold start, and (4) reinforcement learning for enhanced generalisation.
We instantiate this framework in a web agent based on the ReAct format, WebDancer.
Empirical evaluations on the challenging GAIA and WebWalkerQA benchmarks demonstrate the strong performance of WebDancer, achieving considerable results and highlighting the efficacy of our training paradigm.
Further analysis of agent training provides valuable insights and actionable, systematic pathways for developing more capable agentic models. Jialong Wu 0007, Baixuan Li, Runnan Fang, Wenbiao Yin, Zhenglin Wang, Zhengwei Tao, Dingchu Zhang, Zekun Xi, Robert Tang, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Jingren Zhou 0001 |
NeurIPS | 11 |
| 2024 | Three Heads Are Better than One: Improving Cross-Domain NER with Progressive Decomposed NetworkabstractCross-domain named entity recognition (NER) tasks encourage NER models to transfer knowledge from data-rich source domains to sparsely labeled target domains. Previous works adopt the paradigms of pre-training on the source domain followed by fine-tuning on the target domain. However, these works ignore that general labeled NER source domain data can be easily retrieved in the real world, and soliciting more source domains could bring more benefits. Unfortunately, previous paradigms cannot efficiently transfer knowledge from multiple source domains. In this work, to transfer multiple source domains' knowledge, we decouple the NER task into the pipeline tasks of mention detection and entity typing, where the mention detection unifies the training object across domains, thus providing the entity typing with higher-quality entity mentions. Additionally, we request multiple general source domain models to suggest the potential named entities for sentences in the target domain explicitly, and transfer their knowledge to the target domain models through the knowledge progressive networks implicitly. Furthermore, we propose two methods to analyze in which source domain knowledge transfer occurs, thus helping us judge which source domain brings the greatest benefit. In our experiment, we develop a Chinese cross-domain NER dataset. Our model improved the F1 score by an average of 12.50% across 8 Chinese and English datasets compared to models without source domain data. Xuming Hu, Zhaochen Hong, Yong Jiang 0005, Zhichao Lin, Xiaobin Wang, Pengjun Xie, Philip S. Yu |
AAAI | 3 |
| 2024 | EcomGPT: Instruction-Tuning Large Language Models with Chain-of-Task Tasks for E-commerceabstractRecently, instruction-following Large Language Models (LLMs) , represented by ChatGPT, have exhibited exceptional performance in general Natural Language Processing (NLP) tasks. However, the unique characteristics of E-commerce data pose significant challenges to general LLMs. An LLM tailored specifically for E-commerce scenarios, possessing robust cross-dataset/task generalization capabilities, is a pressing necessity. To solve this issue, in this work, we proposed the first E-commerce instruction dataset EcomInstruct, with a total of 2.5 million instruction data. EcomInstruct scales up the data size and task diversity by constructing atomic tasks with E-commerce basic data types, such as product information, user reviews. Atomic tasks are defined as intermediate tasks implicitly involved in solving a final task, which we also call Chain-of-Task tasks. We developed EcomGPT with different parameter scales by training the backbone model BLOOMZ with the EcomInstruct. Benefiting from the fundamental semantic understanding capabilities acquired from the Chain-of-Task tasks, EcomGPT exhibits excellent zero-shot generalization capabilities. Extensive experiments and human evaluations demonstrate that EcomGPT outperforms ChatGPT in term of cross-dataset/task generalization on E-commerce tasks. The EcomGPT will be public at https://github.com/Alibaba-NLP/EcomGPT. Yangning Li, Shirong Ma, Xiaobin Wang, Shen Huang, Chengyue Jiang, Hai-Tao Zheng 0002, Pengjun Xie, Fei Huang 0002, Yong Jiang 0005 |
AAAI | 9 |
| 2024 | SeqGPT: An Out-of-the-Box Large Language Model for Open Domain Sequence UnderstandingabstractLarge language models (LLMs) have shown impressive abilities for open-domain NLP tasks. However, LLMs are sometimes too footloose for natural language understanding (NLU) tasks which always have restricted output and input format. Their performances on NLU tasks are highly related to prompts or demonstrations and are shown to be poor at performing several representative NLU tasks, such as event extraction and entity typing. To this end, we present SeqGPT, a bilingual (i.e., English and Chinese) open-source autoregressive model specially enhanced for open-domain natural language understanding. We express all NLU tasks with two atomic tasks, which define fixed instructions to restrict the input and output format but still ``open'' for arbitrarily varied label sets. The model is first instruction-tuned with extremely fine-grained labeled data synthesized by ChatGPT and then further fine-tuned by 233 different atomic tasks from 152 datasets across various domains. The experimental results show that SeqGPT has decent classification and extraction ability, and is capable of performing language understanding tasks on unseen domains. We also conduct empirical studies on the scaling of data and model size as well as on the transfer across tasks. Our models are accessible at https://github.com/Alibaba-NLP/SeqGPT. Tianyu Yu 0002, Chengyue Jiang, Chao Lou, Shen Huang, Xiaobin Wang, Wei Liu 0131, Jiong Cai, Yangning Li, Kewei Tu, Hai-Tao Zheng 0002, Ningyu Zhang 0001, Pengjun Xie, Fei Huang 0002, Yong Jiang 0005 |
AAAI | 15 |
| 2024 | Effective Demonstration Annotation for In-Context Learning via Language Model-Based Determinantal Point ProcessabstractIn-context learning (ICL) is a few-shot learning paradigm that involves learning mappings through input-output pairs and appropriately applying them to new instances.Despite the remarkable ICL capabilities demonstrated by Large Language Models (LLMs), existing works are highly dependent on large-scale labeled support sets, not always feasible in practical scenarios.To refine this approach, we focus primarily on an innovative selective annotation mechanism, which precedes the standard demonstration retrieval.We introduce the Language Model-based Determinant Point Process (LM-DPP) that simultaneously considers the uncertainty and diversity of unlabeled instances for optimal selection.Consequently, this yields a subset for annotation that strikes a trade-off between the two factors.We apply LM-DPP to various language models, including GPT-J, LlaMA, and GPT-3.Experimental results on 9 NLU and 2 Generation datasets demonstrate that LM-DPP can effectively select canonical examples.Further analysis reveals that LLMs benefit most significantly from subsets that are both low uncertainty and high diversity. Peng Wang 0104, Xiaobin Wang, Chao Lou, Shengyu Mao, Pengjun Xie, Yong Jiang 0005 |
EMNLP | 6 |
| 2024 | Retrieved In-Context Principles from Previous MistakesabstractIn-context learning (ICL) has been instrumental in adapting large language models (LLMs) to downstream tasks using correct input-output examples.Recent advances have attempted to improve model performance through principles derived from mistakes, yet these approaches suffer from lack of customization and inadequate error coverage.To address these limitations, we propose Retrieved In-Context Principles (RICP), a novel teacherstudent framework.In RICP, the teacher model analyzes mistakes from the student model to generate reasons and insights for preventing similar mistakes.These mistakes are clustered based on their underlying reasons for developing task-level principles, enhancing the error coverage of principles.During inference, the most relevant mistakes for each question are retrieved to create question-level principles, improving the customization of the provided guidance.RICP is orthogonal to existing prompting methods and does not require intervention from the teacher model during inference.Experimental results across seven reasoning benchmarks reveal that RICP effectively enhances performance when applied to various prompting strategies. Hao Sun 0015, Yong Jiang 0005, Bo Wang 0134, Yingyan Hou, Yan Zhang 0117, Pengjun Xie, Fei Huang 0002 |
EMNLP | 2 |
| 2024 | FactCHD: Benchmarking Fact-Conflicting Hallucination Detection
Xiang Chen 0016, Duanzheng Song, Honghao Gui, Ningyu Zhang 0001, Yong Jiang 0005, Fei Huang 0002, Chengfei Lyu, Huajun Chen |
IJCAI | 6 |
| 2024 | Exploring Key Point Analysis with Pairwise Generation and Graph PartitioningabstractXiao Li, Yong Jiang, Shen Huang, Pengjun Xie, Gong Cheng, Fei Huang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xiao Li 0043, Yong Jiang 0005, Shen Huang, Pengjun Xie, Gong Cheng 0001, Fei Huang 0002 |
NAACL-HLT | 2 |
| 2024 | WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language ModelsabstractLarge language models (LLMs) need knowledge updates to meet the ever-growing world facts and correct the hallucinated responses, facilitating the methods of lifelong model editing. Where the updated knowledge resides in memories is a fundamental question for model editing. In this paper, we find that editing either long-term memory (direct model parameters) or working memory (non-parametric knowledge of neural network activations/representations by retrieval) will result in an impossible triangle---reliability, generalization, and locality can not be realized together in the lifelong editing settings. For long-term memory, directly editing the parameters will cause conflicts with irrelevant pretrained knowledge or previous edits (poor reliability and locality). For working memory, retrieval-based activations can hardly make the model understand the edits and generalize (poor generalization). Therefore, we propose WISE to bridge the gap between memories. In WISE, we design a dual parametric memory scheme, which consists of the main memory for the pretrained knowledge and a side memory for the edited knowledge. We only edit the knowledge in the side memory and train a router to decide which memory to go through when given a query. For continual editing, we devise a knowledge-sharding mechanism where different sets of edits reside in distinct subspaces of parameters, and are subsequently merged into a shared memory without conflicts. Extensive experiments show that WISE can outperform previous model editing methods and overcome the impossible triangle under lifelong model editing of question answering, hallucination, and out-of-distribution settings across trending LLM architectures, e.g., GPT, LLaMA, and Mistral. Peng Wang 0104, Ningyu Zhang 0001, Ziwen Xu, Yunzhi Yao, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Huajun Chen |
NeurIPS | 6 |
| 2024 | Agent Planning with World Knowledge ModelabstractRecent endeavors towards directly using large language models (LLMs) as agent models to execute interactive planning tasks have shown commendable results. Despite their achievements, however, they still struggle with brainless trial-and-error in global planning and generating hallucinatory actions in local planning due to their poor understanding of the "real" physical world. Imitating humans' mental world knowledge model which provides global prior knowledge before the task and maintains local dynamic knowledge during the task, in this paper, we introduce parametric World Knowledge Model (WKM) to facilitate agent planning. Concretely, we steer the agent model to self-synthesize knowledge from both expert and sampled trajectories. Then we develop WKM, providing prior task knowledge to guide the global planning and dynamic state knowledge to assist the local planning. Experimental results on three real-world simulated datasets with Mistral-7B, Gemma-7B, and Llama-3-8B demonstrate that our method can achieve superior performance compared to various strong baselines. Besides, we analyze to illustrate that our WKM can effectively alleviate the blind trial-and-error and hallucinatory action issues, providing strong support for the agent's understanding of the world. Other interesting findings include: 1) our instance-level task knowledge can generalize better to unseen tasks, 2) weak WKM can guide strong agent model planning, and 3) unified WKM training has promising potential for further development. Shuofei Qiao, Runnan Fang, Ningyu Zhang 0001, Xiang Chen 0016, Shumin Deng, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Huajun Chen |
NeurIPS | 7 |
| 2024 | Editing Personality For Large Language Models
Shengyu Mao, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Ningyu Zhang 0001 |
NLPCC (2) | 4 |
| 2024 | Model-Agnostic Knowledge Distillation Between Heterogeneous Models
Jiaxin Shen, Yanyao Liu, Yong Jiang 0005, Yufeng Chen 0005, Wenjuan Han |
NLPCC (1) | 3 |
| 2023 | MANNER: A Variational Memory-Augmented Model for Cross Domain Few-Shot Named Entity RecognitionabstractThis paper focuses on the task of cross domain few-shot named entity recognition (NER), which aims to adapt the knowledge learned from source domain to recognize named entities in target domain with only a few labeled examples.To address this challenging task, we propose MANNER, a variational memoryaugmented few-shot NER model.Specifically, MANNER uses a memory module to store information from the source domain and then retrieve relevant information from the memory to augment few-shot tasks in the target domain.In order to effectively utilize the information from memory, MANNER uses optimal transport to retrieve and process information from memory, which can explicitly adapt the retrieved information from source domain to target domain and improve the performance in the cross domain few-shot setting.We conduct experiments on both English and Chinese cross domain fewshot NER datasets, and the experimental results demonstrate that MANNER can achieve superior performance 1 . Jinyuan Fang, Xiaobin Wang, Zaiqiao Meng, Pengjun Xie, Fei Huang 0002, Yong Jiang 0005 |
ACL (1) | 6 |
| 2023 | Recall, Expand, and Multi-Candidate Cross-Encode: Fast and Accurate Ultra-Fine Entity TypingabstractUltra-fine entity typing (UFET) predicts extremely free-formed types (e.g., president, politician) of a given entity mention (e.g., Joe Biden) in context.State-of-the-art (SOTA) methods use the cross-encoder (CE) based architecture.CE concatenates a mention (and its context) with each type and feeds the pair into a pretrained language model (PLM) to score their relevance.It brings deeper interaction between the mention and the type to reach better performance but has to perform N (the type set size) forward passes to infer all the types of a single mention.CE is therefore very slow in inference when the type set is large (e.g., N = 10k for UFET).To this end, we propose to perform entity typing in a recall-expand-filter manner.The recall and expansion stages prune the large type set and generate K (typically much smaller than N ) most relevant type candidates for each mention.At the filter stage, we use a novel model called MCCE to concurrently encode and score all these K candidates in only one forward pass to obtain the final type prediction.We investigate different model options for each stage and conduct extensive experiments to compare each option, experiments show that our method reaches SOTA performance on UFET and is thousands of times faster than the CE-based architecture.We also found our method is very effective in fine-grained (130 types) and coarse-grained (9 types) entity typing. Chengyue Jiang, Wenyang Hui, Yong Jiang 0005, Xiaobin Wang, Pengjun Xie, Kewei Tu |
ACL (1) | 3 |
| 2023 | Do PLMs Know and Understand Ontological Knowledge?abstractOntological knowledge, which comprises classes and properties and their relationships, is integral to world knowledge.It is significant to explore whether Pretrained Language Models (PLMs) know and understand such knowledge.However, existing PLM-probing studies focus mainly on factual knowledge, lacking a systematic probing of ontological knowledge.In this paper, we focus on probing whether PLMs store ontological knowledge and have a semantic understanding of the knowledge rather than rote memorization of the surface form.To probe whether PLMs know ontological knowledge, we investigate how well PLMs memorize: (1) types of entities; (2) hierarchical relationships among classes and properties, e.g., Person is a subclass of Animal and Member of Sports Team is a subproperty of Member of ; (3) domain and range constraints of properties, e.g., the subject of Member of Sports Team should be a Person and the object should be a Sports Team.To further probe whether PLMs truly understand ontological knowledge beyond memorization, we comprehensively study whether they can reliably perform logical reasoning with given knowledge according to ontological entailment rules.Our probing results show that PLMs can memorize certain ontological knowledge and utilize implicit knowledge in reasoning.However, both the memorizing and reasoning performances are less than perfect, indicating incomplete knowledge and understanding. Weiqi Wu, Chengyue Jiang, Yong Jiang 0005, Pengjun Xie, Kewei Tu |
ACL (1) | 3 |
| 2023 | COMBO: A Complete Benchmark for Open KG CanonicalizationabstractOpen knowledge graph (KG) consists of (subject, relation, object) triples extracted from millions of raw text.The subject and object noun phrases and the relation in open KG have severe redundancy and ambiguity and need to be canonicalized.Existing datasets for open KG canonicalization only provide gold entitylevel canonicalization for noun phrases.In this paper, we present COMBO, a Complete Benchmark for Open KG canonicalization.Compared with existing datasets, we additionally provide gold canonicalization for relation phrases, gold ontology-level canonicalization for noun phrases, as well as source sentences from which triples are extracted.We also propose metrics for evaluating each type of canonicalization.On the COMBO dataset, we empirically compare previously proposed canonicalization methods as well as a few simple baseline methods based on pretrained language models.We find that properly encoding the phrases in a triple using pretrained language models results in better relation canonicalization and ontology-level canonicalization of the noun phrase.We release our dataset, baselines, and evaluation scripts at Chengyue Jiang, Yong Jiang 0005, Weiqi Wu, Pengjun Xie, Kewei Tu |
EACL | 2 |
| 2023 | One Model for All Domains: Collaborative Domain-Prefix Tuning for Cross-Domain NERabstractCross-domain NER is a challenging task to address the low-resource problem in practical scenarios. Previous typical solutions mainly obtain a NER model by pre-trained language models (PLMs) with data from a rich-resource domain and adapt it to the target domain. Owing to the mismatch issue among entity types in different domains, previous approaches normally tune all parameters of PLMs, ending up with an entirely new NER model for each domain. Moreover, current models only focus on leveraging knowledge in one general source domain while failing to successfully transfer knowledge from multiple sources to the target. To address these issues, we introduce Collaborative Domain-Prefix Tuning for cross-domain NER (CP-NER) based on text-to-text generative PLMs. Specifically, we present text-to-text generation grounding domain-related instructors to transfer knowledge to new domain NER tasks without structural modifications. We utilize frozen PLMs and conduct collaborative domain-prefix tuning to stimulate the potential of PLMs to handle NER tasks across various domains. Experimental results on the Cross-NER benchmark show that the proposed approach has flexible transfer ability and performs better on both one-source and multiple-source cross-domain NER tasks. Xiang Chen 0016, Lei Li 0040, Shuofei Qiao, Ningyu Zhang 0001, Chuanqi Tan, Yong Jiang 0005, Fei Huang 0002, Huajun Chen |
IJCAI | 6 |
| 2022 | Domain-Specific NER via Retrieving Correlated SamplesabstractSuccessful Machine Learning based Named Entity Recognition models could fail on texts from some special domains, for instance, Chinese addresses and e-commerce titles, where requires adequate background knowledge. Such texts are also difficult for human annotators. In fact, we can obtain some potentially helpful information from correlated texts, which have some common entities, to help the text understanding. Then, one can easily reason out the correct answer by referencing correlated samples. In this paper, we suggest enhancing NER models with correlated samples. We draw correlated samples by the sparse BM25 retriever from large-scale in-domain unlabeled data. To explicitly simulate the human reasoning process, we perform a training-free entity type calibrating by majority voting. To capture correlation features in the training stage, we suggest to model correlated samples by the transformer-based multi-instance cross-encoder. Empirical results on datasets of the above two domains show the efficacy of our methods. Xin Zhang 0097, Yong Jiang 0005, Xiaobin Wang, Xuming Hu, Yueheng Sun, Pengjun Xie, Meishan Zhang |
COLING | 2 |
| 2022 | Modeling Label Correlations for Ultra-Fine Entity Typing with Neural Pairwise Conditional Random FieldabstractUltra-fine entity typing (UFET) aims to predict a wide range of type phrases that correctly describe the categories of a given entity mention in a sentence.Most recent works infer each entity type independently, ignoring the correlations between types, e.g., when an entity is inferred as a president, it should also be a politician and a leader.To this end, we use an undirected graphical model called pairwise conditional random field (PCRF) to formulate the UFET problem, in which the type variables are not only unarily influenced by the input but also pairwisely relate to all the other type variables.We use various modern backbones for entity typing to compute unary potentials, and derive pairwise potentials from type phrase representations that both capture prior semantic information and facilitate accelerated inference.We use mean-field variational inference for efficient type inference on very large type sets and unfold it as a neural network module to enable end-to-end training.Experiments on UFET show that the Neural-PCRF consistently outperforms its backbones with little cost and results in a competitive performance against crossencoder based SOTA while being thousands of times faster.We also find Neural-PCRF effective on a widely used fine-grained entity typing dataset with a smaller type set.We pack Neural-PCRF as a network module that can be plugged onto multi-label type classifiers with ease and release it in github.com/modelscope/ adaseq/examples/NPCRF. Chengyue Jiang, Yong Jiang 0005, Weiqi Wu, Pengjun Xie, Kewei Tu |
EMNLP | 2 |
| 2022 | CAT-MNER: Multimodal Named Entity Recognition with Knowledge-Refined Cross-Modal AttentionabstractMultimodal named entity recognition (MNER) aims to detect and classify named entities in multimodal scenarios. It requires bridging the gap between natural language and visual context, which presents two-fold challenges: the cross-modal alignment is diversified, and the cross-modal interaction is sometimes implicit. Existing MNER methods are vulnerable to some implicit interactions and are prone to overlook the involved significant features. To tackle this problem, we novelly propose to refine the cross-modal attention by identifying and highlighting some task-salient features. The saliency of each feature is measured according to its correlation with the expanded entity label words derived from external knowledge bases. We further propose an end-to-end Transformer-based MNER framework, which holds neater architecture yet achieves better performance than previous methods. Extensive experiments are conducted to validate the merits of our method. Moreover, our method reveals a significant advantage in data efficiency and generalization ability. Xuwu Wang, Jiabo Ye, Zhixu Li, Yong Jiang 0005, Ming Yan 0008, Ji Zhang 0011, Yanghua Xiao |
ICME | 5 |
| 2022 | ITA: Image-Text Alignments for Multi-Modal Named Entity RecognitionabstractXinyu Wang, Min Gui, Yong Jiang, Zixia Jia, Nguyen Bach, Tao Wang, Zhongqiang Huang, Kewei Tu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Xinyu Wang 0013, Min Gui, Yong Jiang 0005, Zixia Jia, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Kewei Tu |
NAACL-HLT | 3 |
| 2022 | DAMO-NLP at NLPCC-2022 Task 2: Knowledge Enhanced Robust NER for Speech Entity Linking
Shen Huang, Yuchen Zhai, Xinwei Long, Yong Jiang 0005, Xiaobin Wang, Yin Zhang 0006, Pengjun Xie |
NLPCC (2) | 4 |
| 2021 | Multi-View Cross-Lingual Structured Prediction with Minimum SupervisionabstractZechuan Hu, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zechuan Hu, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu |
ACL/IJCNLP (1) | 2 |
| 2021 | Risk Minimization for Zero-shot Sequence LabelingabstractZechuan Hu, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zechuan Hu, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu |
ACL/IJCNLP (1) | 2 |
| 2021 | Improving Named Entity Recognition by External Context Retrieving and Cooperative LearningabstractXinyu Wang, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xinyu Wang 0013, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu |
ACL/IJCNLP (1) | 2 |
| 2021 | Automated Concatenation of Embeddings for Structured PredictionabstractXinyu Wang, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xinyu Wang 0013, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu |
ACL/IJCNLP (1) | 2 |
| 2021 | Structural Knowledge Distillation: Tractably Distilling Information for Structured PredictorabstractXinyu Wang, Yong Jiang, Zhaohui Yan, Zixia Jia, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xinyu Wang 0013, Yong Jiang 0005, Zhaohui Yan 0001, Zixia Jia, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu |
ACL/IJCNLP (1) | 2 |
| 2021 | Word Reordering for Zero-shot Cross-lingual Structured PredictionabstractAdapting word order from one language to another is a key problem in cross-lingual structured prediction.Current sentence encoders (e.g., RNN, Transformer with position embeddings) are usually word order sensitive.Even with uniform word form representations (MUSE, mBERT), word order discrepancies may hurt the adaptation of models.This paper builds structured prediction models with bag-of-words inputs.It introduces a new reordering module to organize words following the source language order, which learns taskspecific reordering strategies from a generalpurpose order predictor model.Experiments on zero-shot cross-lingual dependency parsing, POS tagging, and morphological tagging show that our model can significantly improve target language performances, especially for languages that are distant from the source language.1 Yong Jiang 0005, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Yuanbin Wu |
EMNLP (1) | 2 |
| 2021 | A Unified Encoding of Structures in Transition SystemsabstractTransition systems usually contain various dynamic structures (e.g., stacks, buffers).An ideal transition-based model should encode these structures completely and efficiently.Previous works relying on templates or neural network structures either only encode partial structure information or suffer from computation efficiency.In this paper, we propose a novel attention-based encoder unifying representation of all structures in a transition system.Specifically, we separate two views of items on structures, namely structure-invariant view and structure-dependent view.With the help of parallel-friendly attention network, we are able to encoding transition states with O(1) additional complexity (with respect to basic feature extractors).Experiments on the PTB and UD show that our proposed method significantly improves the test speed and achieves the best transition-based model, and is comparable to state-of-the-art methods. 1 Yong Jiang 0005, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Yuanbin Wu |
EMNLP (1) | 2 |
| 2021 | MuVER: Improving First-Stage Entity Retrieval with Multi-View Entity RepresentationsabstractEntity retrieval, which aims at disambiguating mentions to canonical entities from massive KBs, is essential for many tasks in natural language processing.Recent progress in entity retrieval shows that the dual-encoder structure is a powerful and efficient framework to nominate candidates if entities are only identified by descriptions.However, they ignore the property that meanings of entity mentions diverge in different contexts and are related to various portions of descriptions, which are treated equally in previous works.In this work, we propose Multi-View Entity Representations (MuVER), a novel approach for entity retrieval that constructs multi-view representations for entity descriptions and approximates the optimal view for mentions via a heuristic searching method.Our method achieves the state-ofthe-art performance on ZESHEL and improves the quality of candidates on three standard Entity Linking datasets 1 . Xinyin Ma, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Weiming Lu 0001 |
EMNLP (1) | 2 |
| 2020 | An Empirical Comparison of Unsupervised Constituency Parsing MethodsabstractUnsupervised constituency parsing aims to learn a constituency parser from a training corpus without parse tree annotations.While many methods have been proposed to tackle the problem, including statistical and neural methods, their experimental results are often not directly comparable due to discrepancies in datasets, data preprocessing, lexicalization, and evaluation metrics.In this paper, we first examine experimental settings used in previous work and propose to standardize the settings for better comparability between methods.We then empirically compare several existing methods, including decade-old and newly proposed ones, under the standardized settings on English and Japanese, two languages with different branching tendencies.We find that recent models do not show a clear advantage over decade-old models in our experiments.We hope our work can provide new insights into existing methods and facilitate future empirical evaluation of unsupervised constituency parsing. Jiong Cai, Yong Jiang 0005, Kewei Tu |
ACL | 4 |
| 2020 | Structure-Level Knowledge Distillation For Multilingual Sequence LabelingabstractMultilingual sequence labeling is a task of predicting label sequences using a single unified model for multiple languages.Compared with relying on multiple monolingual models, using a multilingual model has the benefit of a smaller model size, easier in online serving, and generalizability to low-resource languages.However, current multilingual models still underperform individual monolingual models significantly due to model capacity limitations.In this paper, we propose to reduce the gap between monolingual models and the unified multilingual model by distilling the structural knowledge of several monolingual models (teachers) to the unified multilingual model (student).We propose two novel KD methods based on structure-level information:(1) approximately minimizes the distance between the student's and the teachers' structurelevel probability distributions, (2) aggregates the structure-level knowledge to local distributions and minimizes the distance between two local probability distributions.Our experiments on 4 multilingual tasks with 25 datasets show that our approaches outperform several strong baselines and have stronger zero-shot generalizability than both the baseline model and teacher models. Xinyu Wang 0013, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Fei Huang 0002, Kewei Tu |
ACL | 2 |
| 2020 | A Survey of Unsupervised Dependency ParsingabstractSyntactic dependency parsing is an important task in natural language processing.Unsupervised dependency parsing aims to learn a dependency parser from sentences that have no annotation of their correct parse trees.Despite its difficulty, unsupervised parsing is an interesting research direction because of its capability of utilizing almost unlimited unannotated text data.It also serves as the basis for other research in low-resource parsing.In this paper, we survey existing approaches to unsupervised dependency parsing, identify two major classes of approaches, and discuss recent trends.We hope that our survey can provide insights for researchers and facilitate future research on this topic. Wenjuan Han, Yong Jiang 0005, Hwee Tou Ng, Kewei Tu |
COLING | 2 |
| 2020 | Second-Order Unsupervised Neural Dependency ParsingabstractMost of the unsupervised dependency parsers are based on first-order probabilistic generative models that only consider local parent-child information.Inspired by second-order supervised dependency parsing, we proposed a second-order extension of unsupervised neural dependency models that incorporate grandparent-child or sibling information.We also propose novel design of the neural parameterization and optimization methods of the dependency models.In secondorder models, the number of grammar rules grows cubically with the increase of vocabulary size, making it difficult to train lexicalized models that may contain thousands of words.To circumvent this problem while still benefiting from both second-order parsing and lexicalization, we use the agreement-based learning framework to jointly train a second-order unlexicalized model and a first-order lexicalized model.Experiments on multiple datasets show the effectiveness of our second-order models compared with recent state-of-the-art methods.Our joint model achieves a 10% improvement over the previous state-of-the-art parser on the full WSJ test set. Yong Jiang 0005, Wenjuan Han, Kewei Tu |
COLING | 2 |
| 2020 | Adversarial Attack and Defense of Structured Prediction ModelsabstractBuilding an effective adversarial attacker and elaborating on countermeasures for adversarial attacks for natural language processing (NLP) have attracted a lot of research in recent years.However, most of the existing approaches focus on classification problems.In this paper, we investigate attacks and defenses for structured prediction tasks in NLP.Besides the difficulty of perturbing discrete words and the sentence fluency problem faced by attackers in any NLP tasks, there is a specific challenge to attackers of structured prediction models: the structured output of structured prediction models is sensitive to small perturbations in the input.To address these problems, we propose a novel and unified framework that learns to attack a structured prediction model using a sequence-to-sequence model with feedbacks from multiple reference models of the same structured prediction task.Based on the proposed attack, we further reinforce the victim model with adversarial training, making its prediction more robust and accurate.We evaluate the proposed framework in dependency parsing and part-of-speech tagging.Automatic and human evaluations show that our proposed framework succeeds in both attacking state-of-the-art structured prediction models and boosting them with adversarial training. Wenjuan Han, Yong Jiang 0005, Kewei Tu |
EMNLP (1) | 3 |
| 2020 | AIN: Fast and Accurate Sequence Labeling with Approximate Inference NetworkabstractThe linear-chain Conditional Random Field (CRF) model is one of the most widely-used neural sequence labeling approaches.Exact probabilistic inference algorithms such as the forward-backward and Viterbi algorithms are typically applied in training and prediction stages of the CRF model.However, these algorithms require sequential computation that makes parallelization impossible.In this paper, we propose to employ a parallelizable approximate variational inference algorithm for the CRF model.Based on this algorithm, we design an approximate inference network that can be connected with the encoder of the neural CRF model to form an end-to-end network, which is amenable to parallelization for faster training and prediction.The empirical results show that our proposed approaches achieve a 12.7-fold improvement in decoding speed with long sentences and a competitive accuracy compared with the traditional CRF approach. Xinyu Wang 0013, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu |
EMNLP (1) | 2 |
| 2019 | Bidirectional Transition-Based Dependency ParsingabstractTransition-based dependency parsing is a fast and effective approach for dependency parsing. Traditionally, a transitionbased dependency parser processes an input sentence and predicts a sequence of parsing actions in a left-to-right manner. During this process, an early prediction error may negatively impact the prediction of subsequent actions. In this paper, we propose a simple framework for bidirectional transitionbased parsing. During training, we learn a left-to-right parser and a right-to-left parser separately. To parse a sentence, we perform joint decoding with the two parsers. We propose three joint decoding algorithms that are based on joint scoring, dual decomposition, and dynamic oracle respectively. Empirical results show that our methods lead to competitive parsing accuracy and our method based on dynamic oracle consistently achieves the best performance. Yunzhe Yuan, Yong Jiang 0005, Kewei Tu |
AAAI | 2 |
| 2019 | Enhancing Unsupervised Generative Dependency Parser with Contextual InformationabstractMost of the unsupervised dependency parsers are based on probabilistic generative models that learn the joint distribution of the given sentence and its parse.Probabilistic generative models usually explicit decompose the desired dependency tree into factorized grammar rules, which lack the global features of the entire sentence.In this paper, we propose a novel probabilistic model called discriminative neural dependency model with valence (D-NDMV) that generates a sentence and its parse from a continuous latent representation, which encodes global contextual information of the generated sentence.We propose two approaches to model the latent representation: the first deterministically summarizes the representation from the sentence and the second probabilistically models the representation conditioned on the sentence.Our approach can be regarded as a new type of autoencoder model to unsupervised dependency parsing that combines the benefits of both generative and discriminative techniques.In particular, our approach breaks the context-free independence assumption in previous generative approaches and therefore becomes more expressive.Our extensive experimental results on seventeen datasets from various sources show that our approach achieves competitive accuracy compared with both generative and discriminative state-of-the-art unsupervised dependency parsers. Wenjuan Han, Yong Jiang 0005, Kewei Tu |
ACL (1) | 2 |
| 2019 | Multilingual Grammar Induction with Continuous Language IdentificationabstractWenjuan Han, Ge Wang, Yong Jiang, Kewei Tu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Wenjuan Han, Ge Wang 0005, Yong Jiang 0005, Kewei Tu |
EMNLP/IJCNLP (1) | 3 |
| 2019 | A Regularization-based Framework for Bilingual Grammar InductionabstractYong Jiang, Wenjuan Han, Kewei Tu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yong Jiang 0005, Wenjuan Han, Kewei Tu |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Lexicalized Neural Unsupervised Dependency Parsing
Wenjuan Han, Yong Jiang 0005, Kewei Tu |
Neurocomputing | 2 |
| 2019 | Learning and evaluation of latent dependency forest models
Yong Jiang 0005, Kewei Tu |
Neural Comput. Appl. | 1 |
| 2018 | Maximum A Posteriori Inference in Sum-Product NetworksabstractSum-product networks (SPNs) are a class of probabilistic graphical models that allow tractable marginal inference. However, the maximum a posteriori (MAP) inference in SPNs is NP-hard. We investigate MAP inference in SPNs from both theoretical and algorithmic perspectives. For the theoretical part, we reduce general MAP inference to its special case without evidence and hidden variables; we also show that it is NP-hard to approximate the MAP problem to 2nε for fixed 0 ≤ ε < 1, where n is the input size. For the algorithmic part, we first present an exact MAP solver that runs reasonably fast and could handle SPNs with up to 1k variables and 150k arcs in our experiments. We then present a new approximate MAP solver with a good balance between speed and accuracy, and our comprehensive experiments on real-world datasets show that it has better overall performance than existing approximate solvers. Jun Mei, Yong Jiang 0005, Kewei Tu |
AAAI | 2 |
| 2017 | Latent Dependency Forest ModelsabstractProbabilistic modeling is one of the foundations of modern machine learning and artificial intelligence. In this paper, we propose a novel type of probabilistic models named latent dependency forest models (LDFMs). A LDFM models the dependencies between random variables with a forest structure that can change dynamically based on the variable values. It is therefore capable of modeling context-specific independence. We parameterize a LDFM using a first-order non-projective dependency grammar. Learning LDFMs from data can be formulated purely as a parameter learning problem, and hence the difficult problem of model structure learning is circumvented. Our experimental results show that LDFMs are competitive with existing probabilistic models. Shanbo Chu, Yong Jiang 0005, Kewei Tu |
AAAI | 2 |
| 2017 | CRF Autoencoder for Unsupervised Dependency ParsingabstractUnsupervised dependency parsing, which tries to discover linguistic dependency structures from unannotated data, is a very challenging task.Almost all previous work on this task focuses on learning generative models.In this paper, we develop an unsupervised dependency parsing model based on the CRF autoencoder.The encoder part of our model is discriminative and globally normalized which allows us to use rich features as well as universal linguistic priors.We propose an exact algorithm for parsing as well as a tractable learning algorithm.We evaluated the performance of our model on eight multilingual treebanks and found that our model achieved comparable performance with state-of-the-art approaches. Jiong Cai, Yong Jiang 0005, Kewei Tu |
EMNLP | 2 |
| 2017 | Dependency Grammar Induction with Neural Lexicalization and Big Training DataabstractWe study the impact of big models (in terms of the degree of lexicalization) and big data (in terms of the training corpus size) on dependency grammar induction.We experimented with L-DMV, a lexicalized version of Dependency Model with Valence (Klein and Manning, 2004) and L-NDMV, our lexicalized extension of the Neural Dependency Model with Valence (Jiang et al., 2016).We find that L-DMV only benefits from very small degrees of lexicalization and moderate sizes of training corpora.L-NDMV can benefit from big training data and lexicalization of greater degrees, especially when enhanced with good model initialization, and it achieves a result that is competitive with the current state-of-the-art. Wenjuan Han, Yong Jiang 0005, Kewei Tu |
EMNLP | 2 |
| 2017 | Combining Generative and Discriminative Approaches to Unsupervised Dependency Parsing via Dual DecompositionabstractUnsupervised dependency parsing aims to learn a dependency parser from unannotated sentences.Existing work focuses on either learning generative models using the expectation-maximization algorithm and its variants, or learning discriminative models using the discriminative clustering algorithm.In this paper, we propose a new learning strategy that learns a generative model and a discriminative model jointly based on the dual decomposition method.Our method is simple and general, yet effective to capture the advantages of both models and improve their learning results.We tested our method on the UD treebank and achieved a state-ofthe-art performance on thirty languages. Yong Jiang 0005, Wenjuan Han, Kewei Tu |
EMNLP | 1 |
| 2017 | Semi-supervised Structured Prediction with Neural CRF AutoencoderabstractIn this paper we propose an end-to-end neural CRF autoencoder (NCRF-AE) model for semi-supervised learning of sequential structured prediction problems. Our NCRF-AE consists of two parts: an encoder which is a CRF model enhanced by deep neural networks, and a decoder which is a generative model trying to reconstruct the input. Our model has a unified structure with different loss functions for labeled and unlabeled data with shared parameters. We developed a variation of the EM algorithm for optimizing both the encoder and the decoder simultaneously by decoupling their parameters. Our Experimental results over the Part-of-Speech (POS) tagging task on eight different languages, show that our model can outperform competitive systems in both supervised and semi-supervised scenarios. Xiao Zhang 0017, Yong Jiang 0005, Kewei Tu, Dan Goldwasser |
EMNLP | 2 |
| 2016 | Unsupervised Neural Dependency ParsingabstractUnsupervised dependency parsing aims to learn a dependency grammar from text annotated with only POS tags.Various features and inductive biases are often used to incorporate prior knowledge into learning.One useful type of prior information is that there exist correlations between the parameters of grammar rules involving different POS tags.Previous work employed manually designed features or special prior distributions to encode such information.In this paper, we propose a novel approach to unsupervised dependency parsing that uses a neural model to predict grammar rule probabilities based on distributed representation of POS tags.The distributed representation is automatically learned from data and captures the correlations between POS tags.Our experiments show that our approach outperforms previous approaches utilizing POS correlations and is competitive with recent state-of-the-art approaches on nine different languages. Yong Jiang 0005, Wenjuan Han, Kewei Tu |
EMNLP | 1 |