Wenhao Yu 0002

dblp:159/8117-2 · DBLP profile ↗
← Back
45ranked-venue papers
10as first author
40since 2021 · last 2026
0000-0002-4075-5980ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 41 · 7 first-author · 39 since 2021Databases, data management, data science and information retrieval · 10 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ReCode: Updating Code API Knowledge with Reinforcement Learning
abstract
Large Language Models (LLMs) exhibit remarkable code generation capabilities but falter when adapting to frequent updates in external library APIs. This critical limitation, stemming from reliance on outdated API knowledge from their training data, even with access to current documentation, impedes reliable code generation in dynamic environments. To tackle this issue, we propose ReCode (rule-based Reinforcement learning for Code Update), a novel framework that mimics human programmer adaptation to API changes. Specifically, we construct a dataset of approximately 2,000 data entries to train the LLMs to perform version migration based on updated information. Then, we introduce a modified string similarity metric for code evaluation as the reward for reinforcement learning. Our experiments demonstrate that ReCode substantially boosts LLMs' code generation performance in dynamic API scenarios, especially on the unseen CodeUpdateArena task. Crucially, compared to supervised fine-tuning, ReCode has less impact on LLMs' general code generation abilities. We apply ReCode on various LLMs and reinforcement learning algorithms (GRPO and DAPO), all achieving consistent improvements. Notably, after training, Qwen2.5-Coder-7B outperforms that of the 32B parameter code instruction-tuned model and the reasoning model with the same architecture.
Yunzhi Yao, Wenhao Yu 0002, Ningyu Zhang 0001
AAAI3
2026 Bidirectional LMs are Better Knowledge Memorizers? A Benchmark for Real-world Knowledge Injection
abstract
Yuwei Zhang, Wenhao Yu, Shangbin Feng, Yifan Zhu, Letian Peng, Jayanth Srinivasa, Gaowen Liu, Jingbo Shang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuwei Zhang 0001, Wenhao Yu 0002, Shangbin Feng, Letian Peng, Jayanth Srinivasa, Gaowen Liu, Jingbo Shang
ACL (1)2
2025 OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization
abstract
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Hongming Zhang, Tianqing Fang, Zhenzhong Lan, Dong Yu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Hongliang He 0002, Wenlin Yao, Kaixin Ma, Wenhao Yu 0002, Hongming Zhang 0009, Tianqing Fang, Zhen-Zhong Lan, Dong Yu 0001
ACL (1)4
2025 WebEvolver: Enhancing Web Agent Self-Improvement with Co-evolving World Model
abstract
Agent self-improvement, where agents autonomously train their underlying Large Language Model (LLM) on self-sampled trajectories, shows promising results but often stagnates in web environments due to limited exploration and under-utilization of pretrained web knowledge.To improve the performance of self-improvement, we propose a novel framework that introduces a co-evolving World Model LLM.This world model predicts the next observation based on the current observation and action within the web environment.The World Model serves dual roles: (1) as a virtual web server generating self-instructed training data to continuously refine the agent's policy, and (2) as an imagination engine during inference, enabling look-ahead simulation to guide action selection for the agent LLM.Experiments in real-world web environments (Mind2Web-Live, WebVoyager, and GAIAweb) show a 10% performance gain over existing self-evolving agents, demonstrating the efficacy and generalizability of our approach, without using any distillation from more powerful close-sourced models 1 .
Tianqing Fang, Hongming Zhang 0009, Zhisong Zhang, Kaixin Ma, Wenhao Yu 0002, Haitao Mi, Dong Yu 0001
EMNLP5
2025 Retrieval-augmented GUI Agents with Generative Guidelines
abstract
GUI agents powered by vision-language models (VLMs) show promise in automating complex digital tasks.However, their effectiveness in real-world applications is often limited by scarce training data and the inherent complexity of these tasks, which frequently require longtailed knowledge covering rare, unseen scenarios.We propose RAG-GUI , a lightweight VLM that leverages web tutorials at inference time.RAG-GUI is first warm-started via supervised finetuning (SFT) and further refined through self-guided rejection sampling finetuning (RSF).Designed to be model-agnostic, RAG-GUI functions as a generic plug-in that enhances any VLM-based agent.Evaluated across three distinct tasks, it consistently outperforms baseline agents and surpasses other inference baselines by 2.6% to 13.3% across two model sizes, demonstrating strong generalization and practical plug-and-play capabilities in real-world scenarios.
Ran Xu 0002, Kaixin Ma, Wenhao Yu 0002, Hongming Zhang 0009, Joyce C. Ho, Carl Yang 0001, Dong Yu 0001
EMNLP3
2025 DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?
abstract
Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) have demonstrated impressive language/vision reasoning abilities, igniting the recent trend of building agents for targeted applications such as shopping assistants or AI software engineers. Recently, many data science benchmarks have been proposed to investigate their performance in the data science domain. However, existing data science benchmarks still fall short when compared to real-world data science applications due to their simplified settings. To bridge this gap, we introduce DSBench, a comprehensive benchmark designed to evaluate data science agents with realistic tasks. This benchmark includes 466 data analysis tasks and 74 data modeling tasks, sourced from Eloquence and Kaggle competitions. DSBench offers a realistic setting by encompassing long contexts, multimodal task backgrounds, reasoning with large data files and multi-table structures, and performing end-to-end data modeling tasks. Our evaluation of state-of-the-art LLMs, LVLMs, and agents shows that they struggle with most tasks, with the best agent solving only 34.12% of data analysis tasks and achieving a 34.74% Relative Performance Gap (RPG). These findings underscore the need for further advancements in developing more practical, intelligent, and autonomous data science agents.
Liqiang Jing, Zhehui Huang, Xiaoyang Wang 0001, Wenlin Yao, Wenhao Yu 0002, Kaixin Ma, Hongming Zhang 0009, Xinya Du, Dong Yu 0001
ICLR5
2025 RepoGraph: Enhancing AI Software Engineering with Repository-level Code Graph
abstract
Large Language Models (LLMs) excel in code generation yet struggle with modern AI software engineering tasks. Unlike traditional function-level or file-level coding tasks, AI software engineering requires not only basic coding proficiency but also advanced skills in managing and interacting with code repositories. However, existing methods often overlook the need for repository-level code understanding, which is crucial for accurately grasping the broader context and developing effective solutions. On this basis, we present RepoGraph, a plug-in module that manages a repository-level structure for modern AI software engineering solutions. RepoGraph offers the desired guidance and serves as a repository-wide navigation for AI software engineers. We evaluate RepoGraph on the SWE-bench by plugging it into four different methods of two lines of approaches, where RepoGraph substantially boosts the performance of all systems, leading to a new state-of-the-art among open-source frameworks. Our analyses also demonstrate the extensibility and flexibility of RepoGraph by testing on another repo-level coding benchmark, CrossCodeEval. Our code is available at https://github.com/ozyyshr/RepoGraph.
Siru Ouyang, Wenhao Yu 0002, Kaixin Ma, Zilin Xiao, Zhihan Zhang 0001, Mengzhao Jia, Jiawei Han 0001, Hongming Zhang 0009, Dong Yu 0001
ICLR2
2025 LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
abstract
Recent large language model (LLM)-driven chat assistant systems have integrated memory components to track user-assistant chat histories, enabling more accurate and personalized responses. However, their long-term memory capabilities in sustained interactions remain underexplored. We introduce LongMemEval, a comprehensive benchmark designed to evaluate five core long-term memory abilities of chat assistants: information extraction, multi-session reasoning, temporal reasoning, knowledge updates, and abstention. With 500 meticulously curated questions embedded within freely scalable user-assistant chat histories, LongMemEval presents a significant challenge to existing long-term memory systems, with commercial chat assistants and long-context LLMs showing a 30% accuracy drop on memorizing information across sustained interactions. We then present a unified framework that breaks down the long-term memory design into three stages: indexing, retrieval, and reading. Built upon key experimental insights, we propose several memory design optimizations including session decomposition for value granularity, fact-augmented key expansion for indexing, and time-aware query expansion for refining the search scope. Extensive experiments show that these optimizations greatly improve both memory recall and downstream question answering on LongMemEval. Overall, our study provides valuable resources and guidance for advancing the long-term memory capabilities of LLM-based chat assistants, paving the way toward more personalized and reliable conversational AI. Our benchmark and code are publicly available at https://github.com/xiaowu0162/LongMemEval.
Di Wu 0054, Wenhao Yu 0002, Yuwei Zhang 0001, Kai-Wei Chang 0001, Dong Yu 0001
ICLR3
2025 BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
abstract
Task automation has been greatly empowered by the recent advances in Large Language Models (LLMs) via Python code, where the tasks range from software engineering development to general-purpose reasoning. While current benchmarks have shown that LLMs can solve tasks using programs like human developers, the majority of their evaluations are limited to short and self-contained algorithmic tasks or standalone function calls. Solving challenging and practical tasks requires the capability of utilizing **diverse function calls as tools** to efficiently implement functionalities like data analysis and web development. In addition, using multiple tools to solve a task needs compositional reasoning by accurately understanding **complex instructions**. Fulfilling both of these characteristics can pose a great challenge for LLMs. To assess how well LLMs can solve challenging and practical tasks via programs, we introduce BigCodeBench, a benchmark that challenges LLMs to invoke multiple function calls as tools from 139 libraries and 7 domains for 1,140 fine-grained tasks. To evaluate LLMs rigorously, each task encompasses 5.6 test cases with an average branch coverage of 99%. In addition, we propose a natural-language-oriented variant of BigCodeBench, BigCodeBench-Instruct, that automatically transforms the original docstrings into short instructions containing only essential information. Our extensive evaluation of 60 LLMs shows that **LLMs are not yet capable of following complex instructions to use function calls precisely, with scores up to 60%, significantly lower than the human performance of 97%**. The results underscore the need for further advancements in this area.
Terry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu 0011, Wenhao Yu 0002, Ratnadira Widyasari, Imam Nur Bani Yusuf, Haolan Zhan, Junda He, Indraneil Paul, Simon Brunner, Chen Gong 0005, James Hoang, Armel Zebaze, Xiaoheng Hong, Wen-Ding Li, Jean Kaddour, Zhihan Zhang 0001, Prateek Yadav
ICLR5
2025 Self-Regularization with Sparse Autoencoders for Controllable LLM-based Classification
abstract
Modern text classification methods heavily rely on contextual embeddings from large language models (LLMs). Compared to human-engineered features, these embeddings provide automatic and effective representations for classification model training. However, they also introduce a challenge: we lose the ability to manually remove unintended features, such as sensitive or task-irrelevant features, to guarantee regulatory compliance or improve the generalizability of classification models. This limitation arises because LLM embeddings are opaque and difficult to interpret. In this paper, we propose a novel framework to identify and regularize unintended features in the LLM latent space. Specifically, we first pre-train a sparse autoencoder (SAE) to extract interpretable features from LLM latent spaces. To ensure the SAE can capture task-specific features, we further fine-tune it on task-specific datasets. In training the classification model, we propose a simple and effective regularizer, by minimizing the similarity between the classifier weights and the identified unintended feature, to remove the impact of these unintended features on classification. We evaluate the proposed framework on three real-world tasks, including toxic chat detection, reward modeling, and disease diagnosis. Results show that the proposed self-regularization framework can improve the classifier's generalizability by regularizing those features that are not semantically correlated to the task. This work pioneers controllable text classification on LLM latent spaces by leveraging interpreted features to address generalizability, fairness, and privacy challenges. The code and data are publicly available at https://github.com/JacksonWuxs/Controllable_LLM_Classifier.
Xuansheng Wu, Wenhao Yu 0002, Xiaoming Zhai, Ninghao Liu 0001
KDD (2)2
2025 MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
abstract
The rapid extension of context windows in large vision-language models has given rise to long-context vision-language models (LCVLMs), which are capable of handling hundreds of images with interleaved text tokens in a single forward pass. In this work, we introduce MMLongBench, the first benchmark covering a diverse set of long-context vision-language tasks, to evaluate LCVLMs effectively and thoroughly. MMLongBench is composed of 13,331 examples spanning five different categories of downstream tasks, such as Visual RAG and Many-Shot ICL. It also provides broad coverage of image types, including various natural and synthetic images. To assess the robustness of the models to different input lengths, all examples are delivered at five standardized input lengths (8K-128K tokens) via a cross-modal tokenization scheme that combines vision patches and text tokens. Through a thorough benchmarking of 46 closed-source and open-source LCVLMs, we provide a comprehensive analysis of the current models' vision-language long-context ability. Our results show that: i) performance on a single task is a weak proxy for overall long-context capability; ii) both closed-source and open-source models face challenges in long-context vision-language tasks, indicating substantial room for future improvement; iii) models with stronger reasoning ability tend to exhibit better long-context performance. By offering wide task coverage, various image types, and rigorous length control, MMLongBench provides the missing foundation for diagnosing and advancing the next generation of LCVLMs.
Zhaowei Wang 0003, Wenhao Yu 0002, Xiyu Ren, Yu Zhao 0043, Rohit Saxena, Ginny Y. Wong, Simon See, Pasquale Minervini, Yangqiu Song, Mark Steedman
NeurIPS2
2024 WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
abstract
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, Dong Yu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Hongliang He 0002, Wenlin Yao, Kaixin Ma, Wenhao Yu 0002, Yong Dai 0001, Hongming Zhang 0009, Zhen-Zhong Lan, Dong Yu 0001
ACL (1)4
2024 PLUG: Leveraging Pivot Language in Cross-Lingual Instruction Tuning
abstract
Zhihan Zhang, Dong-Ho Lee, Yuwei Fang, Wenhao Yu, Mengzhao Jia, Meng Jiang, Francesco Barbieri. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zhihan Zhang 0001, Yuwei Fang, Wenhao Yu 0002, Mengzhao Jia, Meng Jiang 0001, Francesco Barbieri
ACL (1)4
2024 Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models
abstract
Retrieval-augmented language model (RALM) represents a significant advancement in mitigating factual hallucination by leveraging external knowledge sources.However, the reliability of the retrieved information is not always guaranteed, and the retrieval of irrelevant data can mislead the response generation.Moreover, standard RALMs frequently neglect their intrinsic knowledge due to the interference from retrieved information.In instances where the retrieved information is irrelevant, RALMs should ideally utilize their intrinsic knowledge or, in the absence of both intrinsic and retrieved knowledge, opt to respond with "unknown" to avoid hallucination.In this paper, we introduces CHAIN-OF-NOTE (CON), a novel approach to improve robustness of RALMs in facing noisy, irrelevant documents and in handling unknown scenarios.The core idea of CON is to generate sequential reading notes for each retrieved document, enabling a thorough evaluation of their relevance to the given question and integrating this information to formulate the final answer.Our experimental results show that GPT-4, when equipped with CON, outperforms the CHAIN-OF-THOUGHT approach.Besides, we utilized GPT-4 to create 10K CON data, subsequently trained on LLaMa-2 7B model.Our experiments across four open-domain QA benchmarks show that fine-tuned RALMs equipped with CON significantly outperform standard fine-tuned RALMs.
Wenhao Yu 0002, Hongming Zhang 0009, Xiaoman Pan, Peixin Cao, Kaixin Ma, Hongwei Wang 0001, Dong Yu 0001
EMNLP1
2024 Dense X Retrieval: What Retrieval Granularity Should We Use?
abstract
Dense retrieval has become a prominent method to obtain relevant context or world knowledge in open-domain NLP tasks.When we use a learned dense retriever on a retrieval corpus at inference time, an often-overlooked design choice is the retrieval unit in which the corpus is indexed, e.g.document, passage, or sentence.We discover that the retrieval unit choice significantly impacts the performance of both retrieval and downstream tasks.Distinct from the typical approach of using passages or sentences, we introduce a novel retrieval unit, proposition, for dense retrieval.Propositions are defined as atomic expressions within text, each encapsulating a distinct factoid and presented in a concise, self-contained natural language format.We conduct an empirical comparison of different retrieval granularity.Our experiments reveal that indexing a corpus by fine-grained units such as propositions significantly outperforms passage-level units in retrieval tasks.Moreover, constructing prompts with fine-grained retrieved units for retrievalaugmented language models improves the performance of downstream QA tasks given a specific computation budget.
Hongwei Wang 0010, Wenhao Yu 0002, Kaixin Ma, Hongming Zhang 0009, Dong Yu 0001
EMNLP4
2024 Learn Beyond The Answer: Training Language Models with Reflection for Mathematical Reasoning
abstract
Supervised fine-tuning enhances the problemsolving abilities of language models across various mathematical reasoning tasks.To maximize such benefits, existing research focuses on broadening the training set with various data augmentation techniques, which is effective for standard single-round question-answering settings.Our work introduces a novel technique aimed at cultivating a deeper understanding of the training problems at hand, enhancing performance not only in standard settings but also in more complex scenarios that require reflective thinking.Specifically, we propose reflective augmentation, a method that embeds problem reflection into each training instance.It trains the model to consider alternative perspectives and engage with abstractions and analogies, thereby fostering a thorough comprehension through reflective reasoning.Extensive experiments validate the achievement of our aim, underscoring the unique advantages of our method and its complementary nature relative to existing augmentation techniques. 1 Question Answer Question
Zhihan Zhang 0001, Tao Ge 0001, Zhenwen Liang, Wenhao Yu 0002, Dian Yu 0001, Mengzhao Jia, Dong Yu 0001, Meng Jiang 0001
EMNLP4
2024 Sub-Sentence Encoder: Contrastive Learning of Propositional Semantic Representations
abstract
Sihao Chen, Hongming Zhang, Tong Chen, Ben Zhou, Wenhao Yu, Dian Yu, Baolin Peng, Hongwei Wang, Dan Roth, Dong Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Hongming Zhang 0009, Ben Zhou, Wenhao Yu 0002, Dian Yu 0001, Baolin Peng, Hongwei Wang 0010, Dan Roth 0001, Dong Yu 0001
NAACL-HLT5
2023 APOLLO: A Simple Approach for Adaptive Pretraining of Language Models for Logical Reasoning
abstract
Soumya Sanyal, Yichong Xu, Shuohang Wang, Ziyi Yang, Reid Pryzant, Wenhao Yu, Chenguang Zhu, Xiang Ren. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Soumya Sanyal 0001, Yichong Xu, Shuohang Wang, Ziyi Yang 0011, Reid Pryzant, Wenhao Yu 0002, Chenguang Zhu 0001, Xiang Ren 0001
ACL (1)6
2023 A Survey of Deep Learning for Mathematical Reasoning
abstract
Mathematical reasoning is a fundamental aspect of human intelligence and is applicable in various fields, including science, engineering, finance, and everyday life.The development of artificial intelligence (AI) systems capable of solving math problems and proving theorems in language has garnered significant interest in the fields of machine learning and natural language processing.For example, mathematics serves as a testbed for aspects of reasoning that are challenging for powerful deep learning models, driving new algorithmic and modeling advances.On the other hand, recent advances in large-scale neural language models have opened up new benchmarks and opportunities to use deep learning for mathematical reasoning.In this survey paper, we review the key tasks, datasets, and methods at the intersection of mathematical reasoning and deep learning over the past decade.We also evaluate existing benchmarks and methods, and discuss future research directions in this domain.Question: Bod has 2 apples and David has 5 apples.How many apples do they have in total?Rationale: x = 2 + 5 Solution: 7
Pan Lu, Liang Qiu 0001, Wenhao Yu 0002, Sean Welleck, Kai-Wei Chang 0001
ACL (1)3
2023 A Survey of Multi-task Learning in Natural Language Processing: Regarding Task Relatedness and Training Methods
abstract
Multi-task learning (MTL) has become increasingly popular in natural language processing (NLP) because it improves the performance of related tasks by exploiting their commonalities and differences.Nevertheless, it is still not understood very well how multi-task learning can be implemented based on the relatedness of training tasks.In this survey, we review recent advances of multi-task learning methods in NLP, with the aim of summarizing them into two general multi-task training methods based on their task relatedness: (i) joint training and (ii) multi-step training.We present examples in various NLP downstream applications, summarize the task relationships and discuss future directions of this promising topic.
Zhihan Zhang 0001, Wenhao Yu 0002, Mengxia Yu, Zhichun Guo, Meng Jiang 0001
EACL2
2023 IfQA: A Dataset for Open-domain Question Answering under Counterfactual Presuppositions
abstract
Although counterfactual reasoning is a fundamental aspect of intelligence, the lack of largescale counterfactual open-domain questionanswering (QA) benchmarks makes it difficult to evaluate and improve models on this ability.To address this void, we introduce the first such dataset, named IfQA, where each question is based on a counterfactual presupposition via an "if" clause.Such questions require models to go beyond retrieving direct factual knowledge from the Web: they must identify the right information to retrieve and reason about an imagined situation that may even go against the facts built into their parameters.The IfQA dataset contains 3,800 questions that were annotated by crowdworkers on relevant Wikipedia passages.Empirical analysis reveals that the IfQA dataset is highly challenging for existing open-domain QA methods, including supervised retrieve-then-read pipeline methods (F1 score 44.5), as well as recent few-shot approaches such as chain-of-thought prompting with ChatGPT (F1 score 57.2).We hope the unique challenges posed by IfQA will push open-domain QA research on both retrieval and reasoning fronts, while also helping endow counterfactual reasoning abilities to today's language understanding models.The IfQA dataset can be found and downloaded at https://allenai.org/data/ifqa.
Wenhao Yu 0002, Meng Jiang 0001, Peter Clark, Ashish Sabharwal
EMNLP1
2023 Let GPT be a Math Tutor: Teaching Math Word Problem Solvers with Customized Exercise Generation
abstract
In this paper, we present a novel approach for distilling math word problem solving capabilities from large language models (LLMs) into smaller, more efficient student models. Our approach is designed to consider the student model's weaknesses and foster a tailored learning experience by generating targeted exercises aligned with educational science principles, such as knowledge tracing and personalized learning. Concretely, we let GPT-3 be a math tutor and run two steps iteratively: 1) assessing the student model's current learning status on a GPT-generated exercise book, and 2) improving the student model by training it with tailored exercise samples generated by GPT-3. Experimental results reveal that our approach outperforms LLMs (e.g., GPT-3 and PaLM) in accuracy across three distinct benchmarks while employing significantly fewer parameters. Furthermore, we provide a comprehensive analysis of the various components within our methodology to substantiate their efficacy.
Zhenwen Liang, Wenhao Yu 0002, Tanmay Rajpurohit, Peter Clark, Xiangliang Zhang 0001, Ashwin Kalyan
EMNLP2
2023 Pre-training Language Models for Comparative Reasoning
abstract
Comparative reasoning is a process of comparing objects, concepts, or entities to draw conclusions, which constitutes a fundamental cognitive ability.In this paper, we propose a novel framework to pre-train language models for enhancing their abilities of comparative reasoning over texts.While there have been approaches for NLP tasks that require comparative reasoning, they suffer from costly manual data labeling and limited generalizability to different tasks.Our approach introduces a novel method of collecting scalable data for text-based entity comparison, which leverages both structured and unstructured data.Moreover, we present a framework of pre-training language models via three novel objectives on comparative reasoning.Evaluation on downstream tasks including comparative question answering, question generation, and summarization shows that our pre-training framework significantly improves the comparative reasoning abilities of language models, especially under low-resource conditions.This work also releases the first integrated benchmark for comparative reasoning.
Mengxia Yu, Zhihan Zhang 0001, Wenhao Yu 0002, Meng Jiang 0001
EMNLP3
2023 Generate rather than Retrieve: Large Language Models are Strong Context Generators
Wenhao Yu 0002, Dan Iter, Shuohang Wang, Yichong Xu, Mingxuan Ju, Soumya Sanyal 0001, Chenguang Zhu 0001, Michael Zeng 0001, Meng Jiang 0001
ICLR1
2023 Multi-task Self-supervised Graph Neural Networks Enable Stronger Task Generalization
Mingxuan Ju, Tong Zhao 0003, Qianlong Wen, Wenhao Yu 0002, Neil Shah, Yanfang Ye 0001, Chuxu Zhang
ICLR4
2023 The Second Workshop on Knowledge-Augmented Methods for Natural Language Processing
abstract
Language models are being developed and deployed in many applications, "small"-scale and large-scale, generic and specialized, text-only and multimodal, etc. Meanwhile, the missingness of important knowledge causes limitations and safety challenges. The knowledge includes commonsense, world facts, domain expertise, personalization, and especially the unique patterns that need to be discovered from big data applications. Training and inference processes of the language models can be and should be augmented with the knowledge. The first KnowledgeNLP at AAAI 2023 attracted scientists on knowledge augmentation methods towards higher language intelligence. This workshop offers a broad platform to share ideas and discuss various topics, such as (1) synergy between knowledge and language model, (2) scalable architectures that integrate NLP, knowledge graph, and graph learning technologies, (3) KnowledgeNLP for e-commerce, education, and healthcare, (4) human factors and social good in KnowledgeNLP.
Wenhao Yu 0002, Lingbo Tong, Nanyun Peng 0001, Meng Jiang 0001
KDD1
2023 GraphPatcher: Mitigating Degree Bias for Graph Neural Networks via Test-time Augmentation
abstract
Recent studies have shown that graph neural networks (GNNs) exhibit strong biases towards the node degree: they usually perform satisfactorily on high-degree nodes with rich neighbor information but struggle with low-degree nodes. Existing works tackle this problem by deriving either designated GNN architectures or training strategies specifically for low-degree nodes. Though effective, these approaches unintentionally create an artificial out-of-distribution scenario, where models mainly or even only observe low-degree nodes during the training, leading to a downgraded performance for high-degree nodes that GNNs originally perform well at. In light of this, we propose a test-time augmentation framework, namely GraphPatcher, to enhance test-time generalization of any GNNs on low-degree nodes. Specifically, GraphPatcher iteratively generates virtual nodes to patch artificially created low-degree nodes via corruptions, aiming at progressively reconstructing target GNN's predictions over a sequence of increasingly corrupted nodes. Through this scheme, GraphPatcher not only learns how to enhance low-degree nodes (when the neighborhoods are heavily corrupted) but also preserves the original superior performance of GNNs on high-degree nodes (when lightly corrupted). Additionally, GraphPatcher is model-agnostic and can also mitigate the degree bias for either self-supervised or supervised GNNs. Comprehensive experiments are conducted over seven benchmark datasets and GraphPatcher consistently enhances common GNNs' overall performance by up to 3.6% and low-degree performance by up to 6.5%, significantly outperforming state-of-the-art baselines. The source code is publicly available at https://github.com/jumxglhf/GraphPatcher.
Mingxuan Ju, Tong Zhao 0003, Wenhao Yu 0002, Neil Shah, Yanfang Ye 0001
NeurIPS3
2023 Knowledge-Augmented Methods for Natural Language Processing
abstract
Knowledge in NLP has been a rising trend especially after the advent of large-scale pre-trained models. Knowledge is critical to equip statistics-based models with common sense, logic and other external information. In this tutorial, we will introduce recent state-of-the-art works in applying knowledge in language understanding, language generation and commonsense reasoning.
Chenguang Zhu 0001, Yichong Xu, Xiang Ren 0001, Bill Y. Lin, Meng Jiang 0001, Wenhao Yu 0002
WSDM6
2023 Exploring Contrast Consistency of Open-Domain Question Answering Systems on Minimally Edited Questions
abstract
Abstract Contrast consistency, the ability of a model to make consistently correct predictions in the presence of perturbations, is an essential aspect in NLP. While studied in tasks such as sentiment analysis and reading comprehension, it remains unexplored in open-domain question answering (OpenQA) due to the difficulty of collecting perturbed questions that satisfy factuality requirements. In this work, we collect minimally edited questions as challenging contrast sets to evaluate OpenQA models. Our collection approach combines both human annotation and large language model generation. We find that the widely used dense passage retriever (DPR) performs poorly on our contrast sets, despite fitting the training set well and performing competitively on standard test sets. To address this issue, we introduce a simple and effective query-side contrastive loss with the aid of data augmentation to improve DPR training. Our experiments on the contrast sets demonstrate that DPR’s contrast consistency is improved without sacrificing its accuracy on the standard test sets.1
Zhihan Zhang 0001, Wenhao Yu 0002, Zheng Ning, Mingxuan Ju, Meng Jiang 0001
Trans. Assoc. Comput. Linguistics2
2023 Deep Multimodal Complementarity Learning
abstract
Complementarity plays a significant role in the synergistic effect created by different components of a complex data object. Complementarity learning on multimodal data has fundamental challenges of representation learning because the complementarity exists along with multiple modalities and one or multiple items of each modality. Also, an appropriate metric is needed for measuring the complementarity in the representation space. Existing methods that rely on similarity-based metrics cannot adequately capture the complementarity. In this work, we propose a novel deep architecture for systematically learning the complementarity of components from multimodal multi-item data. The proposed model consists of three major modules: 1) unimodal aggregation for extracting the intramodal complementarity; 2) cross-modal fusion for extracting the intermodal complementarity at the modality level; and 3) interactive aggregation for extracting the intermodal complementarity at the item level. To quantify complementarity, we utilize the TUBE distance metric to measure the difference between the composited data object and its label in the representation space. Experiments on three real datasets show that our model outperforms the state-of-the-art by +6.8% of mean reciprocal rank (MRR) on object classification and +3.0% of MRR on hold-out item prediction. Qualitative analyses reveal that complementarity is significantly different from similarity.
Daheng Wang, Tong Zhao 0003, Wenhao Yu 0002, Nitesh V. Chawla, Meng Jiang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2022 KG-FiD: Infusing Knowledge Graph in Fusion-in-Decoder for Open-Domain Question Answering
abstract
Donghan Yu, Chenguang Zhu, Yuwei Fang, Wenhao Yu, Shuohang Wang, Yichong Xu, Xiang Ren, Yiming Yang, Michael Zeng. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Donghan Yu, Chenguang Zhu 0001, Yuwei Fang, Wenhao Yu 0002, Shuohang Wang, Yichong Xu, Xiang Ren 0001, Yiming Yang 0002, Michael Zeng 0001
ACL (1)4
2022 Retrieval Augmentation for Commonsense Reasoning: A Unified Approach
abstract
A common thread of retrieval-augmented methods in the existing literature focuses on retrieving encyclopedic knowledge, such as Wikipedia, which facilitates well-defined entity and relation spaces that can be modeled.However, applying such methods to commonsense reasoning tasks faces two unique challenges, i.e., the lack of a general large-scale corpus for retrieval and a corresponding effective commonsense retriever.In this paper, we systematically investigate how to leverage commonsense knowledge retrieval to improve commonsense reasoning tasks.We proposed a unified framework of Retrieval-Augmented Commonsense reasoning (called RACO), including a newly constructed commonsense corpus with over 20 million documents and novel strategies for training a commonsense retriever.We conducted experiments on four different commonsense reasoning tasks.Extensive evaluation results showed that our proposed RACO can significantly outperform other knowledgeenhanced method counterparts, achieving new SoTA performance on the CommonGen 1 and CREAK 2 leaderboards.Our code is available at https://github.com/wyu97/RACo.
Wenhao Yu 0002, Chenguang Zhu 0001, Zhihan Zhang 0001, Shuohang Wang, Zhuosheng Zhang 0001, Yuwei Fang, Meng Jiang 0001
EMNLP1
2022 Empowering Language Models with Knowledge Graph Reasoning for Open-Domain Question Answering
abstract
Answering open-domain questions requires world knowledge about in-context entities.As pre-trained Language Models (LMs) lack the power to store all required knowledge, external knowledge sources, such as knowledge graphs, are often used to augment LMs.In this work, we propose knOwledge REasOning empowered Language Model (OREOLM), which consists of a novel Knowledge Interaction Layer that can be flexibly plugged into existing Transformer-based LMs to interact with a differentiable Knowledge Graph Reasoning module collaboratively.In this way, LM guides KG to walk towards the desired answer, while the retrieved knowledge improves LM.By adopting OREOLM to RoBERTa and T5, we show significant performance gain, achieving state-of-art results in the Closed-Book setting.The performance enhancement is mainly from the KG reasoning's capacity to infer missing relational facts.In addition, OREOLM provides reasoning paths as rationales to interpret the model's decision.
Ziniu Hu, Yichong Xu, Wenhao Yu 0002, Shuohang Wang, Ziyi Yang 0011, Chenguang Zhu 0001, Kai-Wei Chang 0001, Yizhou Sun
EMNLP3
2022 A Unified Encoder-Decoder Framework with Entity Memory
abstract
Entities, as important carriers of real-world knowledge, play a key role in many NLP tasks.We focus on incorporating entity knowledge into an encoder-decoder framework for informative text generation.Existing approaches tried to index, retrieve, and read external documents as evidence, but they suffered from a large computational overhead.In this work, we propose an Encoder-Decoder framework with an entity Memory, namely EDMem.The entity knowledge is stored in the memory as latent representations, and the memory is pre-trained on Wikipedia along with encoder-decoder parameters.To precisely generate entity names, we design three decoding methods to constrain entity generation by linking entities in the memory.EDMem is a unified framework that can be used on various entity-intensive question answering and generation tasks.Extensive experimental results show that EDMem outperforms both memory-based auto-encoder models and non-memory encoder-decoder models.
Zhihan Zhang 0001, Wenhao Yu 0002, Chenguang Zhu 0001, Meng Jiang 0001
EMNLP2
2022 Learning from Counterfactual Links for Link Prediction
abstract
Learning to predict missing links is important for many graph-based applications. Existing methods were designed to learn the association between observed graph structure and existence of link between a pair of nodes. However, the causal relationship between the two variables was largely ignored for learning to predict links on a graph. In this work, we visit this factor by asking a counterfactual question: "would the link still exist if the graph structure became different from observation?" Its answer, counterfactual links, will be able to augment the graph data for representation learning. To create these links, we employ causal models that consider the information (i.e., learned representations) of node pairs as context, global graph structural properties as treatment, and link existence as outcome. We propose a novel data augmentation-based link prediction method that creates counterfactual links and learns representations from both the observed and counterfactual links. Experiments on benchmark data show that our graph learning method achieves state-of-the-art performance on the task of link prediction.
Tong Zhao 0003, Gang Liu 0025, Daheng Wang, Wenhao Yu 0002, Meng Jiang 0001
ICML4
2021 Action Sequence Augmentation for Early Graph-based Anomaly Detection
abstract
The proliferation of web platforms has created incentives for online abuse. Many graph-based anomaly detection techniques are proposed to identify the suspicious accounts and behaviors. However, most of them detect the anomalies once the users have performed many such behaviors. Their performance is substantially hindered when the users' observed data is limited at an early stage, which needs to be improved to minimize financial loss. In this work, we propose Eland, a novel framework that uses action sequence augmentation for early anomaly detection. Eland utilizes a sequence predictor to predict next actions of every user and exploits the mutual enhancement between action sequence augmentation and user-action graph anomaly detection. Experiments on three real-world datasets show that Eland improves the performance of a variety of graph-based anomaly detection methods. With Eland, anomaly detection performance at an earlier stage is better than non-augmented methods that need significantly more observed data by up to 15% on the Area under the ROC curve.
Tong Zhao 0003, Bo Ni, Wenhao Yu 0002, Zhichun Guo, Neil Shah, Meng Jiang 0001
CIKM3
2021 Injecting Entity Types into Entity-Guided Text Generation
abstract
Recent successes in deep generative modeling have led to significant advances in natural language generation (NLG).Incorporating entities into neural generation models has demonstrated great improvements by assisting to infer the summary topic and to generate coherent content.To enhance the role of entity in NLG, in this paper, we aim to model the entity type in the decoding phase to generate contextual words accurately.We develop a novel NLG model to produce a target sequence based on a given list of entities.Our model has a multistep decoder that injects the entity types into the process of entity mention generation.Experiments on two public news datasets demonstrate type injection performs better than existing type embedding concatenation baselines.
Xiangyu Dong 0004, Wenhao Yu 0002, Chenguang Zhu 0001, Meng Jiang 0001
EMNLP (1)2
2021 Sentence-Permuted Paragraph Generation
abstract
Generating paragraphs of diverse contents is important in many applications.Existing generation models produce similar contents from homogenized contexts due to the fixed left-toright sentence order.Our idea is permuting the sentence orders to improve the content diversity of multi-sentence paragraph.We propose a novel framework PermGen whose objective is to maximize the expected log-likelihood of output paragraph distributions with respect to all possible sentence orders.PermGen uses hierarchical positional embedding and designs new procedures for both training phase and inference phase.Experiments on three paragraph generation benchmarks demonstrate Per-mGen generates more diverse outputs with a higher quality than existing models.
Wenhao Yu 0002, Chenguang Zhu 0001, Tong Zhao 0003, Zhichun Guo, Meng Jiang 0001
EMNLP (1)1
2021 Enhancing Taxonomy Completion with Concept Generation via Fusing Relational Representations
abstract
Automatic construction of a taxonomy supports many applications in e-commerce, web search, and question answering. Existing taxonomy expansion or completion methods assume that new concepts have been accurately extracted and their embedding vectors learned from the text corpus. However, one critical and fundamental challenge in fixing the incompleteness of taxonomies is the incompleteness of the extracted concepts, especially for those whose names have multiple words and consequently low frequency in the corpus. To resolve the limitations of extraction-based methods, we propose GenTaxo to enhance taxonomy completion by identifying positions in existing taxonomies that need new concepts and then generating appropriate concept names. Instead of relying on the corpus for concept embeddings, GenTaxo learns the contextual embeddings from their surrounding graph-based and language-based relational information, and leverages the corpus for pre-training a concept name generator. Experimental results demonstrate that GenTaxo improves the completeness of taxonomies over existing methods.
Qingkai Zeng 0001, Jinfeng Lin, Wenhao Yu 0002, Jane Cleland-Huang, Meng Jiang 0001
KDD3
2021 Few-Shot Graph Learning for Molecular Property Prediction
abstract
The recent success of graph neural networks has significantly boosted molecular property prediction, advancing activities such as drug discovery. The existing deep neural network methods usually require large training dataset for each property, impairing their performance in cases (especially for new molecular properties) with a limited amount of experimental data, which are common in real situations. To this end, we propose Meta-MGNN, a novel model for few-shot molecular property prediction. Meta-MGNN applies molecular graph neural network to learn molecular representations and builds a meta-learning framework for model optimization. To exploit unlabeled molecular information and address task heterogeneity of different molecular properties, Meta-MGNN further incorporates molecular structures, attribute based self-supervised modules and self-attentive task weights into the former framework, strengthening the whole learning model. Extensive experiments on two public multi-property datasets demonstrate that Meta-MGNN outperforms a variety of state-of-the-art methods.
Zhichun Guo, Chuxu Zhang, Wenhao Yu 0002, John Herr, Olaf Wiest, Meng Jiang 0001, Nitesh V. Chawla
WWW3
2020 Crossing Variational Autoencoders for Answer Retrieval
abstract
Answer retrieval is to find the most aligned answer from a large set of candidates given a question.Learning vector representations of questions/answers is the key factor.Questionanswer alignment and question/answer semantics are two important signals for learning the representations.Existing methods learned semantic representations with dual encoders or dual variational auto-encoders.The semantic information was learned from language models or question-to-question (answer-to-answer) generative processes.However, the alignment and semantics were too separate to capture the aligned semantics between question and answer.In this work, we propose to cross variational auto-encoders by generating questions with aligned answers and generating answers with aligned questions.Experiments show that our method outperforms the state-of-theart answer retrieval method on SQuAD.Question Answer Decoder 𝑝(𝑞|𝒛 𝒂 ) 𝑝(𝑎|𝒛 𝒒 ) 𝑝(𝑦|𝑧 !, 𝑧 " ) 𝑝(𝑦|𝑧 !, 𝑧 " ) Question Answer Question Answer Decoder Encoder Encoder Decoder Decoder Encoder 𝑝(𝑧 !|𝑞) 𝑝(𝑧 " |𝑎) Encoder Encoder (a) Dual-Encoders (Yang et al., 2019)Question Answer Decoder 𝑝(𝑞|𝒛 𝒂 ) 𝑝(𝑎|𝒛 𝒒 ) 𝑝(𝑦|𝑧 !, 𝑧 " ) 𝑝(𝑦|𝑧 !, 𝑧 " ) Question Answer Question Answer Decoder Encoder Encoder Decoder Decoder Encoder 𝑝(𝑧 !|𝑞) 𝑝(𝑧 " |𝑎) Encoder Encoder (b) Dual-VAEs (Shen et al., 2018) 𝑧 !~𝑝(𝑧 ! ) 𝑧 " ~𝑝(𝑧 " ) Question Answer 𝑧 !~𝑝(𝑧 ! ) 𝑝(𝑧 !|𝑞) 𝑧 " ~𝑝(𝑧 " ) 𝑝(𝑧 " |𝑎) Question Answer 𝑝(𝑦|𝑧 !, 𝑧 " ) 𝑝(𝑞|𝑧 ! ) 𝑝(𝑎|𝑧 " )
Wenhao Yu 0002, Lingfei Wu 0001, Qingkai Zeng 0001, Shu Tao, Yu Deng 0004, Meng Jiang 0001
ACL1
2020 GraSeq: Graph and Sequence Fusion Learning for Molecular Property Prediction
abstract
With the recent advancement of deep learning, molecular representation learning -- automating the discovery of feature representation of molecular structure, has attracted significant attention from both chemists and machine learning researchers. Deep learning can facilitate a variety of downstream applications, including bio-property prediction, chemical reaction prediction, etc. Despite the fact that current SMILES string or molecular graph molecular representation learning algorithms (via sequence modeling and graph neural networks, respectively) have achieved promising results, there is no work to integrate the capabilities of both approaches in preserving molecular characteristics (e.g, atomic cluster, chemical bond) for further improvement. In this paper, we propose GraSeq, a joint graph and sequence representation learning model for molecular property prediction. Specifically, GraSeq makes a complementary combination of graph neural networks and recurrent neural networks for modeling two types of molecular inputs, respectively. In addition, it is trained by the multitask loss of unsupervised reconstruction and various downstream tasks, using limited size of labeled datasets. In a variety of chemical property prediction tests, we demonstrate that our GraSeq model achieves better performance than state-of-the-art approaches.
Zhichun Guo, Wenhao Yu 0002, Chuxu Zhang, Meng Jiang 0001, Nitesh V. Chawla
CIKM2
2020 Experimental Evidence Extraction System in Data Science with Hybrid Table Features and Ensemble Learning
abstract
Data Science has been one of the most popular fields in higher education and research activities. It takes tons of time to read the experimental section of thousands of papers and figure out the performance of the data science techniques. In this work, we build an experimental evidence extraction system to automate the integration of tables (in the paper PDFs) into a database of experimental results. First, it crops the tables and recognizes the templates. Second, it classifies the column names and row names into “method”, “dataset”, or “evaluation metric”, and then unified all the table cells into (method, dataset, metric, score)-quadruples. We propose hybrid features including structural and semantic table features as well as an ensemble learning approach for column/row name classification and table unification. SQL statements can be used to answer questions such as whether a method is the state-of-the-art or whether the reported numbers are conflicting.
Wenhao Yu 0002, Yu Shu, Qingkai Zeng 0001, Meng Jiang 0001
WWW1
2020 Identifying Referential Intention with Heterogeneous Contexts
abstract
Citing, quoting, and forwarding & commenting behaviors are widely seen in academia, news media, and social media. Existing behavior modeling approaches focused on mining content and describing preferences of authors, speakers, and users. However, behavioral intention plays an important role in generating content on the platforms. In this work, we propose to identify the referential intention which motivates the action of using the referred (e.g., cited, quoted, and retweeted) source and content to support their claims. We adopt a theory in sociology to develop a schema of four types of intentions. The challenge lies in the heterogeneity of observed contextual information surrounding the referential behavior, such as referred content (e.g., a cited paper), local context (e.g., the sentence citing the paper), neighboring context (e.g., the former and latter sentences), and network context (e.g., the academic network of authors, affiliations, and keywords). We propose a new neural framework with Interactive Hierarchical Attention (IHA) to identify the intention of referential behavior by properly aggregating the heterogeneous contexts. Experiments demonstrate that the proposed method can effectively identify the type of intention of citing behaviors (on academic data) and retweeting behaviors (on Twitter). And learning the heterogeneous contexts collectively can improve the performance. This work opens a door for understanding content generation from a fundamental perspective of behavior sciences.
Wenhao Yu 0002, Mengxia Yu, Tong Zhao 0003, Meng Jiang 0001
WWW1
2019 Tablepedia: Automating PDF Table Reading in an Experimental Evidence Exploration and Analytic System
abstract
Web research, data science, and artificial intelligence have been rapidly changing our life and society. Researchers and practitioners in the fields take a large amount of time to read literature and compare existing approaches. It would significantly improve their efficiency if there was a system that extracted and managed experimental evidences (say, a specific method achieves a score of a specific metric on a specific dataset) from tables of paper PDFs for search, exploration, and analytic. We build such a demonstration system, called Tablepedia, that use rule-based and learning-based methods to automate the “reading” of PDF tables. It has three modules: template recognition, unification, and SQL operations. We implement three functions to facilitate research and practice: (1) finding related methods and datasets, (2) finding top-performing baseline methods, and (3) finding conflicting reported numbers. A pointer to a screencast on Vimeo: https://vimeo.com/310162310
Wenhao Yu 0002, Qingkai Zeng 0001, Meng Jiang 0001
WWW1