VLDB 2026 Research / reviewers in the wild / expert
Xiang Ren 0001
dblp:36/360-1 · also Sean (Xiang) Ren
· DBLP profile ↗
165ranked-venue papers
12as first author
89since 2021 · last 2026
0000-0001-8655-663XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 145 · 9 first-author · 86 since 2021Databases, data management, data science and information retrieval · 38 · 10 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DRInQ: Evaluating Conversational Implicature with Controlled Context VariationabstractHuman conversation relies heavily on conversational implicature, in which speakers convey meanings that are suggested rather than explicitly stated.Although recent large language models (LLMs) exhibit strong conversational fluency, they remain unreliable when interpretation depends on reasoning that integrates social and contextual cues, a process rarely articulated in text.We introduce DRinQ, a benchmark for evaluating pragmatic reasoning about conversational implicature in question utterances, designed to isolate pragmatic variation while holding each question's surface form fixed.To support scalable evaluation, we propose a semi-automated pipeline that produces question-context-interpretation instances with systematic variation.Across evaluations, we find a consistent generation-inference asymmetry: while state-of-the-art models can generate plausible pragmatic scenarios when guided, they often fail to recover the intended implication at inference time.For smaller models, structured prompting improves alignment with human judgments.A comparative writing study further reveals complementary strengths: human authors tend to produce safer, predictable contexts, whereas models generate varied scenarios with interpretations that sometimes exceed contextual support.These findings highlight persistent challenges in modeling conversational implicature and motivate more contextsensitive evaluation frameworks. Hirona Jacqueline Arai, Xiang Ren 0001 |
ACL (1) | 2 |
| 2026 | Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model ExplanationsabstractKeyu He, Tejas Srinivasan, Brihi Joshi, Xiang Ren, Jesse Thomason, Swabha Swayamdipta. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Keyu He, Tejas Srinivasan, Brihi Joshi, Xiang Ren 0001, Jesse Thomason, Swabha Swayamdipta |
ACL (1) | 4 |
| 2025 | Diagnosing Memorization in Chain-of-Thought Reasoning, One Token at a TimeabstractLarge Language Models (LLMs) perform well on reasoning benchmarks but often fail when inputs alter slightly, raising concerns about the extent to which their success relies on memorization.This issue is especially acute in Chain-of-Thought (CoT) reasoning, where spurious memorized patterns can trigger intermediate errors that cascade into incorrect final answers.We introduce STIM, a novel framework for Source-aware Token-level Identification of Memorization, which attributes each token in a reasoning chain to one of multiple memorization sources -local, mid-range, or long-rangebased on their statistical co-occurrence with the token in the pretraining corpus.Our token-level analysis across tasks and distributional settings reveals that models rely more on memorization in complex or long-tail cases, and that local memorization is often the dominant driver of errors, leading to up to 67% of wrong tokens.We also show that memorization scores from STIM can be effective in predicting the wrong tokens in the wrong reasoning step.STIM offers a powerful tool for diagnosing and improving model reasoning and can generalize to other structured step-wise generation tasks. 1 * Equal Contribution † Work done in Amazon AGI 1 https://github.com/INK-USC/STIMModel Reasoning Steps: Rick killed 10 + 15 = 25 animals last night.Today Rick killed 3 * 3 = 9 wolves.Today Rick killed 15 -3 = 12 cougars.Today Rick killed 9 + 12 = 21 animals.Total animals killed by Rick is 25 + 21 = 46.So the answer is 46. Examples of Simultaneous Memorization Sources of Different StrengthQuestion: Here is a list: [pitanga, pitanga, yumberry, yumberry, pitanga, yumberry, pitanga, yumberry, pitanga, pitanga].How many times does 'pitanga' appear on it?Model Output: 1.The first element is 'pitanga'.... 10.The tenth element is 'pitanga'.Now, counting the occurrences of 'pitanga': We have 'pitanga' appearing 5 times at positions 1, 2, 5, 7, and 10.So the answer is 5.Analysis: Token 5 is incorrect and has the highest influence from local context, where the n-gram "appearing 5" has a high pretraining frequency.STIM scores are: local = 0.81 , mid = 0.24, long = 0.03, suggesting that the model was influenced by locally frequent patterns rather than actual counting.Question: What is ( 100101100 -100101100 * 0.1) * 111110100 + 100101100 * 1111 / 1100100 * 111110100 * 0.1 + (100101100 -100101100 * 0.1 -100101100 * 1111 / 1100100) * 111110100 equal to? (base-2) Model Output:To solve this, we first handle the multiplication and division in base-2, converting to base-10 as necessary:1. (100101100 × 0.1base-2 = 10010110base-2) (which is 180 in base-10) ... So the answer is 26807.536.Analysis: Token 180 is incorrect and has the highest influence from input tokens, having higher pretraining co-occurrence frequency with "111110100", "*" and "100101100" in the input.STIM scores are local=-0.19,mid=0.09,long=0.156, indicating long-range memorization being the primary influence. Huihan Li 0001, Ninareh Mehrabi, Rahul Gupta 0001, Xiang Ren 0001 |
EMNLP | 7 |
| 2025 | Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM JuriesabstractToday, large language models are widely used as judges to evaluate responses from other language models.Hence, it is imperative to benchmark and improve these LLM-judges on realworld language model usage: a typical humanassistant conversation is lengthy, and shows significant diversity in topics, intents, and requirements across turns, e.g.social interactions, task requests, feedback.We present AMULET, a framework that leverages pertinent linguistic concepts of dialog-acts and maxims to improve the accuracy of LLM-judges on preference data with complex, multi-turn conversational context.AMULET presents valuable insights about (a) the communicative structures and intents present in the conversation (dialog acts), and (b) the satisfaction of conversational principles (maxims) by the preference pair responses, and uses them to make judgments.On 4 challenging datasets, AMULET shows that (a) humans frequently (60-70% of the time) change their intents from one turn of the conversation to the next, and (b) in ∼75% of instances, the preference pair responses can be differentiated via dialog acts and/or maxims, reiterating the latter's significance in judging such data.AMULET can be used either as a judge by applying the framework to a single LLM, or integrated into a jury with different LLM judges; our judges and juries show strong improvements on relevant baselines for all 4 datasets.(code, data). Sahana Ramnath, Anurag Mudgil, Brihi Joshi, Skyler Hallinan, Xiang Ren 0001 |
EMNLP | 5 |
| 2025 | Mixing Inference-time Experts for Enhancing LLM ReasoningabstractLarge Language Models (LLMs) have demonstrated impressive reasoning abilities, but their generated rationales often suffer from issues such as reasoning inconsistency and factual errors, undermining their reliability.Prior work has explored improving rationale quality via multi-reward fine-tuning or reinforcement learning (RL), where models are optimized for diverse objectives.While effective, these approaches train the model in a fixed manner and do not have any inference-time adaptability, nor can they generalize reasoning requirements for new test-time inputs.Another approach is to train specialized reasoning experts using reward signals and use them to improve generation at inference time.Existing methods in this paradigm are limited to using only a single expert and cannot improve upon multiple reasoning aspects.To address this, we propose MIXIE, a novel inference-time expertmixing framework that dynamically determines mixing proportions for each expert, enabling contextualized and flexible fusion.We demonstrate the effectiveness of MIXIE on improving chain-of-thought reasoning in LLMs by merging commonsense and entailment reasoning experts finetuned on reward-filtered data.Our approach outperforms existing baselines on three question-answering datasets: StrategyQA, CommonsenseQA, and ARC, highlighting its potential to enhance LLM reasoning with efficient, adaptable expert integration. Soumya Sanyal 0001, Tianyi Xiao, Xiang Ren 0001 |
EMNLP | 3 |
| 2025 | Stepwise Informativeness Search for Improving LLM ReasoningabstractAdvances in Large Language Models (LLMs) have improved multi-step reasoning by generating free-text rationales, but these models tend to lose focus over the middle of long contexts.This raises concerns that as reasoning progresses, LLMs may overlook information in earlier steps when decoding subsequent steps, leading to unreliable and redundant rationales.To address this, we propose guiding LLMs to generate more accurate and concise rationales by (1) proactively referencing information from underutilized prior steps, and (2) minimizing redundant information between new and existing steps.We introduce stepwise informativeness search, an inference-time tree search framework incorporating two selection heuristics: grounding-guided selection which prioritizes steps paying higher attention over underutilized steps; and novelty-guided selection which encourages steps with novel conclusions.We further utilize a self-grounding strategy that prompts LLMs to explicitly reference relevant prior steps as premises before deduction at each step, mitigating distraction from irrelevant content.Experiments on five reasoning datasets across five LLMs show the effectiveness and efficiency of our approach to improve reasoning with reduced errors and redundancy 1 . Enda Zhao, Xiang Ren 0001 |
EMNLP | 3 |
| 2025 | Rethinking Backdoor Detection Evaluation for Language ModelsabstractBackdoor attacks, in which a model behaves maliciously when given an attacker-specified trigger, pose a major security risk for practitioners who depend on publicly released language models.As a countermeasure, backdoor detection methods aim to detect whether a released model contains a backdoor.While existing backdoor detection methods have high accuracy in detecting backdoored models on standard benchmarks, it is unclear whether they can robustly identify backdoors in the wild.In this paper, we examine the robustness of backdoor detectors by manipulating different factors during backdoor planting.We find that the success of existing methods based on trigger inversion or meta classifiers highly depends on how intensely the model is trained on poisoned data.Specifically, backdoors planted with more aggressive or more conservative training are significantly more difficult to detect than the default ones.Our results highlight a lack of robustness of existing backdoor detectors and the limitations in current benchmark construction. Jun Yan 0012, Wenjie Mo 0001, Xiang Ren 0001, Robin Jia |
EMNLP | 3 |
| 2025 | Attributing Culture-Conditioned Generations to Pretraining CorporaabstractIn open-ended generative tasks like narrative writing or dialogue, large language models often exhibit cultural biases, showing limited knowledge and generating templated outputs for less prevalent cultures. Recent works show that these biases may stem from uneven cultural representation in pretraining corpora. This work investigates how pretraining leads to biased culture-conditioned generations
by analyzing how models associate entities with cultures based on pretraining data patterns. We propose the MEMOED framework (MEMOrization from prEtraining Document) to determine whether a generation for a culture arises from memorization. Using MEMOED on culture-conditioned generations about food and clothing for 110 cultures, we find that high-frequency cultures in pretraining data yield more generations with memorized symbols, while some low-frequency cultures produce none. Additionally, the model favors generating entities with extraordinarily high frequency regardless of the conditioned culture, reflecting biases toward frequent pretraining terms irrespective of relevance. We hope that the MEMOED framework and our insights will inspire more works on attributing model performance on pretraining data. Huihan Li 0001, Arnav Goel, Keyu He, Xiang Ren 0001 |
ICLR | 4 |
| 2025 | Diverging Preferences: When do Annotators Disagree and do Models Know?abstractWe examine diverging preferences in human-labeled preference datasets. We develop a taxonomy of disagreement sources spanning ten categories across four high-level classes and find that the majority of disagreements are due to factors such as task underspecification or response style. Our findings challenge a standard assumption in reward modeling methods that annotator disagreements can be attributed to simple noise. We then explore how these findings impact two areas of LLM development: reward modeling training and evaluation. In our experiments, we demonstrate how standard reward modeling (e.g., Bradley-Terry) and LLM-as-Judge evaluation methods fail to account for divergence between annotators. These findings highlight challenges in LLM evaluations, which are greatly influenced by divisive features like response style, and in developing pluralistically aligned LLMs. To address these issues, we develop methods for identifying diverging preferences to mitigate their influence in evaluations and during LLM training. Michael J. Q. Zhang, Zhilin Wang, Jena D. Hwang, Yi Dong 0003, Olivier Delalleau, Yejin Choi 0001, Eunsol Choi, Xiang Ren 0001, Valentina Pyatkin |
ICML | 8 |
| 2025 | CAVE: Controllable Authorship Verification ExplanationsabstractSahana Ramnath, Kartik Pandey, Elizabeth Boschee, Xiang Ren. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sahana Ramnath, Kartik Pandey, Elizabeth Boschee, Xiang Ren 0001 |
NAACL (Long Papers) | 4 |
| 2025 | REL-A.I.: An Interaction-Centered Approach To Measuring Human-LM RelianceabstractKaitlyn Zhou, Jena D. Hwang, Xiang Ren, Nouha Dziri, Dan Jurafsky, Maarten Sap. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kaitlyn Zhou, Jena D. Hwang, Xiang Ren 0001, Nouha Dziri, Daniel Jurafsky, Maarten Sap |
NAACL (Long Papers) | 3 |
| 2025 | Demystifying Language Model Forgetting with Low-rank Example AssociationsabstractLarge Language models (LLMs) suffer from forgetting of upstream knowledge when fine-tuned. Despite efforts on mitigating forgetting, few have investigated how forgotten upstream examples are dependent on newly learned tasks. Insights on such dependencies enable efficient and targeted mitigation of forgetting. In this paper, we empirically analyze forgetting that occurs in $N$ upstream examples of language modeling or instruction-tuning after fine-tuning LLMs on one of $M$ new tasks, visualized in $M\times N$ matrices. We show that the matrices are often well-approximated with low-rank matrices, indicating the dominance of simple associations between the learned tasks and forgotten upstream examples. Leveraging the analysis, we predict forgetting of upstream examples when fine-tuning LLMs on unseen tasks with matrix completion over the empirical associations. This enables fast identification of most forgotten examples without expensive inference on the entire upstream data. Despite simplicity, the approach outperforms prior approaches that learn semantic relationships of learned tasks and upstream examples with LMs. We demonstrate the practical utility of our analysis by showing statistically significantly reduced forgetting as we upweight predicted examples for replay during fine-tuning. Xisen Jin, Xiang Ren 0001 |
NeurIPS | 2 |
| 2025 | Better Language Model Inversion by Compactly Representing Next-Token DistributionsabstractLanguage model inversion seeks to recover hidden prompts using only language model outputs. This capability has implications for security and accountability in language model deployments, such as leaking private information from an API-protected language model’s system message. We propose a new method – prompt inversion from logprob sequences (PILS) – that recovers hidden prompts by gleaning clues from the model’s next-token probabilities over the course of multiple generation steps. Our method is enabled by a key insight: The vector-valued outputs of a language model occupy a low-dimensional subspace. This enables us to losslessly compress the full next-token probability distribution over multiple generation steps using a linear map, allowing more output information to be used for inversion. Our approach yields massive gains over previous state-of-the-art methods for recovering hidden prompts, achieving 2–3.5 times higher exact recovery rates across test sets, in one case increasing the recovery rate from 17% to 60%. Our method also exhibits surprisingly good generalization behavior; for instance, an inverter trained on 16 generations steps gets 5–27% higher prompt recovery when we increase the number of steps to 32 at test time. Furthermore, we demonstrate strong performance of our method on the more challenging task of recovering hidden system messages. We also analyze the role of verbatim repetition in prompt recovery and propose a new method for cross-family model transfer for logit-based inverters. Our findings suggest that next-token probabilities are a considerably more vulnerable attack surface for inversion attacks than previously known. Murtaza Nazir, Matthew Finlayson, John X. Morris, Xiang Ren 0001, Swabha Swayamdipta |
NeurIPS | 4 |
| 2024 | Can LLMs Reason with Rules? Logic Scaffolding for Stress-Testing and Improving LLMsabstractLarge language models (LLMs) have achieved impressive human-like performance across various reasoning tasks.However, their mastery of underlying inferential rules still falls short of human capabilities.To investigate this, we propose a logic scaffolding inferential rule generation framework, to construct an inferential rule base, ULogic, comprising both primitive and compositional rules across five domains.Our analysis of GPT-series models over a rule subset reveals significant gaps in LLMs' logic understanding compared to human performance, especially in compositional and structural complex rules with certain bias patterns.We further distill these rules into a smaller-scale inference engine for flexible rule generation and enhancing downstream reasoning.Through a multijudger evaluation, our inference engine proves effective in generating accurate, complex and abstract conclusions and premises, and improve various commonsense reasoning tasks.Overall, our work sheds light on LLMs' limitations in grasping inferential rule and suggests ways to enhance their logical reasoning abilities 1 . Zhongyu Wei, Yejin Choi 0001, Xiang Ren 0001 |
ACL (1) | 4 |
| 2024 | Relying on the Unreliable: The Impact of Language Models' Reluctance to Express UncertaintyabstractAs natural language becomes the default interface for human-AI interaction, there is a need for LMs to appropriately communicate uncertainties in downstream applications.In this work, we investigate how LMs incorporate confidence in responses via natural language and how downstream users behave in response to LM-articulated uncertainties.We examine publicly deployed models and find that LMs are reluctant to express uncertainties when answering questions even when they produce incorrect responses.LMs can be explicitly prompted to express confidences, but tend to be overconfident, resulting in high error rates (an average of 47%) among confident responses.We test the risks of LM overconfidence by conducting human experiments and show that users rely heavily on LM generations, whether or not they are marked by certainty.Lastly, we investigate the preference-annotated datasets used in post training alignment and find that humans are biased against texts with uncertainty.Our work highlights new safety harms facing human-LM interactions and proposes design recommendations and mitigating strategies moving forward. Kaitlyn Zhou, Jena D. Hwang, Xiang Ren 0001, Maarten Sap |
ACL (1) | 3 |
| 2024 | Resource-rational moral judgment
Sarah A. Wu, Xiang Ren 0001, Tobias Gerstenberg, Yejin Choi 0001, Sydney Levine |
CogSci | 2 |
| 2024 | In Search of the Long-Tail: Systematic Generation of Long-Tail Inferential Knowledge via Logical Rule Guided SearchabstractHuihan Li, Yuting Ning, Zeyi Liao, Siyuan Wang, Xiang Lorraine Li, Ximing Lu, Wenting Zhao, Faeze Brahman, Yejin Choi, Xiang Ren. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Huihan Li 0001, Yuting Ning, Zeyi Liao, Xiang Li 0069, Ximing Lu, Faeze Brahman, Yejin Choi 0001, Xiang Ren 0001 |
EMNLP | 10 |
| 2024 | Symbolic Working Memory Enhances Language Models for Complex Rule ApplicationabstractLarge Language Models (LLMs) have shown remarkable reasoning performance but struggle with multi-step deductive reasoning involving a series of rule application steps, especially when rules are presented non-sequentially.Our preliminary analysis shows that while LLMs excel in single-step rule application, their performance drops significantly in multi-step scenarios due to the challenge in rule grounding.It requires anchoring the applicable rule and supporting facts at each step, amidst multiple input rules, facts, and inferred facts.To address this, we propose augmenting LLMs with external working memory and introduce a neurosymbolic framework for rule application.The memory stores facts and rules in both natural language and symbolic forms, enabling precise tracking.Utilizing this memory, our framework iteratively performs symbolic rule grounding and LLM-based rule implementation.The former matches predicates and variables of symbolic rules and facts to ground applicable rules at each step.Experiments indicate our framework's effectiveness in rule application and its robustness across various steps and settings 1 . Zhongyu Wei, Yejin Choi 0001, Xiang Ren 0001 |
EMNLP | 4 |
| 2024 | PlaSma: Procedural Knowledge Models for Language-based Planning and Re-PlanningabstractProcedural planning, which entails decomposing a high-level goal into a sequence of temporally ordered steps, is an important yet intricate task for machines. It involves integrating common-sense knowledge to reason about complex and often contextualized situations, e.g. ``scheduling a doctor's appointment without a phone''. While current approaches show encouraging results using large language models (LLMs), they are hindered by drawbacks such as costly API calls and reproducibility issues. In this paper, we advocate planning using smaller language models. We present PlaSma, a novel two-pronged approach to endow small language models with procedural knowledge and (constrained) language-based planning capabilities. More concretely, we develop *symbolic procedural knowledge distillation* to enhance the commonsense knowledge in small language models and an *inference-time algorithm* to facilitate more structured and accurate reasoning. In addition, we introduce a new related task, *Replanning*, that requires a revision of a plan to cope with a constrained situation. In both the planning and replanning settings, we show that orders-of-magnitude smaller models (770M-11B parameters) can compete and often surpass their larger teacher models' capabilities. Finally, we showcase successful application of PlaSma in an embodied environment, VirtualHome. Faeze Brahman, Chandra Bhagavatula, Valentina Pyatkin, Jena D. Hwang, Xiang Li 0069, Hirona Jacqueline Arai, Soumya Sanyal 0001, Keisuke Sakaguchi, Xiang Ren 0001, Yejin Choi 0001 |
ICLR | 9 |
| 2024 | Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis RefinementabstractThe ability to derive underlying principles from a handful of observations and then generalize to novel situations---known as inductive reasoning---is central to human intelligence. Prior work suggests that language models (LMs) often fall short on inductive reasoning, despite achieving impressive success on research benchmarks. In this work, we conduct a systematic study of the inductive reasoning capabilities of LMs through $\textit{iterative hypothesis refinement}$, a technique that more closely mirrors the human inductive process than standard input-output prompting. Iterative hypothesis refinement employs a three-step process: proposing, selecting, and refining hypotheses in the form of textual rules. By examining the intermediate rules, we observe that LMs are phenomenal $\textit{hypothesis proposers}$ (i.e., generating candidate rules), and when coupled with a (task-specific) symbolic interpreter that is able to systematically filter the proposed set of rules, this hybrid approach achieves strong results across inductive reasoning benchmarks that require inducing causal relations, language-like instructions, and symbolic concepts. However, they also behave as puzzling $\textit{inductive reasoners}$, showing notable performance gaps between rule induction (i.e., identifying plausible rules) and rule application (i.e., applying proposed rules to instances), suggesting that LMs are proposing hypotheses without being able to actually apply the rules. Through empirical and human analyses, we further reveal several discrepancies between the inductive reasoning processes of LMs and humans, shedding light on both the potentials and limitations of using LMs in inductive reasoning tasks. Linlu Qiu, Ximing Lu, Melanie Sclar, Valentina Pyatkin, Chandra Bhagavatula, Bailin Wang, Yejin Choi 0001, Nouha Dziri, Xiang Ren 0001 |
ICLR | 11 |
| 2024 | Tailoring Self-Rationalizers with Multi-Reward DistillationabstractLarge language models (LMs) are capable of generating free-text rationales to aid question answering. However, prior work 1) suggests that useful self-rationalization is emergent only at significant scales (e.g., 175B parameter GPT-3); and 2) focuses largely on downstream performance, ignoring the semantics of the rationales themselves, e.g., are they faithful, true, and helpful for humans? In this work, we enable small-scale LMs (∼200x smaller than GPT-3) to generate rationales that not only improve downstream task performance, but are also more plausible, consistent, and diverse, assessed both by automatic and human evaluation. Our method, MaRio (Multi-rewArd RatIOnalization), is a multi-reward conditioned self-rationalization algorithm that optimizes multiple distinct properties like plausibility, diversity and consistency. Results on three difficult question-answering datasets StrategyQA, QuaRel and OpenBookQA show that not only does MaRio improve task accuracy, but it also improves the self-rationalization quality of small LMs across the aforementioned axes better than a supervised fine-tuning (SFT) baseline. Extensive human evaluations confirm that MaRio rationales are preferred vs. SFT rationales, as well as qualitative improvements in plausibility and consistency. Sahana Ramnath, Brihi Joshi, Skyler Hallinan, Ximing Lu, Liunian Harold Li, Aaron Chan, Jack Hessel, Yejin Choi 0001, Xiang Ren 0001 |
ICLR | 9 |
| 2024 | WildChat: 1M ChatGPT Interaction Logs in the WildabstractChatbots such as GPT-4 and ChatGPT are now serving millions of users. Despite their widespread use, there remains a lack of public datasets showcasing how these tools are used by a population of users in practice. To bridge this gap, we offered free access to ChatGPT for online users in exchange for their affirmative, consensual opt-in to anonymously collect their chat transcripts and request headers. From this, we compiled WildChat, a corpus of 1 million user-ChatGPT conversations, which consists of over 2.5 million interaction turns. We compare WildChat with other popular user-chatbot interaction datasets, and find that our dataset offers the most diverse user prompts, contains the largest number of languages, and presents the richest variety of potentially toxic use-cases for researchers to study. In addition to timestamped chat transcripts, we enrich the dataset with demographic data, including state, country, and hashed IP addresses, alongside request headers. This augmentation allows for more detailed analysis of user behaviors across different geographical regions and temporal dimensions. Finally, because it captures a broad range of use cases, we demonstrate the dataset’s potential utility in fine-tuning instruction-following models. WildChat is released at https://wildchat.allen.ai under AI2 ImpACT Licenses. Xiang Ren 0001, Jack Hessel, Claire Cardie, Yejin Choi 0001, Yuntian Deng |
ICLR | 2 |
| 2024 | What Will My Model Forget? Forecasting Forgotten Examples in Language Model RefinementabstractLanguage models deployed in the wild make errors. However, simply updating the model with the corrected error instances causes catastrophic forgetting---the updated model makes errors on instances learned during the instruction tuning or upstream training phase. Randomly replaying upstream data yields unsatisfactory performance and often comes with high variance and poor controllability. To this end, we try to forecast upstream examples that will be forgotten due to a model update for improved controllability of the replay process and interpretability. We train forecasting models given a collection of online learned examples and corresponding forgotten upstream pre-training examples. We propose a partially interpretable forecasting model based on the observation that changes in pre-softmax logit scores of pretraining examples resemble that of online learned examples, which performs decently on BART but fails on T5 models. We further show a black-box classifier based on inner products of example representations achieves better forecasting performance over a series of setups. Finally, we show that we reduce forgetting of upstream pretraining examples by replaying examples that are forecasted to be forgotten, demonstrating the practical utility of forecasting example forgetting. Xisen Jin, Xiang Ren 0001 |
ICML | 2 |
| 2024 | Stress-Testing Long-Context Language Models with Lifelong ICL and Task HaystackabstractWe introduce Lifelong ICL, a problem setting that challenges long-context language models (LMs) to learn a sequence of language tasks through in-context learning (ICL). We further introduce Task Haystack, an evaluation suite dedicated to assessing and diagnosing how long-context LMs utilizes contexts in Lifelong ICL. When given a task instruction and test inputs, long-context LMs are expectedto leverage the relevant demonstrations in the Lifelong ICL prompt, avoid distraction and interference from other tasks, and achieve test accuracies that are not significantly worse than those of the Single-task ICL baseline.Task Haystack draws inspiration from the widely-adopted “needle-in-a-haystack” (NIAH) evaluation, but presents distinct new challenges. It requires models (1) to utilize the contexts at a deeper level, rather than resorting to simple copying and pasting; (2) to navigate through long streams of evolving topics and tasks, proxying the complexities and dynamism of contexts in real-world scenarios. Additionally, Task Haystack inherits the controllability of NIAH, providing model developers with tools and visualizations to identify model vulnerabilities effectively.We benchmark 14 long-context LMs using Task Haystack, finding that frontier models like GPT-4o still struggle with the setting, failing on 15% of cases on average. Most open-weight models further lack behind by a large margin, with failure rates reaching up to 61%. In our controlled analysis, we identify factors such as distraction and recency bias as contributors to these failure cases. Further, performance declines when task instructions are paraphrased at test time or when ICL demonstrations are repeated excessively, raising concerns about the robustness, instruction understanding, and true context utilization of long-context LMs. We release our code and data to encourage future research that investigates and addresses these limitations. Xiaoyue Xu, Qinyuan Ye, Xiang Ren 0001 |
NeurIPS | 3 |
| 2024 | SELF-DISCOVER: Large Language Models Self-Compose Reasoning StructuresabstractWe introduce SELF-DISCOVER, a general framework for LLMs to self-discover the task-intrinsic reasoning structures to tackle complex reasoning problems that are challenging for typical prompting methods. Core to the framework is a self-discovery process where LLMs select multiple atomic reasoning modules such as critical thinking and step-by-step thinking, and compose them into an explicit reasoning structure for LLMs to follow during decoding. SELF-DISCOVER substantially improves GPT-4 and PaLM 2’s performance on challenging reasoning benchmarks such as BigBench-Hard, grounded agent reasoning, and MATH, by as much as 32% compared to Chain of Thought (CoT). Furthermore, SELF-DISCOVER outperforms inference-intensive methods such as CoT-Self-Consistency by more than 20%, while requiring 10-40x fewer inference compute. Finally, we show that the self-discovered reasoning structures are universally applicable across model families: from PaLM 2-L to GPT-4, and from GPT-4 to Llama2, and share commonalities with human reasoning patterns. Jay Pujara, Xiang Ren 0001, Heng-Tze Cheng, Quoc V. Le, Ed H. Chi, Denny Zhou, Swaroop Mishra, Huaixiu Steven Zheng |
NeurIPS | 3 |
| 2023 | On Grounded Planning for Embodied Tasks with Language ModelsabstractLanguage models (LMs) have demonstrated their capability in possessing commonsense knowledge of the physical world, a crucial aspect of performing tasks in everyday life. However, it remains unclear whether they have the capacity to generate grounded, executable plans for embodied tasks. This is a challenging task as LMs lack the ability to perceive the environment through vision and feedback from the physical environment. In this paper, we address this important research question and present the first investigation into the topic. Our novel problem formulation, named G-PlanET, inputs a high-level goal and a data table about objects in a specific environment, and then outputs a step-by-step actionable plan for a robotic agent to follow. To facilitate the study, we establish an evaluation protocol and design a dedicated metric, KAS, to assess the quality of the plans. Our experiments demonstrate that the use of tables for encoding the environment and an iterative decoding strategy can significantly enhance the LMs' ability in grounded planning. Our analysis also reveals interesting and non-trivial findings. Bill Y. Lin, Chengsong Huang, Qian Liu 0033, Wenda Gu, Sam Sommerer, Xiang Ren 0001 |
AAAI | 6 |
| 2023 | APOLLO: A Simple Approach for Adaptive Pretraining of Language Models for Logical ReasoningabstractSoumya Sanyal, Yichong Xu, Shuohang Wang, Ziyi Yang, Reid Pryzant, Wenhao Yu, Chenguang Zhu, Xiang Ren. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Soumya Sanyal 0001, Yichong Xu, Shuohang Wang, Ziyi Yang 0011, Reid Pryzant, Wenhao Yu 0002, Chenguang Zhu 0001, Xiang Ren 0001 |
ACL (1) | 8 |
| 2023 | BITE: Textual Backdoor Attacks with Iterative Trigger InjectionabstractBackdoor attacks have become an emerging threat to NLP systems.By providing poisoned training data, the adversary can embed a "backdoor" into the victim model, which allows input instances satisfying certain textual patterns (e.g., containing a keyword) to be predicted as a target label of the adversary's choice.In this paper, we demonstrate that it is possible to design a backdoor attack that is both stealthy (i.e., hard to notice) and effective (i.e., has a high attack success rate).We propose BITE, a backdoor attack that poisons the training data to establish strong correlations between the target label and a set of "trigger words".These trigger words are iteratively identified and injected into the target-label instances through natural word-level perturbations.The poisoned training data instruct the victim model to predict the target label on inputs containing trigger words, forming the backdoor.Experiments on four text classification datasets show that our proposed attack is significantly more effective than baseline methods while maintaining decent stealthiness, raising alarm on the usage of untrusted training data.We further propose a defense method named DeBITE based on potential trigger word removal, which outperforms existing methods in defending against BITE and generalizes well to handling other backdoor attacks. 1 Jun Yan 0012, Vansh Gupta, Xiang Ren 0001 |
ACL (1) | 3 |
| 2023 | REV: Information-Theoretic Evaluation of Free-Text RationalesabstractHanjie Chen, Faeze Brahman, Xiang Ren, Yangfeng Ji, Yejin Choi, Swabha Swayamdipta. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Faeze Brahman, Xiang Ren 0001, Yangfeng Ji, Yejin Choi 0001, Swabha Swayamdipta |
ACL (1) | 3 |
| 2023 | LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative FusionabstractWe present LLM-BL E N D E R, an ensembling framework designed to attain consistently superior performance by leveraging the diverse strengths of multiple open-source large language models (LLMs).Our framework consists of two modules: PAIRRANKER and GEN-FUSER, addressing the observation that optimal LLMs for different examples can significantly vary.PAIRRANKER employs a specialized pairwise comparison method to distinguish subtle differences between candidate outputs.It jointly encodes the input text and a pair of candidates, using cross-attention encoders to determine the superior one.Our results demonstrate that PAIRRANKER exhibits the highest correlation with ChatGPT-based ranking.Then, GENFUSER aims to merge the top-ranked candidates, generating an improved output by capitalizing on their strengths and mitigating their weaknesses.To facilitate largescale evaluation, we introduce a benchmark dataset, MixInstruct, which is a mixture of multiple instruction datasets featuring oracle pairwise comparisons.Our LLM-BL E N D E R significantly outperform individual LLMs and baseline methods across various metrics, establishing a substantial performance gap. 1 2 Open Assistant 12.61% Koala 6.71% Alpaca 11.61% Baize 11.61% StableLM 1.90% FLAN-T5 0.80% Vicuna 21.22% Dolly V2 4.50% MOSS 12.91% ChatGLM 8.51% MPT 7.61% Percentage of Examples Where Each Model Ranks First Which LLM should I use for my input?All!I can ensemble! Dongfu Jiang, Xiang Ren 0001, Bill Y. Lin |
ACL (1) | 2 |
| 2023 | Are Machine Rationales (Not) Useful to Humans? Measuring and Improving Human Utility of Free-text RationalesabstractBrihi Joshi, Ziyi Liu, Sahana Ramnath, Aaron Chan, Zhewei Tong, Shaoliang Nie, Qifan Wang, Yejin Choi, Xiang Ren. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Brihi Joshi, Ziyi Liu 0007, Sahana Ramnath, Aaron Chan, Zhewei Tong, Shaoliang Nie, Qifan Wang 0001, Yejin Choi 0001, Xiang Ren 0001 |
ACL (1) | 9 |
| 2023 | Symbolic Chain-of-Thought Distillation: Small Models Can Also "Think" Step-by-StepabstractLiunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren, Kai-Wei Chang, Yejin Choi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Liunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren 0001, Kai-Wei Chang 0001, Yejin Choi 0001 |
ACL (1) | 4 |
| 2023 | Cross-lingual Continual LearningabstractThe longstanding goal of multi-lingual learning has been to develop a universal cross-lingual model that can withstand the changes in multilingual data distributions.There has been a large amount of work to adapt such multilingual models to unseen target languages.However, the majority of work in this direction focuses on the standard one-hop transfer learning pipeline from source to target languages, whereas in realistic scenarios, new languages can be incorporated at any time in a sequential manner.In this paper, we present a principled Cross-lingual Continual Learning (CCL) evaluation paradigm, where we analyze different categories of approaches used to continually adapt to emerging data from different languages.We provide insights into what makes multilingual sequential learning particularly challenging.To surmount such challenges, we benchmark a representative set of cross-lingual continual learning algorithms and analyze their knowledge preservation, accumulation, and generalization capabilities compared to baselines on carefully curated datastreams.The implications of this analysis include a recipe for how to measure and balance different cross-lingual continual learning desiderata, which go beyond conventional transfer learning. Meryem M'hamdi, Xiang Ren 0001, Jonathan May |
ACL (1) | 2 |
| 2023 | SCOTT: Self-Consistent Chain-of-Thought DistillationabstractLarge language models (LMs) beyond a certain scale, demonstrate the emergent capability of generating free-text rationales for their predictions via chain-of-thought (CoT) prompting.While CoT can yield dramatically improved performance, such gains are only observed for sufficiently large LMs.Even more concerning, there is little guarantee that the generated rationales are consistent with LM's predictions or faithfully justify the decisions.In this work, we propose SCOTT, a faithful knowledge distillation method to learn a small, self-consistent CoT model from a teacher model that is orders of magnitude larger.To form better supervision, we elicit rationales supporting the gold answers from a large LM (teacher) by contrastive decoding, which encourages the teacher to generate tokens that become more plausible only when the answer is considered.To ensure faithful distillation, we use the teacher-generated rationales to learn a student LM with a counterfactual reasoning objective, which prevents the student from ignoring the rationales to make inconsistent predictions.Experiments show that, while yielding comparable end-task performance, our method can generate CoT rationales that are more faithful than baselines do.Further analysis suggests that such a model respects the rationales more when making decisions; thus, we can improve its performance more by refining its rationales. Peifeng Wang, Zheng Li 0018, Yifan Gao 0001, Xiang Ren 0001 |
ACL (1) | 6 |
| 2023 | Contrastive Novelty-Augmented Learning: Anticipating Outliers with Large Language ModelsabstractIn many task settings, text classification models are likely to encounter examples from novel classes on which they cannot predict correctly.Selective prediction, in which models abstain on low-confidence examples, provides a possible solution, but existing models are often overly confident on unseen classes.To remedy this overconfidence, we introduce Contrastive Novelty-Augmented Learning (CoNAL), a twostep method that generates OOD examples representative of novel classes, then trains to decrease confidence on them.First, we generate OOD examples by prompting a large language model twice: we prompt it to enumerate relevant novel classes, then generate examples from each novel class matching the task format.Second, we train a classifier with a novel contrastive objective that encourages lower confidence on generated OOD examples than training examples.When trained with CoNAL, classifiers improve in their ability to detect and abstain on novel class examples over prior methods by an average of 2.3% in terms of accuracy under the accuracy-coverage curve (AUAC) and 5.5% AUROC across 4 NLP datasets, with no cost to in-distribution accuracy.1 Albert Xu, Xiang Ren 0001, Robin Jia |
ACL (1) | 2 |
| 2023 | FiD-ICL: A Fusion-in-Decoder Approach for Efficient In-Context LearningabstractLarge pre-trained models are capable of fewshot in-context learning (ICL), i.e., performing a new task by prepending a few demonstrations before the test input.However, the concatenated demonstrations are often excessively long and induce additional computation.Inspired by fusion-in-decoder (FiD) models which efficiently aggregate more passages and thus outperforms concatenation-based models in opendomain QA, we hypothesize that similar techniques can be applied to improve the efficiency and end-task performance of ICL.To verify this, we present a comprehensive study on applying three fusion methods-concatenationbased (early fusion), FiD (intermediate), and ensemble-based (late)-to ICL.We adopt a meta-learning setup where a model is first trained to perform ICL on a mixture of tasks using one selected fusion method, then evaluated on held-out tasks for ICL.Results on 11 heldout tasks show that FiD-ICL matches or outperforms the other two fusion methods.Additionally, we show that FiD-ICL (1) is 10x faster at inference time compared to concat-based and ensemble-based ICL, as we can easily precompute the representations of in-context examples and reuse them; (2) enables scaling up to meta-training 3B-sized models, which would fail for concat-based ICL.1 Qinyuan Ye, Iz Beltagy, Matthew E. Peters, Xiang Ren 0001, Hannaneh Hajishirzi |
ACL (1) | 4 |
| 2023 | I Cast Detect Thoughts: Learning to Converse and Guide with Intents and Theory-of-Mind in Dungeons and DragonsabstractPei Zhou, Andrew Zhu, Jennifer Hu, Jay Pujara, Xiang Ren, Chris Callison-Burch, Yejin Choi, Prithviraj Ammanabrolu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Andrew Zhu, Jennifer Hu 0001, Jay Pujara, Xiang Ren 0001, Chris Callison-Burch, Yejin Choi 0001, Prithviraj Ammanabrolu |
ACL (1) | 5 |
| 2023 | AutoTriggER: Label-Efficient and Robust Named Entity Recognition with Auxiliary Trigger ExtractionabstractDong-Ho Lee, Ravi Kiran Selvam, Sheikh Muhammad Sarwar, Bill Yuchen Lin, Fred Morstatter, Jay Pujara, Elizabeth Boschee, James Allan, Xiang Ren. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Ravi Kiran Selvam, Sheikh Muhammad Sarwar, Bill Y. Lin, Fred Morstatter, Jay Pujara, Elizabeth Boschee, James Allan 0001, Xiang Ren 0001 |
EACL | 9 |
| 2023 | Inference-Time Policy Adapters (IPA): Tailoring Extreme-Scale LMs without Fine-tuningabstractXiming Lu, Faeze Brahman, Peter West, Jaehun Jung, Khyathi Chandu, Abhilasha Ravichander, Prithviraj Ammanabrolu, Liwei Jiang, Sahana Ramnath, Nouha Dziri, Jillian Fisher, Bill Lin, Skyler Hallinan, Lianhui Qin, Xiang Ren, Sean Welleck, Yejin Choi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Ximing Lu, Faeze Brahman, Peter West, Jaehun Jung, Khyathi Raghavi Chandu, Abhilasha Ravichander, Prithviraj Ammanabrolu, Sahana Ramnath, Nouha Dziri, Jillian Fisher, Bill Y. Lin, Skyler Hallinan, Lianhui Qin, Xiang Ren 0001, Sean Welleck, Yejin Choi 0001 |
EMNLP | 15 |
| 2023 | Dataless Knowledge Fusion by Merging Weights of Language Models
Xisen Jin, Xiang Ren 0001, Daniel Preotiuc-Pietro, Pengxiang Cheng 0001 |
ICLR | 2 |
| 2023 | PINTO: Faithful Language Reasoning Using Prompt-Generated Rationales
Peifeng Wang, Aaron Chan, Filip Ilievski, Muhao Chen 0001, Xiang Ren 0001 |
ICLR | 5 |
| 2023 | Retweet-BERT: Political Leaning Detection Using Language Features and Information Diffusion on Social NetworksabstractEstimating the political leanings of social media users is a challenging and ever more pressing problem given the increase in social media consumption. We introduce Retweet-BERT, a simple and scalable model to estimate the political leanings of Twitter users. Retweet-BERT leverages the retweet network structure and the language used in users' profile descriptions. Our assumptions stem from patterns of networks and linguistics homophily among people who share similar ideologies. Retweet-BERT demonstrates competitive performance against other state-of-the-art baselines, achieving 96%-97% macro-F1 on two recent Twitter datasets (a COVID-19 dataset and a 2020 United States presidential elections dataset). We also perform manual validation to validate the performance of Retweet-BERT on users not in the training data. Finally, in a case study of COVID-19, we illustrate the presence of political echo chambers on Twitter and show that it exists primarily among right-leaning users. Our code is open-sourced and our data is publicly available. Julie Jiang, Xiang Ren 0001, Emilio Ferrara |
ICWSM | 2 |
| 2023 | Faith and Fate: Limits of Transformers on CompositionalityabstractTransformer large language models (LLMs) have sparked admiration for their exceptional performance on tasks that demand intricate multi-step reasoning. Yet, these models simultaneously show failures on surprisingly trivial problems.
This begs the question: Are these errors incidental, or do they signal more substantial limitations?
In an attempt to demystify transformer LLMs, we investigate the limits of these models across three representative compositional tasks---multi-digit multiplication, logic grid puzzles, and a classic dynamic programming problem. These tasks require breaking problems down into sub-steps and synthesizing these steps into a precise answer. We formulate compositional tasks as computation graphs to systematically quantify the level of complexity, and break down reasoning steps into intermediate sub-procedures.
Our empirical findings suggest that transformer LLMs solve compositional tasks by reducing multi-step compositional reasoning into linearized subgraph matching, without necessarily developing systematic problem-solving skills. To round off our empirical study, we provide theoretical arguments on abstract multi-step reasoning problems that highlight how autoregressive generations' performance can rapidly decay with increased task complexity. Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Li 0069, Bill Y. Lin, Sean Welleck, Peter West, Chandra Bhagavatula, Ronan Le Bras 0001, Jena D. Hwang, Soumya Sanyal 0001, Xiang Ren 0001, Allyson Ettinger, Zaïd Harchaoui, Yejin Choi 0001 |
NeurIPS | 13 |
| 2023 | SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive TasksabstractWe introduce SwiftSage, a novel agent framework inspired by the dual-process theory of human cognition, designed to excel in action planning for complex interactive reasoning tasks. SwiftSage integrates the strengths of behavior cloning and prompting large language models (LLMs) to enhance task completion performance. The framework comprises two primary modules: the Swift module, representing fast and intuitive thinking, and the Sage module, emulating deliberate thought processes. The Swift module is a small encoder-decoder LM fine-tuned on the oracle agent's action trajectories, while the Sage module employs LLMs such as GPT-4 for subgoal planning and grounding. We develop a heuristic method to harmoniously integrate the two modules, resulting in a more efficient and robust problem-solving process. In 30 tasks from the ScienceWorld benchmark, SwiftSage significantly outperforms other methods such as SayCan, ReAct, and Reflexion, demonstrating its effectiveness in solving complex interactive tasks. Bill Y. Lin, Yicheng Fu, Karina Yang, Faeze Brahman, Shiyu Huang 0001, Chandra Bhagavatula, Prithviraj Ammanabrolu, Yejin Choi 0001, Xiang Ren 0001 |
NeurIPS | 9 |
| 2023 | Knowledge-Augmented Methods for Natural Language ProcessingabstractKnowledge in NLP has been a rising trend especially after the advent of large-scale pre-trained models. Knowledge is critical to equip statistics-based models with common sense, logic and other external information. In this tutorial, we will introduce recent state-of-the-art works in applying knowledge in language understanding, language generation and commonsense reasoning. Chenguang Zhu 0001, Yichong Xu, Xiang Ren 0001, Bill Y. Lin, Meng Jiang 0001, Wenhao Yu 0002 |
WSDM | 3 |
| 2023 | L-BGNN: Layerwise Trained Bipartite Graph Neural NetworksabstractLearning low-dimensional representations of bipartite graphs enables e-commerce applications, such as recommendation, classification, and link prediction. A layerwise-trained bipartite graph neural network (L-BGNN) embedding method, which is unsupervised, efficient, and scalable, is proposed in this work. To aggregate the information across and within two partitions of a bipartite graph, a customized interdomain message passing (IDMP) operation and an intradomain alignment (IDA) operation are adopted by the proposed L-BGNN method. Furthermore, we develop a layerwise training algorithm for L-BGNN to capture the multihop relationship of large bipartite networks and improve training efficiency. We conduct extensive experiments on several datasets and downstream tasks of various scales to demonstrate the effectiveness and efficiency of the L-BGNN method as compared with state-of-the-art methods. Our codes are publicly available at https://github.com/TianXieUSC/L-BGNN. Tian Xie 0005, Chaoyang He 0001, Xiang Ren 0001, Cyrus Shahabi, C.-C. Jay Kuo |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language ModelsabstractLarge pre-trained vision-language (VL) models can learn a new task with a handful of examples and generalize to a new task without fine-tuning.However, these VL models are hard to deploy for real-world applications due to their impractically huge sizes and slow inference speed.To solve this limitation, we study prompt-based low-resource learning of VL tasks with our proposed method, FEWVLM, relatively smaller than recent fewshot learners.For FEWVLM, we pre-train a sequence-to-sequence transformer model with prefix language modeling (PrefixLM) and masked language modeling (MaskedLM).Furthermore, we analyze the effect of diverse prompts for few-shot tasks.Experimental results on VQA show that FEWVLM with prompt-based learning outperforms Frozen (Tsimpoukelli et al., 2021) which is 31× larger than FEWVLM by 18.2% point and achieves comparable results to a 246× larger model, PICa (Yang et al., 2021).In our analysis, we observe that (1) prompts significantly affect zero-shot performance but marginally affect few-shot performance, (2) models with noisy prompts learn as quickly as hand-crafted prompts given larger training data, and (3) MaskedLM helps VQA tasks while PrefixLM boosts captioning performance.Our code is publicly available at https://github. com/woojeongjin/FewVLM * Work was mainly done while Woojeong Jin 0001, Yu Cheng 0001, Yelong Shen, Weizhu Chen, Xiang Ren 0001 |
ACL (1) | 5 |
| 2022 | Leveraging Visual Knowledge in Language Tasks: An Empirical Study on Intermediate Pre-training for Cross-Modal Knowledge TransferabstractPre-trained language models are still far from human performance in tasks that need understanding of properties (e.g.appearance, measurable quantity) and affordances of everyday objects in the real world since the text lacks such information due to reporting bias.In this work, we study whether integrating visual knowledge into a language model can fill the gap.We investigate two types of knowledge transfer: (1) text knowledge transfer using image captions that may contain enriched visual knowledge and (2) cross-modal knowledge transfer using both images and captions with vision-language training objectives.On 5 downstream tasks that may need visual knowledge to solve the problem, we perform extensive empirical comparisons over the presented objectives.Our experiments show that visual knowledge transfer can improve performance in both low-resource and fully supervised settings.1 Woojeong Jin 0001, Chenguang Zhu 0001, Jay Pujara, Xiang Ren 0001 |
ACL (1) | 5 |
| 2022 | Good Examples Make A Faster Learner: Simple Demonstration-based Learning for Low-resource NERabstractDong-Ho Lee, Akshen Kadakia, Kangmin Tan, Mahak Agarwal, Xinyu Feng, Takashi Shibuya, Ryosuke Mitani, Toshiyuki Sekiya, Jay Pujara, Xiang Ren. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Akshen Kadakia, Kangmin Tan, Mahak Agarwal, Takashi Shibuya 0001, Ryosuke Mitani, Toshiyuki Sekiya, Jay Pujara, Xiang Ren 0001 |
ACL (1) | 10 |
| 2022 | On Continual Model Refinement in Out-of-Distribution Data StreamsabstractReal-world natural language processing (NLP) models need to be continually updated to fix the prediction errors in out-of-distribution (OOD) data streams while overcoming catastrophic forgetting.However, existing continual learning (CL) problem setups cannot cover such a realistic and complex scenario.In response to this, we propose a new CL problem formulation dubbed continual model refinement (CMR).Compared to prior CL settings, CMR is more practical and introduces unique challenges (boundary-agnostic and non-stationary distribution shift, diverse mixtures of multiple OOD data clusters, error-centric streams, etc.).We extend several existing CL approaches to the CMR setting and evaluate them extensively.For benchmarking and analysis, we propose a general sampling algorithm to obtain dynamic OOD data streams with controllable nonstationarity, as well as a suite of metrics measuring various aspects of online performance.Our experiments and detailed analysis reveal the promise and challenges of the CMR problem, supporting that studying CMR in dynamic OOD streams can benefit the longevity of deployed NLP models in production. 1 Bill Y. Lin, Sida I. Wang, Xi Victoria Lin, Robin Jia, Xiang Ren 0001, Scott Yih |
ACL (1) | 6 |
| 2022 | FaiRR: Faithful and Robust Deductive Reasoning over Natural LanguageabstractTransformers have been shown to be able to perform deductive reasoning on a logical rulebase containing rules and statements written in natural language.Recent works show that such models can also produce the reasoning steps (i.e., the proof graph) that emulate the model's logical reasoning process.Currently, these black-box models generate both the proof graph and intermediate inferences within the same model and thus may be unfaithful.In this work, we frame the deductive logical reasoning task by defining three modular components: rule selection, fact selection, and knowledge composition.The rule and fact selection steps select the candidate rule and facts to be used and then the knowledge composition combines them to generate new inferences.This ensures model faithfulness by assured causal relation from the proof step to the inference reasoning.To test our framework, we propose FAIRR (Faithful and Robust Reasoner) where the above three components are independently modeled by transformers.We observe that FAIRR is robust to novel language perturbations, and is faster at inference than previous works on existing reasoning datasets.Additionally, in contrast to black-box generative models, the errors made by FAIRR are more interpretable due to the modular approach.1 Soumya Sanyal 0001, Harman Singh, Xiang Ren 0001 |
ACL (1) | 3 |
| 2022 | KG-FiD: Infusing Knowledge Graph in Fusion-in-Decoder for Open-Domain Question AnsweringabstractDonghan Yu, Chenguang Zhu, Yuwei Fang, Wenhao Yu, Shuohang Wang, Yichong Xu, Xiang Ren, Yiming Yang, Michael Zeng. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Donghan Yu, Chenguang Zhu 0001, Yuwei Fang, Wenhao Yu 0002, Shuohang Wang, Yichong Xu, Xiang Ren 0001, Yiming Yang 0002, Michael Zeng 0001 |
ACL (1) | 7 |
| 2022 | Think Before You Speak: Explicitly Generating Implicit Commonsense Knowledge for Response GenerationabstractPei Zhou, Karthik Gopalakrishnan, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren 0001, Yang Liu 0004, Dilek Hakkani-Tür |
ACL (1) | 6 |
| 2022 | RobustLR: A Diagnostic Benchmark for Evaluating Logical Robustness of Deductive ReasonersabstractTransformers have been shown to be able to perform deductive reasoning on inputs containing rules and statements written in English natural language.However, it is unclear if these models indeed follow rigorous logical reasoning to arrive at the prediction, or rely on spurious correlation patterns in making decision.A strong deductive reasoning model should consistently understand the semantics of different logical operators.To this end, we present ROBUSTLR, a deductive reasoning-based diagnostic benchmark that evaluates the robustness of language models to minimal logical edits in the inputs and different logical equivalence conditions.In our experiments with RoBERTa, T5, and GPT3, we show that the models trained on deductive reasoning datasets with various logical operations do not perform consistently on the RO-BUSTLR test set, thus showing that the models are not robust to our proposed logical perturbations.Further, we observe that the models find it especially hard to learn logical negation operator.Our results demonstrate the shortcomings of current language models in logical reasoning, and call for the development of better inductive biases to teach the logical semantics to language models.All the datasets and code base have been made publicly available.1 Soumya Sanyal 0001, Zeyi Liao, Xiang Ren 0001 |
EMNLP | 3 |
| 2022 | Machine Translation Robustness to Natural Asemantic VariationabstractCurrent Machine Translation (MT) models still struggle with more challenging input, such as noisy data and tail-end words and phrases.Several works have addressed this robustness issue by identifying specific categories of noise and variation then tuning models to perform better on them.An important yet under-studied category involves minor variations in nuance (non-typos) that preserve meaning w.r.t. the target language.We introduce and formalize this category as Natural Asemantic Variation (NAV) and investigate it in the context of MT robustness.We find that existing MT models fail when presented with NAV data, but we demonstrate strategies to improve performance on NAV by fine-tuning them with human-generated variations.We also show that NAV robustness can be transferred across languages and find that synthetic perturbations can achieve some but not all of the benefits of organic NAV data. Jacob Bremerman, Xiang Ren 0001, Jonathan May |
EMNLP | 2 |
| 2022 | Reflect, Not Reflex: Inference-Based Common Ground Improves Dialogue Response QualityabstractHuman communication relies on common ground (CG), the mutual knowledge and beliefs shared by participants, to produce coherent and interesting conversations.In this paper, we demonstrate that current response generation (RG) models produce generic and dull responses in dialogues because they act reflexively, failing to explicitly model CG, both due to the lack of CG in training data and the standard RG training procedure.We introduce Reflect, a dataset that annotates dialogues with explicit CG (materialized as inferences approximating shared knowledge and beliefs) and solicits 9k diverse human-generated responses each following one common ground.Using Reflect, we showcase the limitations of current dialogue data and RG models: less than half of the responses in current data is rated as high quality (sensible, specific, and interesting) and models trained using this data have even lower quality, while most Reflect responses are judged high quality.Next, we analyze whether CG can help models produce better quality responses by using Reflect CG to guide RG models.Surprisingly, we find that simply prompting GPT3 to "think" about CG generates 30% more quality responses, showing promising benefits to integrating CG into the RG process. 1 Hyundong Cho, Pegah Jandaghi, Bill Y. Lin, Jay Pujara, Xiang Ren 0001 |
EMNLP | 7 |
| 2022 | Contextualized Scene Imagination for Generative Commonsense Reasoning
Peifeng Wang, Jonathan Zamora, Filip Ilievski, Muhao Chen 0001, Xiang Ren 0001 |
ICLR | 6 |
| 2022 | UNIREX: A Unified Learning Framework for Language Model Rationale ExtractionabstractAn extractive rationale explains a language model’s (LM’s) prediction on a given task instance by highlighting the text inputs that most influenced the prediction. Ideally, rationale extraction should be faithful (reflective of LM’s actual behavior) and plausible (convincing to humans), without compromising the LM’s (i.e., task model’s) task performance. Although attribution algorithms and select-predict pipelines are commonly used in rationale extraction, they both rely on certain heuristics that hinder them from satisfying all three desiderata. In light of this, we propose UNIREX, a flexible learning framework which generalizes rationale extractor optimization as follows: (1) specify architecture for a learned rationale extractor; (2) select explainability objectives (\ie faithfulness and plausibility criteria); and (3) jointly train the task model and rationale extractor on the task using selected objectives. UNIREX enables replacing prior works’ heuristic design choices with a generic learned rationale extractor in (1) and optimizing it for all three desiderata in (2)-(3). To facilitate comparison between methods w.r.t. multiple desiderata, we introduce the Normalized Relative Gain (NRG) metric. On five English text classification datasets, our best UNIREX configuration outperforms baselines by an average of 32.9% NRG. Plus, UNIREX rationale extractors’ faithfulness can even generalize to unseen datasets and tasks. Aaron Chan, Maziar Sanjabi, Lambert Mathias, Liang Tan 0005, Shaoliang Nie, Xiaochang Peng, Xiang Ren 0001, Hamed Firooz |
ICML | 7 |
| 2022 | Lifelong Pretraining: Continually Adapting Language Models to Emerging CorporaabstractXisen Jin, Dejiao Zhang, Henghui Zhu, Wei Xiao, Shang-Wen Li, Xiaokai Wei, Andrew Arnold, Xiang Ren. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Xisen Jin, Dejiao Zhang, Henghui Zhu, Wei Xiao 0001, Shang-Wen Li 0001, Xiaokai Wei, Andrew O. Arnold, Xiang Ren 0001 |
NAACL-HLT | 8 |
| 2022 | NewsEdits: A News Article Revision Dataset and a Novel Document-Level Reasoning ChallengeabstractAlexander Spangher, Xiang Ren, Jonathan May, Nanyun Peng. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Alexander Spangher, Xiang Ren 0001, Jonathan May, Nanyun Peng 0001 |
NAACL-HLT | 2 |
| 2022 | On the Robustness of Reading Comprehension Models to Entity RenamingabstractJun Yan, Yang Xiao, Sagnik Mukherjee, Bill Yuchen Lin, Robin Jia, Xiang Ren. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jun Yan 0012, Sagnik Mukherjee, Bill Y. Lin, Robin Jia, Xiang Ren 0001 |
NAACL-HLT | 6 |
| 2022 | Sparse Distillation: Speeding Up Text Classification by Using Bigger Student ModelsabstractQinyuan Ye, Madian Khabsa, Mike Lewis, Sinong Wang, Xiang Ren, Aaron Jaech. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Qinyuan Ye, Madian Khabsa, Mike Lewis, Sinong Wang, Xiang Ren 0001, Aaron Jaech |
NAACL-HLT | 5 |
| 2022 | Unsupervised Cross-Task Generalization via Retrieval AugmentationabstractHumans can perform unseen tasks by recalling relevant skills acquired previously and then generalizing them to the target tasks, even if there is no supervision at all. In this paper, we aim to improve this kind of cross-task generalization ability of massive multi-task language models, such as T0 and FLAN, in an unsupervised setting. We propose a retrieval-augmentation method named ReCross that takes a few unlabelled examples as queries to retrieve a small subset of upstream data and uses them to update the multi-task model for better generalization. ReCross is a straightforward yet effective retrieval method that combines both efficient dense retrieval and effective pair-wise reranking. Our results and analysis show that it significantly outperforms both non-retrieval methods and other baseline methods. Bill Y. Lin, Kangmin Tan, Beiwen Tian, Xiang Ren 0001 |
NeurIPS | 5 |
| 2022 | Assessing Scientific Research Papers with Knowledge GraphsabstractIn recent decades, the growing scale of scientific research has led to numerous novel findings. Reproducing these findings is the foundation of future research. However, due to the complexity of experiments, manually assessing scientific research is laborious and time-intensive, especially in social and behavioral sciences. Although increasing reproducibility studies have garnered increased attention in the research community, there is still a lack of systematic ways for evaluating scientific research at scale. In this paper, we propose a novel approach towards automatically assessing scientific publications by constructing a knowledge graph (KG) that captures a holistic view of the research contributions. Specifically, during the KG construction, we combine information from two different perspectives: micro-level features that capture knowledge from published articles such as sample sizes, effect sizes, and experimental models, and macro-level features that comprise relationships between entities such as authorship and reference information. We then learn low-dimensional representations using language models and knowledge graph embeddings for entities (nodes in KGs), which are further used for the assessments. A comprehensive set of experiments on two benchmark datasets shows the usefulness of leveraging KGs for scoring scientific research. Kexuan Sun 0002, Zhiqiang Qiu, Abel Salinas, Yuzhong Huang, Daniel Benjamin, Fred Morstatter, Xiang Ren 0001, Kristina Lerman, Jay Pujara |
SIGIR | 8 |
| 2021 | IsoBN: Fine-Tuning BERT with Isotropic Batch NormalizationabstractFine-tuning pre-trained language models (PTLMs), such as BERT and its better variant RoBERTa, has been a common practice for advancing performance in natural language understanding (NLU) tasks. Recent advance in representation learning shows that isotropic (i.e., unit-variance and uncorrelated) embeddings can significantly improve performance on downstream tasks with faster convergence and better generalization. The isotropy of the pre-trained embeddings in PTLMs, however, is relatively under-explored. In this paper, we analyze the isotropy of the pre-trained [CLS] embeddings of PTLMs with straightforward visualization, and point out two major issues: high variance in their standard deviation, and high correlation between different dimensions. We also propose a new network regularization method, isotropic batch normalization (IsoBN) to address the issues, towards learning more isotropic representations in fine-tuning by dynamically penalizing dominating principal components. This simple yet effective fine-tuning method yields about 1.0 absolute increment on the average of seven NLU tasks. Wenxuan Zhou 0002, Bill Y. Lin, Xiang Ren 0001 |
AAAI | 3 |
| 2021 | ForecastQA: A Question Answering Challenge for Event Forecasting with Temporal Text DataabstractWoojeong Jin, Rahul Khanna, Suji Kim, Dong-Ho Lee, Fred Morstatter, Aram Galstyan, Xiang Ren. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Woojeong Jin 0001, Rahul Khanna, Fred Morstatter, Aram Galstyan, Xiang Ren 0001 |
ACL/IJCNLP (1) | 7 |
| 2021 | Common Sense Beyond English: Evaluating and Improving Multilingual Language Models for Commonsense ReasoningabstractBill Yuchen Lin, Seyeon Lee, Xiaoyang Qiao, Xiang Ren. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Bill Y. Lin, Seyeon Lee, Xiaoyang Qiao, Xiang Ren 0001 |
ACL/IJCNLP (1) | 4 |
| 2021 | AdaTag: Multi-Attribute Value Extraction from Product Profiles with Adaptive DecodingabstractJun Yan, Nasser Zalmout, Yan Liang, Christan Grant, Xiang Ren, Xin Luna Dong. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jun Yan 0012, Nasser Zalmout, Yan Liang 0004, Christan Grant, Xiang Ren 0001, Xin Dong 0001 |
ACL/IJCNLP (1) | 5 |
| 2021 | Cross-Attention is All You Need: Adapting Pretrained Transformers for Machine TranslationabstractWe study the power of cross-attention in the Transformer architecture within the context of transfer learning for machine translation, and extend the findings of studies into crossattention when training from scratch.We conduct a series of experiments through finetuning a translation model on data where either the source or target language has changed.These experiments reveal that fine-tuning only the cross-attention parameters is nearly as effective as fine-tuning all parameters (i.e., the entire translation model).We provide insights into why this is the case and observe that limiting fine-tuning in this manner yields crosslingually aligned embeddings.The implications of this finding for researchers and practitioners include a mitigation of catastrophic forgetting, the potential for zero-shot translation, and the ability to extend machine translation models to several new language pairs with reduced parameter storage overhead.1 1 Our code is available at https://github.com/ MGheini/xattn-transfer-for-mt. Mozhdeh Gheini, Xiang Ren 0001, Jonathan May |
EMNLP (1) | 2 |
| 2021 | ECONET: Effective Continual Pretraining of Language Models for Event Temporal ReasoningabstractWhile pre-trained language models (PTLMs) have achieved noticeable success on many NLP tasks, they still struggle for tasks that require event temporal reasoning, which is essential for event-centric applications.We present a continual pre-training approach that equips PTLMs with targeted knowledge about event temporal relations.We design self-supervised learning objectives to recover masked-out event and temporal indicators and to discriminate sentences from their corrupted counterparts (where event or temporal indicators got replaced).By further pre-training a PTLM with these objectives jointly, we reinforce its attention to event and temporal information, yielding enhanced capability on event temporal reasoning.This Effective CONtinual pre-training framework for Event Temporal reasoning (ECONET) improves the PTLMs' fine-tuning performances across five relation extraction and question answering tasks and achieves new or on-par state-of-the-art performances in most of our downstream tasks. 1 Rujun Han, Xiang Ren 0001, Nanyun Peng 0001 |
EMNLP (1) | 2 |
| 2021 | RockNER: A Simple Method to Create Adversarial Examples for Evaluating the Robustness of Named Entity Recognition ModelsabstractTo audit the robustness of named entity recognition (NER) models, we propose RockNER, a simple yet effective method to create natural adversarial examples.Specifically, at the entity level, we replace target entities with other entities of the same semantic class in Wikidata; at the context level, we use pre-trained language models (e.g., BERT) to generate word substitutions.Together, the two levels of attack produce natural adversarial examples that result in a shifted distribution from the training data on which our target models have been trained.We apply the proposed method to the OntoNotes dataset and create a new benchmark named OntoRock for evaluating the robustness of existing NER models via a systematic evaluation protocol.Our experiments and analysis reveal that even the best model has a significant performance drop, and these models seem to memorize in-domain entity patterns instead of reasoning from the context.Our work also studies the effects of a few simple data augmentation methods to improve the robustness of NER models. 1 Bill Y. Lin, Wenyang Gao, Jun Yan 0012, Ryan Moreno, Xiang Ren 0001 |
EMNLP (1) | 5 |
| 2021 | Extract, Denoise and Enforce: Evaluating and Improving Concept Preservation for Text-to-Text GenerationabstractPrior studies on text-to-text generation typically assume that the model could figure out what to attend to in the input and what to include in the output via seq2seq learning, with only the parallel training data and no additional guidance.However, it remains unclear whether current models can preserve important concepts in the source input, as seq2seq learning does not have explicit focus on the concepts and commonly used evaluation metrics also treat concepts equally important as other tokens.In this paper, we present a systematic analysis that studies whether current seq2seq models, especially pre-trained language models, are good enough for preserving important input concepts and to what extent explicitly guiding generation with the concepts as lexical constraints is beneficial.We answer the above questions by conducting extensive analytical experiments on four representative text-to-text generation tasks.Based on the observations, we then propose a simple yet effective framework to automatically extract, denoise, and enforce important input concepts as lexical constraints.This new method performs comparably or better than its unconstrained counterpart on automatic metrics, demonstrates higher coverage for concept preservation, and receives better ratings in the human evaluation. 1 Yuning Mao, Wenchang Ma, Deren Lei, Jiawei Han 0001, Xiang Ren 0001 |
EMNLP (1) | 5 |
| 2021 | Lawyers are Dishonest? Quantifying Representational Harms in Commonsense Knowledge ResourcesabstractWarning: this paper contains content that may be offensive or upsetting.Commonsense knowledge bases (CSKB) are increasingly used for various natural language processing tasks.Since CSKBs are mostly human-generated and may reflect societal biases, it is important to ensure that such biases are not conflated with the notion of commonsense.Here we focus on two widely used CSKBs, ConceptNet and GenericsKB, and establish the presence of bias in the form of two types of representational harms, overgeneralization of polarized perceptions and representation disparity across different demographic groups in both CSKBs.Next, we find similar representational harms for downstream models that use ConceptNet.Finally, we propose a filtering-based approach for mitigating such harms, and observe that our filtered-based approach can reduce the issues in both resources and models but leads to a performance drop, leaving room for future work to build fairer and stronger commonsense models. Ninareh Mehrabi, Fred Morstatter, Jay Pujara, Xiang Ren 0001, Aram Galstyan |
EMNLP (1) | 5 |
| 2021 | Discretized Integrated Gradients for Explaining Language ModelsabstractAs a prominent attribution-based explanation algorithm, Integrated Gradients (IG) is widely adopted due to its desirable explanation axioms and the ease of gradient computation.It measures feature importance by averaging the model's output gradient interpolated along a straight-line path in the input data space.However, such straight-line interpolated points are not representative of text data due to the inherent discreteness of the word embedding space.This questions the faithfulness of the gradients computed at the interpolated points and consequently, the quality of the generated explanations.Here we propose Discretized Integrated Gradients (DIG), which allows effective attribution along non-linear interpolation paths.We develop two interpolation strategies for the discrete word embedding space that generates interpolation points that lie close to actual words in the embedding space, yielding more faithful gradient computation.We demonstrate the effectiveness of DIG over IG through experimental and human evaluations on multiple sentiment classification datasets.We provide the source code of DIG to encourage reproducible research 1 . Soumya Sanyal 0001, Xiang Ren 0001 |
EMNLP (1) | 2 |
| 2021 | CrossFit: A Few-shot Learning Challenge for Cross-task Generalization in NLPabstractHumans can learn a new language task efficiently with only few examples, by leveraging their knowledge obtained when learning prior tasks.In this paper, we explore whether and how such cross-task generalization ability can be acquired, and further applied to build better few-shot learners across diverse NLP tasks.We introduce CROSSFIT , a problem setup for studying cross-task generalization ability, which standardizes seen/unseen task partitions, data access during different learning stages, and the evaluation protocols.To instantiate different seen/unseen task partitions in CROSS-FIT and facilitate in-depth analysis, we present the NLP Few-shot Gym, a repository of 160 diverse few-shot NLP tasks created from openaccess NLP datasets and converted to a unified text-to-text format.Our analysis reveals that the few-shot learning ability on unseen tasks can be improved via an upstream learning stage using a set of seen tasks.We also observe that the selection of upstream learning tasks can significantly influence few-shot performance on unseen tasks, asking further analysis on task similarity and transferability. 1 Qinyuan Ye, Bill Y. Lin, Xiang Ren 0001 |
EMNLP (1) | 3 |
| 2021 | On the Influence of Masking Policies in Intermediate Pre-trainingabstractCurrent NLP models are predominantly trained through a two-stage "pre-train then fine-tune" pipeline.Prior work has shown that inserting an intermediate pre-training stage, using heuristic masking policies for masked language modeling (MLM), can significantly improve final performance.However, it is still unclear (1) in what cases such intermediate pre-training is helpful, (2) whether hand-crafted heuristic objectives are optimal for a given task, and (3) whether a masking policy designed for one task is generalizable beyond that task.In this paper, we perform a large-scale empirical study to investigate the effect of various masking policies in intermediate pre-training with nine selected tasks across three categories.Crucially, we introduce methods to automate the discovery of optimal masking policies via direct supervision or meta-learning.We conclude that the success of intermediate pre-training is dependent on appropriate pre-train corpus, selection of output format (i.e., masked spans or full sentence), and clear understanding of the role that MLM plays for the downstream task.In addition, we find our learned masking policies outperform the heuristic of masking named entities on TriviaQA, and policies learned from one task can positively transfer to other tasks in certain cases, inviting future research in this direction. Qinyuan Ye, Belinda Z. Li, Sinong Wang, Benjamin Bolte, Hao Ma 0001, Scott Yih, Xiang Ren 0001, Madian Khabsa |
EMNLP (1) | 7 |
| 2021 | RICA: Evaluating Robust Inference Capabilities Based on Commonsense AxiomsabstractPre-trained language models (PTLMs) have achieved impressive performance on commonsense inference benchmarks, but their ability to employ commonsense to make robust inferences, which is crucial for effective communications with humans, is debated.In the pursuit of advancing fluid human-AI communication, we propose a new challenge, RICA: Robust Inference using Commonsense Axioms, that evaluates robust commonsense inference despite textual perturbations.To generate data for this challenge, we develop a systematic and scalable procedure using commonsense knowledge bases and probe PTLMs across two different evaluation settings.Extensive experiments on our generated probe sets with more than 10k statements show that PTLMs perform no better than random guessing on the zero-shot setting, are heavily impacted by statistical biases, and are not robust to perturbation attacks.We also find that fine-tuning on similar statements offer limited gains, as PTLMs still fail to generalize to unseen inferences.Our new large-scale benchmark exposes a significant gap between PTLMs and human-level language understanding and offers a new challenge for PTLMs to demonstrate commonsense. 1 Logical TemplateRel(A,B,r) à Comp(Prop(A,p), Prop(B,p)) Rahul Khanna, Seyeon Lee, Bill Y. Lin, Daniel Ho, Jay Pujara, Xiang Ren 0001 |
EMNLP (1) | 7 |
| 2021 | Learning to Deceive Knowledge Graph Augmented Models via Targeted Perturbation
Mrigank Raman, Aaron Chan, Siddhant Agarwal, Peifeng Wang, Hansen Wang, Sungchul Kim, Ryan Rossi, Handong Zhao, Nedim Lipka, Xiang Ren 0001 |
ICLR | 10 |
| 2021 | Pre-training Text-to-Text Transformers for Concept-centric Common Sense
Wangchunshu Zhou, Ravi Kiran Selvam, Seyeon Lee, Xiang Ren 0001 |
ICLR | 5 |
| 2021 | On Transferability of Bias Mitigation Effects in Language Model Fine-TuningabstractXisen Jin, Francesco Barbieri, Brendan Kennedy, Aida Mostafazadeh Davani, Leonardo Neves, Xiang Ren. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Xisen Jin, Francesco Barbieri, Brendan Kennedy 0001, Aida Mostafazadeh Davani, Leonardo Neves, Xiang Ren 0001 |
NAACL-HLT | 6 |
| 2021 | Differentiable Open-Ended Commonsense ReasoningabstractBill Yuchen Lin, Haitian Sun, Bhuwan Dhingra, Manzil Zaheer, Xiang Ren, William Cohen. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Bill Y. Lin, Haitian Sun, Bhuwan Dhingra, Manzil Zaheer, Xiang Ren 0001, William W. Cohen |
NAACL-HLT | 5 |
| 2021 | X-METRA-ADA: Cross-lingual Meta-Transfer learning Adaptation to Natural Language Understanding and Question AnsweringabstractMeryem M’hamdi, Doo Soon Kim, Franck Dernoncourt, Trung Bui, Xiang Ren, Jonathan May. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Meryem M'hamdi, Doo Soon Kim, Franck Dernoncourt, Trung Bui, Xiang Ren 0001, Jonathan May |
NAACL-HLT | 5 |
| 2021 | TaxoClass: Hierarchical Multi-Label Text Classification Using Only Class NamesabstractJiaming Shen, Wenda Qiu, Yu Meng, Jingbo Shang, Xiang Ren, Jiawei Han. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Wenda Qiu, Yu Meng 0001, Jingbo Shang, Xiang Ren 0001, Jiawei Han 0001 |
NAACL-HLT | 5 |
| 2021 | SalKG: Learning From Knowledge Graph Explanations for Commonsense ReasoningabstractAugmenting pre-trained language models with knowledge graphs (KGs) has achieved success on various commonsense reasoning tasks. However, for a given task instance, the KG, or certain parts of the KG, may not be useful. Although KG-augmented models often use attention to focus on specific KG components, the KG is still always used, and the attention mechanism is never explicitly taught which KG components should be used. Meanwhile, saliency methods can measure how much a KG feature (e.g., graph, node, path) influences the model to make the correct prediction, thus explaining which KG features are useful. This paper explores how saliency explanations can be used to improve KG-augmented models' performance. First, we propose to create coarse (Is the KG useful?) and fine (Which nodes/paths in the KG are useful?) saliency explanations. Second, to motivate saliency-based supervision, we analyze oracle KG-augmented models which directly use saliency explanations as extra inputs for guiding their attention. Third, we propose SalKG, a framework for KG-augmented models to learn from coarse and/or fine saliency explanations. Given saliency explanations created from a task's training set, SalKG jointly trains the model to predict the explanations, then solve the task by attending to KG features highlighted by the predicted explanations. On three popular commonsense QA benchmarks (CSQA, OBQA, CODAH) and a range of KG-augmented models, we show that SalKG can yield considerable performance gains --- up to 2.76% absolute improvement on CSQA. Aaron Chan, Boyuan Long, Soumya Sanyal 0001, Tanishq Gupta, Xiang Ren 0001 |
NeurIPS | 6 |
| 2021 | Gradient-based Editing of Memory Examples for Online Task-free Continual LearningabstractWe explore task-free continual learning (CL), in which a model is trained to avoid catastrophic forgetting in the absence of explicit task boundaries or identities. Among many efforts on task-free CL, a notable family of approaches are memory-based that store and replay a subset of training examples. However, the utility of stored seen examples may diminish over time since CL models are continually updated. Here, we propose Gradient based Memory EDiting (GMED), a framework for editing stored examples in continuous input space via gradient updates, in order to create more "challenging" examples for replay. GMED-edited examples remain similar to their unedited forms, but can yield increased loss in the upcoming model updates, thereby making the future replays more effective in overcoming catastrophic forgetting. By construction, GMED can be seamlessly applied in conjunction with other memory-based CL algorithms to bring further improvement. Experiments validate the effectiveness of GMED, and our best method significantly outperforms baselines and previous state-of-the-art on five out of six datasets. Xisen Jin, Arka Sadhu, Junyi Du, Xiang Ren 0001 |
NeurIPS | 4 |
| 2021 | Refining Language Models with Compositional ExplanationsabstractPre-trained language models have been successful on text classification tasks, but are prone to learning spurious correlations from biased datasets, and are thus vulnerable when making inferences in a new domain. Prior work reveals such spurious patterns via post-hoc explanation algorithms which compute the importance of input features. Further, the model is regularized to align the importance scores with human knowledge, so that the unintended model behaviors are eliminated. However, such a regularization technique lacks flexibility and coverage, since only importance scores towards a pre-defined list of features are adjusted, while more complex human knowledge such as feature interaction and pattern generalization can hardly be incorporated. In this work, we propose to refine a learned language model for a target domain by collecting human-provided compositional explanations regarding observed biases. By parsing these explanations into executable logic rules, the human-specified refinement advice from a small set of explanations can be generalized to more training examples. We additionally introduce a regularization term allowing adjustments for both importance and interaction of features to better rectify model behavior. We demonstrate the effectiveness of the proposed approach on two text classification tasks by showing improved performance in target domain as well as improved model fairness after refinement. Huihan Yao, Qinyuan Ye, Xisen Jin, Xiang Ren 0001 |
NeurIPS | 5 |
| 2021 | Semi-automated protocol disambiguation and code generationabstractFor decades, Internet protocols have been specified using natural language. Given the ambiguity inherent in such text, it is not surprising that protocol implementations have long exhibited bugs. In this paper, we apply natural language processing (NLP) to effect semi-automated generation of protocol implementations from specification text. Our system, Sage, can uncover ambiguous or under-specified sentences in specifications; once these are clarified by the author of the protocol specification, Sage can generate protocol code automatically. Jane Yen, Tamás Lévai, Qinyuan Ye, Xiang Ren 0001, Ramesh Govindan, Barath Raghavan |
SIGCOMM | 4 |
| 2021 | Commonsense-Focused Dialogues for Response Generation: An Empirical StudyabstractPei Zhou, Karthik Gopalakrishnan, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2021. Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren 0001, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 6 |
| 2021 | Time-Series Event Prediction with Evolutionary State GraphabstractThe accurate and interpretable prediction of future events in time-series data often requires the capturing of representative patterns (or referred to as states) underpinning the observed data. To this end, most existing studies focus on the representation and recognition of states, but ignore the changing transitional relations among them. In this paper, we present evolutionary state graph, a dynamic graph structure designed to systematically represent the evolving relations (edges) among states (nodes) along time. We conduct analysis on the dynamic graphs constructed from the time-series data and show that changes on the graph structures (e.g., edges connecting certain state nodes) can inform the occurrences of events (i.e., time-series fluctuation). Inspired by this, we propose a novel graph neural network model, Evolutionary State Graph Network (EvoNet), to encode the evolutionary state graph for accurate and interpretable time-series event prediction. Specifically, EvoNet models both the node-level (state-to-state) and graph-level (segment-to-segment) propagation, and captures the node-graph (state-to-segment) interactions over time. Experimental results based on five real-world datasets show that our approach not only achieves clear improvements compared with 11 baselines, but also provides more insights towards explaining the results of event predictions. Wenjie Hu 0003, Yang Yang 0009, Ziqiang Cheng, Carl Yang 0001, Xiang Ren 0001 |
WSDM | 5 |
| 2020 | Contextualizing Hate Speech Classifiers with Post-hoc ExplanationabstractHate speech classifiers trained on imbalanced datasets struggle to determine if group identifiers like "gay" or "black" are used in offensive or prejudiced ways.Such biases manifest in false positives when these identifiers are present, due to models' inability to learn the contexts which constitute a hateful usage of identifiers.We extract post-hoc explanations from fine-tuned BERT classifiers to detect bias towards identity terms.Then, we propose a novel regularization technique based on these explanations that encourages models to learn from the context of group identifiers in addition to the identifiers themselves.Our approach improved over baselines in limiting false positives on out-of-domain data while maintaining or improving in-domain performance. Brendan Kennedy 0001, Xisen Jin, Aida Mostafazadeh Davani, Morteza Dehghani, Xiang Ren 0001 |
ACL | 5 |
| 2020 | Learning to Contextually Aggregate Multi-Source Supervision for Sequence LabelingabstractSequence labeling is a fundamental task for a range of natural language processing problems. When used in practice, its performance is largely influenced by the annotation quality and quantity, and meanwhile, obtaining ground truth labels is often costly. In many cases, ground truth labels do not exist, but noisy annotations or annotations from different domains are accessible. In this paper, we propose a novel framework Consensus Network (ConNet) that can be trained on annotations from multiple sources (e.g., crowd annotation, cross-domain data). It learns individual representation for every source and dynamically aggregates source-specific knowledge by a context-aware attention module. Finally, it leads to a model reflecting the agreement (consensus) among multiple sources. We evaluate the proposed framework in two practical settings of multi-source learning: learning with crowd annotations and unsupervised cross-domain model adaptation. Extensive experimental results show that our model achieves significant improvements over existing methods in both settings. We also demonstrate that the method can apply to various tasks and cope with different encoders. Ouyu Lan, Bill Y. Lin, Xiang Ren 0001 |
ACL | 6 |
| 2020 | TriggerNER: Learning with Entity Triggers as Explanations for Named Entity RecognitionabstractTraining neural models for named entity recognition (NER) in a new domain often requires additional human annotations that are usually expensive and time-consuming to collect.Thus, a crucial research question is how to obtain supervision in a cost-effective way.In this paper, we introduce "entity triggers," an effective proxy of human explanations for facilitating label-efficient learning of NER models.An entity trigger is defined as a group of words in a sentence that helps to explain why humans would recognize an entity in the sentence.We crowd-sourced 14k entity triggers for two well-studied NER datasets 1 .Our proposed model, Trigger Matching Network, jointly learns trigger representations and soft matching module with self-attention such that can generalize to unseen sentences easily for tagging.The framework is significantly more cost-effective than the traditional frameworks. Bill Y. Lin, Ryan Moreno, Prashant Shiralkar, Xiang Ren 0001 |
ACL | 7 |
| 2020 | Facet-Aware Evaluation for Extractive SummarizationabstractCommonly adopted metrics for extractive summarization focus on lexical overlap at the token level.In this paper, we present a facetaware evaluation setup for better assessment of the information coverage in extracted summaries.Specifically, we treat each sentence in the reference summary as a facet, identify the sentences in the document that express the semantics of each facet as support sentences of the facet, and automatically evaluate extractive summarization methods by comparing the indices of extracted sentences and support sentences of all the facets in the reference summary.To facilitate this new evaluation setup, we construct an extractive version of the CNN/Daily Mail dataset and perform a thorough quantitative investigation, through which we demonstrate that facet-aware evaluation manifests better correlation with human judgment than ROUGE, enables fine-grained evaluation as well as comparative analysis, and reveals valuable insights of state-of-the-art summarization methods. 1 1 Data can be found at https://github.com/ morningmoni/FAR.Reference: Three people in Kansas have died from a listeria outbreak.Lexical Overlap: But they did not appear identical to listeria samples taken from patients infected in the Kansas outbreak.(ROUGE-1 F1=37.0,multiple token matches but totally different semantics) Manual Extract: Five people were infected and three died in the past year in Kansas from listeria that might be linked to blue bell creameries products, according to the CDC.(ROUGE-1 F1=36.9, semantics covered but lower ROUGE due to the presence of other details) Yuning Mao, Qi Zhu 0008, Xiang Ren 0001, Jiawei Han 0001 |
ACL | 4 |
| 2020 | Generating Natural Language Adversarial Examples on a Large Scale with Generative ModelsabstractToday text classification models have been widely used. However, these classifiers are found to be easily fooled by adversarial examples. Fortunately, standard attacking methods generate adversarial texts in a pair-wise way, that is, an adversarial text can only be created from a real-world text by replacing a few words. In many applications, these texts are limited in numbers, therefore their corresponding adversarial examples are often not diverse enough and sometimes hard to read, thus can be easily detected by humans and cannot create chaos at a large scale. In this paper, we propose an end to end solution to efficiently generate adversarial texts from scratch using generative models, which are not restricted to perturbing the given texts. We call it unrestricted adversarial text generation. Specifically, we train a conditional variational autoencoder (VAE) with an additional adversarial loss to guide the generation of adversarial examples. Moreover, to improve the validity of adversarial texts, we utilize discrimators and the training framework of generative adversarial networks (GANs) to make adversarial texts consistent with real data. Experimental results on sentiment analysis demonstrate the scalability and efficiency of our method. It can attack text classification models with a higher success rate than existing methods, and provide acceptable quality for humans in the meantime. Yankun Ren, Jianbin Lin, Siliang Tang, Jun Zhou 0011, Yuan Qi 0001, Xiang Ren 0001 |
ECAI | 7 |
| 2020 | Scalable Multi-Hop Relational Reasoning for Knowledge-Aware Question AnsweringabstractExisting work that augment question answering (QA) models with external knowledge (e.g., knowledge graphs) either struggle to model multi-hop relations efficiently, or lack transparency into the model's prediction rationale.In this paper, we propose a novel knowledge-aware approach that equips pretrained language models (PTLMs) with a multi-hop relational reasoning module, named multi-hop graph relation network (MHGRN).It performs multi-hop, multi-relational reasoning over subgraphs extracted from external knowledge graphs.The proposed reasoning module unifies path-based reasoning methods and graph neural networks and results in better interpretability and scalability.We also empirically show its effectiveness and scalability on CommonsenseQA and OpenbookQA datasets, and interpret its behaviors with case studies, with the code for experiments released 1 . Yanlin Feng, Bill Y. Lin, Peifeng Wang, Jun Yan 0012, Xiang Ren 0001 |
EMNLP (1) | 6 |
| 2020 | Visually Grounded Continual Learning of Compositional PhrasesabstractHumans acquire language continually with much more limited access to data samples at a time, as compared to contemporary NLP systems.To study this human-like language acquisition ability, we present VisCOLL, a visually grounded language learning task, which simulates the continual acquisition of compositional phrases from streaming visual scenes.In the task, models are trained on a paired image-caption stream which has shifting object distribution; while being constantly evaluated by a visually-grounded masked language prediction task on held-out test sets.VisCOLL compounds the challenges of continual learning (i.e., learning from continuously shifting data distribution) and compositional generalization (i.e., generalizing to novel compositions).To facilitate research on VisCOLL, we construct two datasets, COCO-shift and Flickrshift, and benchmark them using different continual learning methods.Results reveal that SoTA continual learning approaches provide little to no improvements on VisCOLL, since storing examples of all possible compositions is infeasible.We conduct further ablations and analysis to guide future work 1 . Xisen Jin, Junyi Du, Arka Sadhu, Ramakant Nevatia, Xiang Ren 0001 |
EMNLP (1) | 5 |
| 2020 | Recurrent Event Network: Autoregressive Structure Inferenceover Temporal Knowledge GraphsabstractKnowledge graph reasoning is a critical task in natural language processing.The task becomes more challenging on temporal knowledge graphs, where each fact is associated with a timestamp.Most existing methods focus on reasoning at past timestamps and they are not able to predict facts happening in the future.This paper proposes Recurrent Event Network (RE-NET), a novel autoregressive architecture for predicting future interactions.The occurrence of a fact (event) is modeled as a probability distribution conditioned on temporal sequences of past knowledge graphs.Specifically, our RE-NET employs a recurrent event encoder to encode past facts, and uses a neighborhood aggregator to model the connection of facts at the same timestamp.Future facts can then be inferred in a sequential manner based on the two modules.We evaluate our proposed method via link prediction at future times on five public datasets.Through extensive experiments, we demonstrate the strength of RE-NET, especially on multi-step inference over future timestamps, and achieve state-of-the-art performance on all five datasets 1 . Woojeong Jin 0001, Meng Qu, Xisen Jin, Xiang Ren 0001 |
EMNLP (1) | 4 |
| 2020 | Learning Collaborative Agents with Rule Guidance for Knowledge Graph ReasoningabstractWalk-based models have shown their advantages in knowledge graph (KG) reasoning by achieving decent performance while providing interpretable decisions.However, the sparse reward signals offered by the KG during traversal are often insufficient to guide a sophisticated walk-based reinforcement learning (RL) model.An alternate approach is to use traditional symbolic methods (e.g., rule induction), which achieve good performance but can be hard to generalize due to the limitation of symbolic representation.In this paper, we propose RuleGuider, which leverages high-quality rules generated by symbolicbased methods to provide reward supervision for walk-based agents.Experiments on benchmark datasets show that RuleGuider improves the performance of walk-based models without losing interpretability. 1 Deren Lei, Gangrong Jiang, Xiaotao Gu, Kexuan Sun 0002, Yuning Mao, Xiang Ren 0001 |
EMNLP (1) | 6 |
| 2020 | Birds have four legs?! NumerSense: Probing Numerical Commonsense Knowledge of Pre-Trained Language ModelsabstractRecent works show that pre-trained language models (PTLMs), such as BERT, possess certain commonsense and factual knowledge.They suggest that it is promising to use PTLMs as "neural knowledge bases" via predicting masked words.Surprisingly, we find that this may not work for numerical commonsense knowledge (e.g., a bird usually has two legs).In this paper, we investigate whether and to what extent we can induce numerical commonsense knowledge from PTLMs as well as the robustness of this process.To study this, we introduce a novel probing task with a diagnostic dataset, NUMERSENSE 1 , containing 13.6k masked-word-prediction probes (10.5k for fine-tuning and 3.1k for testing).Our analysis reveals that: (1) BERT and its stronger variant RoBERTa perform poorly on the diagnostic dataset prior to any fine-tuning; (2) finetuning with distant supervision brings some improvement; (3) the best supervised model still performs poorly as compared to human performance (54.06% vs 96.3% in accuracy). Bill Y. Lin, Seyeon Lee, Rahul Khanna, Xiang Ren 0001 |
EMNLP (1) | 4 |
| 2020 | Multi-document Summarization with Maximal Marginal Relevance-guided Reinforcement LearningabstractWhile neural sequence learning methods have made significant progress in single-document summarization (SDS), they produce unsatisfactory results on multi-document summarization (MDS).We observe two major challenges when adapting SDS advances to MDS: (1) MDS involves larger search space and yet more limited training data, setting obstacles for neural methods to learn adequate representations; (2) MDS needs to resolve higher information redundancy among the source documents, which SDS methods are less effective to handle.To close the gap, we present RL-MMR, Maximal Margin Relevance-guided Reinforcement Learning for MDS, which unifies advanced neural SDS methods and statistical measures used in classical MDS.RL-MMR casts MMR guidance on fewer promising candidates, which restrains the search space and thus leads to better representation learning.Additionally, the explicit redundancy measure in MMR helps the neural representation of the summary to better capture redundancy.Extensive experiments demonstrate that RL-MMR achieves state-of-the-art performance on benchmark MDS datasets.In particular, we show the benefits of incorporating MMR into end-to-end learning when adapting SDS to MDS in terms of both learning effectiveness and efficiency. 1 Yuning Mao, Yanru Qu, Yiqing Xie, Xiang Ren 0001, Jiawei Han 0001 |
EMNLP (1) | 4 |
| 2020 | SynSetExpan: An Iterative Framework for Joint Entity Set Expansion and Synonym DiscoveryabstractEntity set expansion and synonym discovery are two critical NLP tasks.Previous studies accomplish them separately, without exploring their interdependences.In this work, we hypothesize that these two tasks are tightly coupled because two synonymous entities tend to have similar likelihoods of belonging to various semantic classes.This motivates us to design SynSetExpan, a novel framework that enables two tasks to mutually enhance each other.SynSetExpan uses a synonym discovery model to include popular entities' infrequent synonyms into the set, which boosts the set expansion recall.Meanwhile, the set expansion model, being able to determine whether an entity belongs to a semantic class, can generate pseudo training data to fine-tune the synonym discovery model towards better accuracy.To facilitate the research on studying the interplays of these two tasks, we create the first large-scale Synonym-Enhanced Set Expansion (SE2) dataset via crowdsourcing.Extensive experiments on the SE2 dataset and previous benchmarks demonstrate the effectiveness of SynSetExpan for both entity set expansion and synonym discovery tasks. Wenda Qiu, Jingbo Shang, Michelle Vanni, Xiang Ren 0001, Jiawei Han 0001 |
EMNLP (1) | 5 |
| 2020 | Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence Models
Xisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue 0001, Xiang Ren 0001 |
ICLR | 5 |
| 2020 | Learning from Explanations with Neural Execution Tree
Ziqi Wang 0003, Yujia Qin, Wenxuan Zhou 0002, Jun Yan 0012, Qinyuan Ye, Leonardo Neves, Zhiyuan Liu 0001, Xiang Ren 0001 |
ICLR | 8 |
| 2020 | Temporal Attribute Prediction via Joint Modeling of Multi-Relational Structure EvolutionabstractTime series prediction is an important problem in machine learning. Previous methods for time series prediction did not involve additional information. With a lot of dynamic knowledge graphs available, we can use this additional information to predict the time series better. Recently, there has been a focus on the application of deep representation learning on dynamic graphs. These methods predict the structure of the graph by reasoning over the interactions in the graph at previous time steps. In this paper, we propose a new framework to incorporate the information from dynamic knowledge graphs for time series prediction. We show that if the information contained in the graph and the time series data are closely related, then this inter-dependence can be used to predict the time series with improved accuracy. Our framework, DArtNet, learns a static embedding for every node in the graph as well as a dynamic embedding which is dependent on the dynamic attribute value (time-series). Then it captures the information from the neighborhood by taking a relation specific mean and encodes the history information using RNN. We jointly train the model link prediction and attribute prediction. We evaluate our method on five specially curated datasets for this problem and show a consistent improvement in time series prediction results. We release the data and code of model DArtNet for future research. Sankalp Garg, Navodita Sharma, Woojeong Jin 0001, Xiang Ren 0001 |
IJCAI | 4 |
| 2020 | Alleviate Dataset Shift Problem in Fine-grained Entity Typing with Virtual Adversarial TrainingabstractThe recent success of Distant Supervision (DS) brings abundant labeled data for the task of fine-grained entity typing (FET) without human annotation. However, the heuristically generated labels inevitably bring a significant distribution gap, namely dataset shift, between the distantly labeled training set and the manually curated test set. Considerable efforts have been made to alleviate this problem from the label perspective by either intelligently denoising the training labels, or designing noise-aware loss functions. Despite their progress, the dataset shift can hardly be eliminated completely. In this work, complementary to the label perspective, we reconsider this problem from the model perspective: Can we learn a more robust typing model with the existence of dataset shift? To this end, we propose a novel regularization module based on virtual adversarial training (VAT). The proposed approach first uses a self-paced sample selection function to select suitable samples for VAT, then constructs virtual adversarial perturbations based on the selected samples, and finally regularizes the model to be robust to such perturbations. Experiments on two benchmarks demonstrate the effectiveness of the proposed method, with an average 3.8%, 2.5%, and 3.2% improvement in accuracy, Macro F1 and Micro F1 respectively compared to the next best method. Siliang Tang, Xiaotao Gu, Zhigang Chen 0003, Jian Shao 0001, Xiang Ren 0001 |
IJCAI | 7 |
| 2020 | NERO: A Neural Rule Grounding Framework for Label-Efficient Relation ExtractionabstractDeep neural models for relation extraction tend to be less reliable when perfectly labeled data is limited, despite their success in label-sufficient scenarios. Instead of seeking more instance-level labels from human annotators, here we propose to annotate frequent surface patterns to form labeling rules. These rules can be automatically mined from large text corpora and generalized via a soft rule matching mechanism. Prior works use labeling rules in an exact matching fashion, which inherently limits the coverage of sentence matching and results in the low-recall issue. In this paper, we present a neural approach to ground rules for RE, named Nero, which jointly learns a relation extraction module and a soft matching module. One can employ any neural relation extraction models as the instantiation for the RE module. The soft matching module learns to match rules with semantically similar sentences such that raw corpora can be automatically labeled and leveraged by the RE module (in a much better coverage) as augmented supervision, in addition to the exactly matched sentences. Extensive experiments and analysis on two public and widely-used datasets demonstrate the effectiveness of the proposed Nero framework, comparing with both rule-based and semi-supervised methods. Through user studies, we find that the time efficiency for a human to annotate rules and sentences are similar (0.30 vs. 0.35 min per label). In particular, Nero’s performance using 270 rules is comparable to the models trained using 3,000 labeled sentences, yielding a 9.5x speedup. Moreover, Nero can predict for unseen relations at test time and provide interpretable predictions. We release our code1 to the community for future research. Wenxuan Zhou 0002, Bill Y. Lin, Ziqi Wang 0003, Junyi Du, Leonardo Neves, Xiang Ren 0001 |
WWW | 7 |
| 2020 | Dynamic network embedding via incremental skip-gram with negative sampling
Hao Peng 0001, Jianxin Li 0002, Hao Yan 0004, Qiran Gong, Senzhang Wang, Xiang Ren 0001 |
Sci. China Inf. Sci. | 8 |
| 2019 | Mining Entity Synonyms with Efficient Neural Set GenerationabstractMining entity synonym sets (i.e., sets of terms referring to the same entity) is an important task for many entity-leveraging applications. Previous work either rank terms based on their similarity to a given query term, or treats the problem as a two-phase task (i.e., detecting synonymy pairs, followed by organizing these pairs into synonym sets). However, these approaches fail to model the holistic semantics of a set and suffer from the error propagation issue. Here we propose a new framework, named SynSetMine, that efficiently generates entity synonym sets from a given vocabulary, using example sets from external knowledge bases as distant supervision. SynSetMine consists of two novel modules: (1) a set-instance classifier that jointly learns how to represent a permutation invariant synonym set and whether to include a new instance (i.e., a term) into the set, and (2) a set generation algorithm that enumerates the vocabulary only once and applies the learned set-instance classifier to detect all entity synonym sets in it. Experiments on three real datasets from different domains demonstrate both effectiveness and efficiency of SynSetMine for mining entity synonym sets. Ruiliang Lyu, Xiang Ren 0001, Michelle Vanni, Brian M. Sadler, Jiawei Han 0001 |
AAAI | 3 |
| 2019 | Cross-Relation Cross-Bag Attention for Distantly-Supervised Relation ExtractionabstractDistant supervision leverages knowledge bases to automatically label instances, thus allowing us to train relation extractor without human annotations. However, the generated training data typically contain massive noise, and may result in poor performances with the vanilla supervised learning. In this paper, we propose to conduct multi-instance learning with a novel Cross-relation Cross-bag Selective Attention (C2SA), which leads to noise-robust training for distant supervised relation extractor. Specifically, we employ the sentence-level selective attention to reduce the effect of noisy or mismatched sentences, while the correlation among relations were captured to improve the quality of attention weights. Moreover, instead of treating all entity-pairs equally, we try to pay more attention to entity-pairs with a higher quality. Similarly, we adopt the selective attention mechanism to achieve this goal. Experiments with two types of relation extractor demonstrate the superiority of the proposed approach over the state-of-the-art, while further ablation studies verify our intuitions and demonstrate the effectiveness of our proposed two techniques. Yujin Yuan, Siliang Tang, Zhongfei Zhang, Yueting Zhuang, Shiliang Pu, Fei Wu 0001, Xiang Ren 0001 |
AAAI | 8 |
| 2019 | Eliciting Knowledge from Experts: Automatic Transcript Parsing for Cognitive Task AnalysisabstractCognitive task analysis (CTA) is a type of analysis in applied psychology aimed at eliciting and representing the knowledge and thought processes of domain experts. In CTA, often heavy human labor is involved to parse the interview transcript into structured knowledge (e.g., flowchart for different actions). To reduce human efforts and scale the process, automated CTA transcript parsing is desirable. However, this task has unique challenges as (1) it requires the understanding of long-range context information in conversational text; and (2) the amount of labeled data is limited and indirect—i.e., context-aware, noisy, and low-resource. In this paper, we propose a weakly-supervised information extraction framework for automated CTA transcript parsing. We partition the parsing process into a sequence labeling task and a text span-pair relation extraction task, with distant supervision from human-curated protocol files. To model long-range context information for extracting sentence relations, neighbor sentences are involved as a part of input. Different types of models for capturing context dependency are then applied. We manually annotate real-world CTA transcripts to facilitate the evaluation of the parsing tasks. Junyi Du, Xiang Ren 0001 |
ACL (1) | 4 |
| 2019 | Distantly Supervised Biomedical Named Entity Recognition with Dictionary ExpansionabstractState-of-the-art biomedical named entity recognition (BioNER) systems apply supervised machine learning models (i.e., relying on human effort for training data annotation) which are not easy to be generalized to new entity types and datasets. We propose a distantly supervised approach, AutoBioNER, that automatically recognizes biomedical entities from massive corpora with user-input dictionaries. AutoBioNER does not need any human annotated data. It relies on incomplete entity dictionaries to provide seeds for each entity type and performs a novel entity set expansion step for corpus-level new entity recognition and dictionary completion. The expanded dictionaries are used as distant supervision to train a neural model for BioNER. Experimental results show that AutoBioNER achieves the best performance among the methods that only use dictionaries with no additional human effort on BioNER benchmark datasets. It is also demonstrated that the dictionary expansion step plays an important role in the great performances. Xuan Wang 0008, Yu Zhang 0044, Qi Li 0012, Xiang Ren 0001, Jingbo Shang, Jiawei Han 0001 |
BIBM | 4 |
| 2019 | Mining News Events from Comparable News Corpora: A Multi-Attribute Proximity Network Modeling ApproachabstractWe present ProxiModel, a novel event mining framework for extracting high-quality structured event knowledge from large, redundant, and noisy news data sources. The proposed model differentiates itself from other approaches by modeling both the event correlation within each individual document as well as across the corpus. To facilitate this, we introduce the concept of a proximity-network, a novel space-efficient data structure to facilitate scalable event mining. This proximity network captures the corpus-level co-occurence statistics for candidate event descriptors, event attributes, as well as their connections. We probabilistically model the proximity network as a generative process with sparsity-inducing regularization. This allows us to efficiently and effectively extract high-quality and interpretable news events. Experiments on three different news corpora demonstrate that the proposed method is effective and robust at generating high-quality event descriptors and attributes. We briefly detail many interesting applications from our proposed framework such as news summarization, event tracking and multi-dimensional analysis on news. Finally, we explore a case study on visualizing the events for a Japan Tsunami news corpus and demonstrate ProxiModel's ability to automatically summarize emerging news events. Hyungsul Kim, Ahmed El-Kishky, Xiang Ren 0001, Jiawei Han 0001 |
IEEE BigData | 3 |
| 2019 | Reporting the Unreported: Event Extraction for Analyzing the Local Representation of Hate CrimesabstractAida Mostafazadeh Davani, Leigh Yeh, Mohammad Atari, Brendan Kennedy, Gwenyth Portillo Wightman, Elaine Gonzalez, Natalie Delong, Rhea Bhatia, Arineh Mirinjian, Xiang Ren, Morteza Dehghani. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Aida Mostafazadeh Davani, Leigh Yeh, Mohammad Atari, Brendan Kennedy 0001, Gwenyth Portillo-Wightman, Elaine Gonzalez, Natalie Delong, Rhea Bhatia, Arineh Mirinjian, Xiang Ren 0001, Morteza Dehghani |
EMNLP/IJCNLP (1) | 10 |
| 2019 | Collaborative Policy Learning for Open Knowledge Graph ReasoningabstractCong Fu, Tong Chen, Meng Qu, Woojeong Jin, Xiang Ren. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Cong Fu 0001, Meng Qu, Woojeong Jin 0001, Xiang Ren 0001 |
EMNLP/IJCNLP (1) | 5 |
| 2019 | KagNet: Knowledge-Aware Graph Networks for Commonsense ReasoningabstractBill Yuchen Lin, Xinyue Chen, Jamin Chen, Xiang Ren. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Bill Y. Lin, Jamin Chen, Xiang Ren 0001 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Hierarchical Text Classification with Reinforced Label AssignmentabstractYuning Mao, Jingjing Tian, Jiawei Han, Xiang Ren. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yuning Mao, Jiawei Han 0001, Xiang Ren 0001 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | HMEAE: Hierarchical Modular Event Argument ExtractionabstractXiaozhi Wang, Ziqi Wang, Xu Han, Zhiyuan Liu, Juanzi Li, Peng Li, Maosong Sun, Jie Zhou, Xiang Ren. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xiaozhi Wang, Ziqi Wang 0003, Xu Han 0007, Zhiyuan Liu 0001, Juan-Zi Li, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016, Xiang Ren 0001 |
EMNLP/IJCNLP (1) | 9 |
| 2019 | Learning Dynamic Context Augmentation for Global Entity LinkingabstractXiyuan Yang, Xiaotao Gu, Sheng Lin, Siliang Tang, Yueting Zhuang, Fei Wu, Zhigang Chen, Guoping Hu, Xiang Ren. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xiyuan Yang, Xiaotao Gu, Siliang Tang, Yueting Zhuang, Fei Wu 0001, Zhigang Chen 0003, Xiang Ren 0001 |
EMNLP/IJCNLP (1) | 9 |
| 2019 | Looking Beyond Label Noise: Shifted Label Distribution Matters in Distantly Supervised Relation ExtractionabstractQinyuan Ye, Liyuan Liu, Maosen Zhang, Xiang Ren. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Qinyuan Ye, Maosen Zhang, Xiang Ren 0001 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | AtSNE: Efficient and Robust Visualization on GPU through Hierarchical OptimizationabstractVisualization of high-dimensional data is a fundamental yet challenging problem in data mining. These visualization techniques are commonly used to reveal the patterns in the high-dimensional data, such as clusters and the similarity among clusters. Recently, some successful visualization tools (e.g., BH-t-SNE and LargeVis) have been developed. However, there are two limitations with them : (1) they cannot capture the global data structure well. Thus, their visualization results are sensitive to initialization, which may cause confusions to the data analysis. (2) They cannot scale to large-scale datasets. They are not suitable to be implemented on the GPU platform because their complex algorithm logic, high memory cost, and random memory access mode will lead to low hardware utilization. To address the aforementioned problems, we propose a novel visualization approach named as Anchor-t-SNE (AtSNE), which provides efficient GPU-based visualization solution for large-scale and high-dimensional data. Specifically, we generate a number of anchor points from the original data and regard them as the skeleton of the layout, which holds the global structure information. We propose a hierarchical optimization approach to optimize the positions of the anchor points and ordinary data points in the layout simultaneously. Our approach presents much better and robust visual effects on 11 public datasets, and achieve 5 to 28 times speed-up on different datasets, compared with the current state-of-the-art methods. In particular, we deliver a high-quality 2-D layout for a 20 million and 96-dimension dataset within 5 hours, while the current methods fail to give results due to running out of the memory. Cong Fu 0001, Deng Cai 0001, Xiang Ren 0001 |
KDD | 4 |
| 2019 | Characterizing and Forecasting User Engagement with In-App Action Graph: A Case Study of SnapchatabstractWhile mobile social apps have become increasingly important in people's daily life, we have limited understanding on what motivates users to engage with these apps. In this paper, we answer the question whether users' in-app activity patterns help inform their future app engagement (e.g., active days in a future time window)? Previous studies on predicting user app engagement mainly focus on various macroscopic features (e.g., time-series of activity frequency), while ignoring fine-grained inter-dependencies between different in-app actions at the microscopic level. Here we propose to formalize individual user's in-app action transition patterns as a temporally evolving action graph, and analyze its characteristics in terms of informing future user engagement. Our analysis suggested that action graphs are able to characterize user behavior patterns and inform future engagement. We derive a number of high-order graph features to capture in-app usage patterns and construct interpretable models for predicting trends of engagement changes and active rates. To further enhance predictive power, we design an end-to-end, multi-channel neural model to encode both temporal action graphs, activity sequences, and other macroscopic features. Experiments on predicting user engagement for 150k Snapchat new users over a 28-day period demonstrate the effectiveness of the proposed prediction models. The analysis and prediction framework is also deployed at Snapchat to deliver real world business insights. Our proposed framework is also general and can be applied to any online platform. Yozen Liu, Lucas Pierce, Xiang Ren 0001 |
KDD | 4 |
| 2019 | Integrating Local Context and Global Cohesiveness for Open Information ExtractionabstractExtracting entities and their relations from text is an important task for understanding massive text corpora. Open information extraction (IE) systems mine relation tuples (i.e., entity arguments and a predicate string to describe their relation) from sentences. These relation tuples are not confined to a predefined schema for the relations of interests. However, current Open IE systems focus on modeling local context information in a sentence to extract relation tuples, while ignoring the fact that global statistics in a large corpus can be collectively leveraged to identify high-quality sentence-level extractions. In this paper, we propose a novel Open IE system, called ReMine, which integrates local context signals and global structural signals in a unified, distant-supervision framework. Leveraging facts from external knowledge bases as supervision, the new system can be applied to many different domains to facilitate sentence-level tuple extractions using corpus-level statistics. Our system operates by solving a joint optimization problem to unify (1) segmenting entity/relation phrases in individual sentences based on local context; and (2) measuring the quality of tuples extracted from individual sentences with a translating-based objective. Learning the two subtasks jointly helps correct errors produced in each subtask so that they can mutually enhance each other. Experiments on two real-world corpora from different domains demonstrate the effectiveness, generality, and robustness of ReMine when compared to state-of-the-art open IE systems. Qi Zhu 0008, Xiang Ren 0001, Jingbo Shang, Yu Zhang 0044, Ahmed El-Kishky, Jiawei Han 0001 |
WSDM | 2 |
| 2019 | Learning Dual Retrieval Module for Semi-supervised Relation ExtractionabstractRelation extraction is an important task in structuring content of text data, and becomes especially challenging when learning with weak supervision-where only a limited number of labeled sentences are given and a large number of unlabeled sentences are available. Most existing work exploits unlabeled data based on the ideas of self-training (i.e., bootstrapping a model) and self-ensembling (e.g., ensembling multiple model variants). However, these methods either suffer from the issue of semantic drift, or do not fully capture the problem characteristics of relation extraction. In this paper, we leverage a key insight that retrieving sentences expressing a relation is a dual task of predicting the relation label for a given sentence-two tasks are complementary to each other and can be optimized jointly for mutual enhancement. To model this intuition, we propose DualRE, a principled framework that introduces a retrieval module which is jointly trained with the original relation prediction module. In this way, high-quality samples selected by the retrieval module from unlabeled data can be used to improve the prediction module, and vice versa. Experimental results1 on two public datasets as well as case studies demonstrate the effectiveness of the DualRE approach. Jun Yan 0012, Meng Qu, Xiang Ren 0001 |
WWW | 4 |
| 2019 | Jointly Learning Explainable Rules for Recommendation with Knowledge GraphabstractExplainability and effectiveness are two key aspects for building recommender systems. Prior efforts mostly focus on incorporating side information to achieve better recommendation performance. However, these methods have some weaknesses: (1) prediction of neural network-based embedding methods are hard to explain and debug; (2) symbolic, graph-based approaches (e.g., meta path-based models) require manual efforts and domain knowledge to define patterns and rules, and ignore the item association types (e.g. substitutable and complementary). In this paper, we propose a novel joint learning framework to integrate induction of explainable rules from knowledge graph with construction of a rule-guided neural recommendation model. The framework encourages two modules to complement each other in generating effective and explainable recommendation: 1) inductive rules, mined from item-centric knowledge graphs, summarize common multi-hop relational patterns for inferring different item associations and provide human-readable explanation for model prediction; 2) recommendation module can be augmented by induced rules and thus have better generalization ability dealing with the cold-start issue. Extensive experiments1 show that our proposed method has achieved significant improvements in item recommendation over baselines on real-world datasets. Our model demonstrates robust performance over “noisy” item knowledge graphs, generated by linking item names to related entities. Weizhi Ma, Min Zhang 0006, Woojeong Jin 0001, Chenyang Wang 0003, Yiqun Liu 0001, Shaoping Ma, Xiang Ren 0001 |
WWW | 8 |
| 2019 | Cross-type biomedical named entity recognition with deep multi-task learningabstractMOTIVATION: State-of-the-art biomedical named entity recognition (BioNER) systems often require handcrafted features specific to each entity type, such as genes, chemicals and diseases. Although recent studies explored using neural network models for BioNER to free experts from manual feature engineering, the performance remains limited by the available training data for each entity type. RESULTS: We propose a multi-task learning framework for BioNER to collectively use the training data of different types of entities and improve the performance on each of them. In experiments on 15 benchmark BioNER datasets, our multi-task model achieves substantially better performance compared with state-of-the-art BioNER systems and baseline neural sequence labeling models. Further analysis shows that the large performance gains come from sharing character- and word-level information among relevant biomedical entities across differently labeled corpora. AVAILABILITY AND IMPLEMENTATION: Our source code is available at https://github.com/yuzhimanhua/lm-lstm-crf. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xuan Wang 0008, Yu Zhang 0044, Xiang Ren 0001, Yuhao Zhang 0004, Marinka Zitnik, Jingbo Shang, Curt Langlotz, Jiawei Han 0001 |
Bioinform. | 3 |
| 2018 | Empower Sequence Labeling with Task-Aware Neural Language ModelabstractLinguistic sequence labeling is a general approach encompassing a variety of problems, such as part-of-speech tagging and named entity recognition. Recent advances in neural networks (NNs) make it possible to build reliable models without handcrafted features. However, in many cases, it is hard to obtain sufficient annotations to train these models. In this study, we develop a neural framework to extract knowledge from raw texts and empower the sequence labeling task. Besides word-level knowledge contained in pre-trained word embeddings, character-aware neural language models are incorporated to extract character-level knowledge. Transfer learning techniques are further adopted to mediate different components and guide the language model towards the key knowledge. Comparing to previous methods, these task-specific knowledge allows us to adopt a more concise model and conduct more efficient training. Different from most transfer learning methods, the proposed framework does not rely on any additional supervision. It extracts knowledge from self-contained order information of training sequences. Extensive experiments on benchmark datasets demonstrate the effectiveness of leveraging character-level knowledge and the efficiency of co-training. For example, on the CoNLL03 NER task, model training completes in about 6 hours on a single GPU, reaching F_1 score of 91.71+/-0.10 without using any extra annotations. Jingbo Shang, Xiang Ren 0001, Frank F. Xu, Huan Gui, Jian Peng 0001, Jiawei Han 0001 |
AAAI | 3 |
| 2018 | Dynamic Network Embedding by Modeling Triadic Closure ProcessabstractNetwork embedding, which aims to learn the low-dimensional representations of vertices, is an important task and has attracted considerable research efforts recently. In real world, networks, like social network and biological networks, are dynamic and evolving over time. However, almost all the existing network embedding methods focus on static networks while ignore network dynamics. In this paper, we present a novel representation learning approach, DynamicTriad, to preserve both structural information and evolution patterns of a given network. The general idea of our approach is to impose triad, which is a group of three vertices and is one of the basic units of networks. In particular, we model how a closed triad, which consists of three vertices connected with each other, develops from an open triad that has two of three vertices not connected with each other. This triadic closure process is a fundamental mechanism in the formation and evolution of networks, thereby makes our model being able to capture the network dynamics and to learn representation vectors for each vertex at different time steps. Experimental results on three real-world networks demonstrate that, compared with several state-of-the-art techniques, DynamicTriad achieves substantial gains in several application scenarios. For example, our approach can effectively be applied and help to identify telephone frauds in a mobile network, and to predict whether a user will repay her loans or not in a loan network. Le-kui Zhou, Yang Yang 0009, Xiang Ren 0001, Fei Wu 0001, Yueting Zhuang |
AAAI | 3 |
| 2018 | End-to-End Reinforcement Learning for Automatic Taxonomy InductionabstractWe present a novel end-to-end reinforcement learning approach to automatic taxonomy induction from a set of terms.While prior methods treat the problem as a two-phase task (i.e., detecting hypernymy pairs followed by organizing these pairs into a tree-structured hierarchy), we argue that such two-phase methods may suffer from error propagation, and cannot effectively optimize metrics that capture the holistic structure of a taxonomy.In our approach, the representations of term pairs are learned using multiple sources of information and used to determine which term to select and where to place it on the taxonomy via a policy network.All components are trained in an end-to-end manner with cumulative rewards, measured by a holistic tree metric over the training taxonomies.Experiments on two public datasets of different domains show that our approach outperforms prior state-ofthe-art taxonomy induction methods up to 19.6% on ancestor F1. 1 Yuning Mao, Xiang Ren 0001, Xiaotao Gu, Jiawei Han 0001 |
ACL (1) | 2 |
| 2018 | Open-Schema Event Profiling for Massive News CorporaabstractWith the rapid growth of online information services, a sheer volume of news data becomes available. To help people quickly digest the explosive information, we define a new problem - schema-based news event profiling - profiling events reported in open-domain news corpora, with a set of slots and slot-value pairs for each event, where the set of slots forms the schema of an event type. Such profiling not only provides readers with concise views of events, but also facilitates various applications such as information retrieval, knowledge graph construction and question answering. It is however a quite challenging task. The first challenge is to find out events and event types because they are both initially unknown. The second difficulty is the lack of pre-defined event-type schemas. Lastly, even with the schemas extracted, to generate event profiles from them is still essential yet demanding. Quan Yuan 0001, Xiang Ren 0001, Wenqi He, Chao Zhang 0014, Xinhe Geng, Lifu Huang, Heng Ji 0001, Chin-Yew Lin, Jiawei Han 0001 |
CIKM | 2 |
| 2018 | Efficient Contextualized Representation: Language Model Pruning for Sequence LabelingabstractMany efforts have been made to facilitate natural language processing tasks with pre-trained language models (LMs), and brought significant improvements to various applications.To fully leverage the nearly unlimited corpora and capture linguistic information of multifarious levels, large-size LMs are required; but for a specific task, only parts of these information are useful.Such large-sized LMs, even in the inference stage, may cause heavy computation workloads, making them too time-consuming for large-scale applications.Here we propose to compress bulky LMs while preserving useful information with regard to a specific task.As different layers of the model keep different information, we develop a layer selection method for model pruning using sparsityinducing regularization.By introducing the dense connectivity, we can detach any layer without affecting others, and stretch shallow and wide LMs to be deep and narrow.In model training, LMs are learned with layerwise dropouts for better robustness.Experiments on two benchmark datasets demonstrate the effectiveness of our method. Xiang Ren 0001, Jingbo Shang, Xiaotao Gu, Jian Peng 0001, Jiawei Han 0001 |
EMNLP | 2 |
| 2018 | Learning Named Entity Tagger using Domain-Specific DictionaryabstractRecent advances in deep neural models allow us to build reliable named entity recognition (NER) systems without handcrafting features.However, such methods require large amounts of manually-labeled training data.There have been efforts on replacing human annotations with distant supervision (in conjunction with external dictionaries), but the generated noisy labels pose significant challenges on learning effective neural models.Here we propose two neural models to suit noisy distant supervision from the dictionary.First, under the traditional sequence labeling framework, we propose a revised fuzzy CRF layer to handle tokens with multiple possible labels.After identifying the nature of noisy labels in distant supervision, we go beyond the traditional framework and propose a novel, more effective neural model AutoNER with a new Tie or Break scheme.In addition, we discuss how to refine distant supervision for better NER performance.Extensive experiments on three benchmark datasets demonstrate that AutoNER achieves the best performance when only using dictionaries with no additional human effort, and delivers competitive results with state-of-the-art supervised benchmarks. Jingbo Shang, Xiaotao Gu, Xiang Ren 0001, Teng Ren, Jiawei Han 0001 |
EMNLP | 4 |
| 2018 | GraphRNN: Generating Realistic Graphs with Deep Auto-regressive ModelsabstractModeling and generating graphs is fundamental for studying networks in biology, engineering, and social sciences. However, modeling complex distributions over graphs and then efficiently sampling from these distributions is challenging due to the non-unique, high-dimensional nature of graphs and the complex, non-local dependencies that exist between edges in a given graph. Here we propose GraphRNN, a deep autoregressive model that addresses the above challenges and approximates any distribution of graphs with minimal assumptions about their structure. GraphRNN learns to generate graphs by training on a representative set of graphs and decomposes the graph generation process into a sequence of node and edge formations, conditioned on the graph structure generated so far. In order to quantitatively evaluate the performance of GraphRNN, we introduce a benchmark suite of datasets, baselines and novel evaluation metrics based on Maximum Mean Discrepancy, which measure distances between sets of graphs. Our experiments show that GraphRNN significantly outperforms all baselines, learning to generate diverse graphs that match the structural characteristics of a target set, while also scaling to graphs 50 times larger than previous deep models. Jiaxuan You, Rex Ying, Xiang Ren 0001, William L. Hamilton, Jure Leskovec |
ICML | 3 |
| 2018 | HiExpan: Task-Guided Taxonomy Construction by Hierarchical Tree ExpansionabstractTaxonomies are of great value to many knowledge-rich applications. As the manual taxonomy curation costs enormous human effects, automatic taxonomy construction is in great demand. However, most existing automatic taxonomy construction methods can only build hypernymy taxonomies wherein each edge is limited to expressing the is-a relation. Such a restriction limits their applicability to more diverse real-world tasks where the parent-child may carry different relations. In this paper, we aim to construct a task-guided taxonomy from a domain-specific corpus, and allow users to input a seed taxonomy, serving as the task guidance. We propose an expansion-based taxonomy construction framework, namely HiExpan, which automatically generates key term list from the corpus and iteratively grows the seed taxonomy. Specifically, HiExpan views all children under each taxonomy node forming a coherent set and builds the taxonomy by recursively expanding all these sets. Furthermore, HiExpan incorporates a weakly-supervised relation extraction module to extract the initial children of a newly-expanded node and adjusts the taxonomy tree by optimizing its global structure. Our experiments on three real datasets from different domains demonstrate the effectiveness of HiExpan for building task-guided taxonomies. Zeqiu Wu, Dongming Lei, Chao Zhang 0014, Xiang Ren 0001, Michelle Vanni, Brian M. Sadler, Jiawei Han 0001 |
KDD | 5 |
| 2018 | Hierarchical Graph Representation Learning with Differentiable PoolingabstractRecently, graph neural networks (GNNs) have revolutionized the field of graph representation learning through effectively learned node embeddings, and achieved state-of-the-art results in tasks such as node classification and link prediction. However, current GNN methods are inherently flat and do not learn hierarchical representations of graphs---a limitation that is especially problematic for the task of graph classification, where the goal is to predict the label associated with an entire graph. Here we propose DiffPool, a differentiable graph pooling module that can generate hierarchical representations of graphs and can be combined with various graph neural network architectures in an end-to-end fashion. DiffPool learns a differentiable soft cluster assignment for nodes at each layer of a deep GNN, mapping nodes to a set of clusters, which then form the coarsened input for the next GNN layer. Our experimental results show that combining existing GNN methods with DiffPool yields an average improvement of 5-10% accuracy on graph classification benchmarks, compared to all existing pooling approaches, achieving a new state-of-the-art on four out of five benchmark datasets. Rex Ying, Jiaxuan You, Christopher Morris 0001, Xiang Ren 0001, William L. Hamilton, Jure Leskovec |
NeurIPS | 4 |
| 2018 | First Workshop on Knowledge Base Construction, Mining and ReasoningabstractNo abstract available. Xiang Ren 0001, Craig A. Knoblock, William Yang Wang, Yu Su 0001 |
WSDM | 1 |
| 2018 | Indirect Supervision for Relation Extraction using Question-Answer PairsabstractAutomatic relation extraction (E)for types of interest is of great importance for interpreting massive text corpora in an efficient manner. For example, we want to identify the relationship "president_of" between entities "Donald Trump" and "United States" in a sentence expressing such a relation. Traditional RE models have heavily relied on human-annotated corpus for training, which can be costly in generating labeled data and become obstacles when dealing with more relation types. Thus, more RE extraction systems have shifted to be built upon training data automatically acquired by linking to knowledge bases (distant supervision). However, due to the incompleteness of knowledge bases and the context-agnostic labeling, the training data collected via distant supervision (DS) can be very noisy. In recent years, as increasing attention has been brought to tackling question-answering (QA) tasks, user feedback or datasets of such tasks become more accessible. In this paper, we propose a novel framework, ReQuest, to leverage question-answer pairs as an indirect source of supervision for relation extraction, and study how to use such supervision to reduce noise induced from DS. Our model jointly embeds relation mentions, types, QA entity mention pairs and text features in two low-dimensional spaces (RE and QA), where objects with same relation types or semantically similar question-answer pairs have similar representations. Shared features connect these two spaces, carrying clearer semantic knowledge from both sources. ReQuest, then use these learned embeddings to estimate the types of test relation mentions. We formulate a global objective function and adopt a novel margin-based QA loss to reduce noise in DS by exploiting semantic evidence from the QA dataset. Our experimental results achieve an average of 11% improvement in F1 score on two public RE datasets combined with TREC QA dataset. Codes and datasets can be downloaded at https://github.com/ellenmellon/ReQuest. Zeqiu Wu, Xiang Ren 0001, Frank F. Xu, Jiawei Han 0001 |
WSDM | 2 |
| 2018 | Weakly-supervised Relation Extraction by Pattern-enhanced Embedding LearningabstractExtracting relations from text corpora is an important task with wide applications. However, it becomes particularly challenging when focusing on weakly-supervised relation extraction, that is, utilizing a few relation instances (i.e., a pair of entities and their relation) as seeds to extract from corpora more instances of the same relation. Existing distributional approaches leverage the corpus-level co-occurrence statistics of entities to predict their relations, and require a large number of labeled instances to learn effective relation classifiers. Alternatively, pattern-based approaches perform boostrapping or apply neural networks to model the local contexts, but still rely on a large number of labeled instances to build reliable models. In this paper, we study the integration of distributional and pattern-based methods in a weakly-supervised setting such that the two kinds of methods can provide complementary supervision for each other to build an effective, unified model. We propose a novel co-training framework with a distributional module and a pattern module. During training, the distributional module helps the pattern module discriminate between the informative patterns and other patterns, and the pattern module generates some highly-confident instances to improve the distributional module. The whole framework can be effectively optimized by iterating between improving the pattern module and updating the distributional module. We conduct experiments on two tasks: knowledge base completion with text corpora and corpus-level relation extraction. Experimental results prove the effectiveness of our framework over many competitive baselines. Meng Qu, Xiang Ren 0001, Yu Zhang 0044, Jiawei Han 0001 |
WWW | 2 |
| 2018 | Automated Phrase Mining from Massive Text CorporaabstractAs one of the fundamental tasks in text analysis, phrase mining aims at extracting quality phrases from a text corpus and has various downstream applications including information extraction/retrieval, taxonomy construction, and topic modeling. Most existing methods rely on complex, trained linguistic analyzers, and thus likely have unsatisfactory performance on text corpora of new domains and genres without extra but expensive adaption. None of the state-of-the-art models, even data-driven models, is fully automated because they require human experts for designing rules or labeling phrases. In this paper, we propose a novel framework for automated phrase mining, AutoPhrase, which supports any language as long as a general knowledge base (e.g., Wikipedia) in that language is available, while benefiting from, but not requiring, a POS tagger. Compared to the state-of-the-art methods, AutoPhrase has shown significant improvements in both effectiveness and efficiency on five real-world datasets across different domains and languages. Besides, AutoPhrase can be extended to model single-word quality phrases. Jingbo Shang, Meng Jiang 0001, Xiang Ren 0001, Clare R. Voss, Jiawei Han 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2017 | An Attention-based Collaboration Framework for Multi-View Network Representation LearningabstractLearning distributed node representations in networks has been attracting increasing attention recently due to its effectiveness in a variety of applications. Existing approaches usually study networks with a single type of proximity between nodes, which defines a single view of a network. However, in reality there usually exists multiple types of proximities between nodes, yielding networks with multiple views. This paper studies learning node representations for networks with multiple views, which aims to infer robust node representations across different views. We propose a multi-view representation learning approach, which promotes the collaboration of different views and lets them vote for the robust representations. During the voting process, an attention mechanism is introduced, which enables each node to focus on the most informative views. Experimental results on real-world networks show that the proposed approach outperforms existing state-of-the-art approaches for network representation learning with a single view and other competitive approaches with multiple views. Meng Qu, Jian Tang 0005, Jingbo Shang, Xiang Ren 0001, Ming Zhang 0004, Jiawei Han 0001 |
CIKM | 4 |
| 2017 | Heterogeneous Supervision for Relation Extraction: A Representation Learning ApproachabstractRelation extraction is a fundamental task in information extraction.Most existing methods have heavy reliance on annotations labeled by human experts, which are costly and time-consuming.To overcome this drawback, we propose a novel framework, REHESSION, to conduct relation extractor learning using annotations from heterogeneous information source, e.g., knowledge base and domain heuristics.These annotations, referred as heterogeneous supervision, often conflict with each other, which brings a new challenge to the original relation extraction task: how to infer the true label from noisy labels for a given instance.Identifying context information as the backbone of both relation extraction and true label discovery, we adopt embedding techniques to learn the distributed representations of context, which bridges all components with mutual enhancement in an iterative fashion.Extensive experimental results demonstrate the superiority of REHESSION over the state-of-the-art. Xiang Ren 0001, Qi Zhu 0008, Shi Zhi, Huan Gui, Heng Ji 0001, Jiawei Han 0001 |
EMNLP | 2 |
| 2017 | MetaPAD: Meta Pattern Discovery from Massive Text CorporaabstractMining textual patterns in news, tweets, papers, and many other kinds of text corpora has been an active theme in text mining and NLP research. Previous studies adopt a dependency parsing-based pattern discovery approach. However, the parsing results lose rich context around entities in the patterns, and the process is costly for a corpus of large scale. In this study, we propose a novel typed textual pattern structure, called meta pattern, which is extended to a frequent, informative, and precise subsequence pattern in certain context. We propose an efficient framework, called MetaPAD, which discovers meta patterns from massive corpora with three techniques: (1) it develops a context-aware segmentation method to carefully determine the boundaries of patterns with a learnt pattern quality assessment function, which avoids costly dependency parsing and generates high-quality patterns; (2) it identifies and groups synonymous meta patterns from multiple facets---their types, contexts, and extractions; and (3) it examines type distributions of entities in the instances extracted by each group of patterns, and looks for appropriate type levels to make discovered patterns precise. Experiments demonstrate that our proposed framework discovers high-quality typed textual patterns efficiently from different genres of massive corpora and facilitates information extraction. Meng Jiang 0001, Jingbo Shang, Taylor Cassidy, Xiang Ren 0001, Lance M. Kaplan, Tim Hanratty, Jiawei Han 0001 |
KDD | 4 |
| 2017 | Automatic Synonym Discovery with Knowledge BasesabstractRecognizing entity synonyms from text has become a crucial task in many entity-leveraging applications. However, discovering entity synonyms from domain-specific text corpora (e.g., news articles, scientific papers) is rather challenging. Current systems take an entity name string as input to find out other names that are synonymous, ignoring the fact that often times a name string can refer to multiple entities (e.g., "apple" could refer to both Apple Inc. and the fruit apple). Moreover, most existing methods require training data manually created by domain experts to construct supervised-learning systems. In this paper, we study the problem of automatic synonym discovery with knowledge bases, that is, identifying synonyms for knowledge base entities in a given domain-specific corpus. The manually-curated synonyms for each entity stored in a knowledge base not only form a set of name strings to disambiguate the meaning for each other, but also can serve as "distant" supervision to help determine important features for the task. We propose a novel framework, called DPE, to integrate two kinds of mutually-complementing signals for synonym discovery, i.e., distributional features based on corpus-level statistics and textual patterns based on local contexts. In particular, DPE jointly optimizes the two kinds of signals in conjunction with distant supervision, so that they can mutually enhance each other in the training stage. At the inference stage, both signals will be utilized to discover synonyms for the given entities. Experimental results prove the effectiveness of the proposed framework. Meng Qu, Xiang Ren 0001, Jiawei Han 0001 |
KDD | 2 |
| 2017 | SetExpan: Corpus-Based Set Expansion via Context Feature Selection and Rank Ensemble
Zeqiu Wu, Dongming Lei, Jingbo Shang, Xiang Ren 0001, Jiawei Han 0001 |
ECML/PKDD (1) | 5 |
| 2017 | Building Structured Databases of Factual Knowledge from Massive Text CorporaabstractIn today's computerized and information-based society, people are inundated with vast amounts of text data, ranging from news articles, social media post, scientific publications, to a wide range of textual information from various domains (corporate reports, advertisements, legal acts, medical reports). To turn such massive unstructured text data into structured, actionable knowledge, one of the grand challenges is to gain an understanding of the factual information (e.g., entities, attributes, relations) in the text. Xiang Ren 0001, Meng Jiang 0001, Jingbo Shang, Jiawei Han 0001 |
SIGMOD Conference | 1 |
| 2017 | Comparative Document Analysis for Large Text CorporaabstractThis paper presents a novel research problem, Comparative Document Analysis (CDA), that is, joint discovery of commonalities and differences between two individual documents (or two sets of documents) in a large text corpus. Given any pair of documents from a (background) document collection, CDA aims to automatically identify sets of quality phrases to summarize the commonalities of both documents and highlight the distinctions of each with respect to the other informatively and concisely. Our solution uses a general graph-based framework to derive novel measures on phrase semantic commonality and pairwise distinction, where the background corpus is used for computing phrase-document semantic relevance. We use the measures to guide the selection of sets of phrases by solving two joint optimization problems. A scalable iterative algorithm is developed to integrate the maximization of phrase commonality or distinction measure with the learning of phrase-document semantic relevance. Experiments on large text corpora from two different domains---scientific papers and news---demonstrate the effectiveness and robustness of the proposed framework on comparing documents. Analysis on a 10GB+ text corpus demonstrates the scalability of our method, whose computation time grows linearly as the corpus size increases. Our case study on comparing news articles published at different dates shows the power of the proposed method on comparing sets of documents. Xiang Ren 0001, Yuanhua Lv, Kuansan Wang, Jiawei Han 0001 |
WSDM | 1 |
| 2017 | CoType: Joint Extraction of Typed Entities and Relations with Knowledge BasesabstractExtracting entities and relations for types of interest from text is important for understanding massive text corpora. Traditionally, systems of entity relation extraction have relied on human-annotated corpora for training and adopted an incremental pipeline. Such systems require additional human expertise to be ported to a new domain, and are vulnerable to errors cascading down the pipeline. In this paper, we investigate joint extraction of typed entities and relations with labeled data heuristically obtained from knowledge bases (i.e., distant supervision). As our algorithm for type labeling via distant supervision is context-agnostic, noisy training data poses unique challenges for the task. We propose a novel domain-independent framework, called CoType, that runs a data-driven text segmentation algorithm to extract entity mentions, and jointly embeds entity mentions, relation mentions, text features and type labels into two low-dimensional spaces (for entity and relation mentions respectively), where, in each space, objects whose types are close will also have similar representations. CoType, then using these learned embeddings, estimates the types of test (unlinkable) mentions. We formulate a joint optimization problem to learn embeddings from text corpora and knowledge bases, adopting a novel partial-label loss function for noisy labeled data and introducing an object "translation" function to capture the cross-constraints of entities and relations on each other. Experiments on three public datasets demonstrate the effectiveness of CoType across different domains (e.g., news, biomedical), with an average of 25% improvement in F1 score compared to the next best method. Xiang Ren 0001, Zeqiu Wu, Wenqi He, Meng Qu, Clare R. Voss, Heng Ji 0001, Tarek F. Abdelzaher, Jiawei Han 0001 |
WWW | 1 |
| 2016 | Precision Matrix Estimation in High Dimensional Gaussian Graphical Models with Faster RatesabstractIn this paper, we present a new estimator for precision matrix in high dimensional Gaussian graphical models. At the core of the proposed estimator is a collection of node-wise linear regression with nonconvex penalty. In contrast to existing estimators for Gaussian graphical models with O(s\sqrt\log d/n) estimation error bound in terms of spectral norm, where s is the maximum degree of a graph, the proposed estimator could attain O(s/\sqrtn+\sqrt\log d/n) spectral norm based convergence rate in the best case, and it is no worse than exiting estimators in general. In addition, our proposed estimator enjoys the oracle property under a milder condition than existing estimators. We show through extensive experiments on both synthetic and real datasets that our estimator outperforms the state-of-the art estimators. Lingxiao Wang 0001, Xiang Ren 0001, Quanquan Gu |
AISTATS | 2 |
| 2016 | FacetGist: Collective Extraction of Document Facets in Large Technical CorporaabstractGiven the large volume of technical documents available, it is crucial to automatically organize and categorize these documents to be able to understand and extract value from them. Towards this end, we introduce a new research problem called Facet Extraction. Given a collection of technical documents, the goal of Facet Extraction is to automatically label each document with a set of concepts for the key facets (e.g., application, technique, evaluation metrics, and dataset) that people may be interested in. Facet Extraction has numerous applications, including document summarization, literature search, patent search and business intelligence. The major challenge in performing Facet Extraction arises from multiple sources: concept extraction, concept to facet matching, and facet disambiguation. To tackle these challenges, we develop FacetGist, a framework for facet extraction. Facet Extraction involves constructing a graph-based heterogeneous network to capture information available across multiple local sentence-level features, as well as global context features. We then formulate a joint optimization problem, and propose an efficient algorithm for graph-based label propagation to estimate the facet of each concept mention. Experimental results on technical corpora from two domains demonstrate that Facet Extraction can lead to an improvement of over 25% in both precision and recall over competing schemes. Tarique Siddiqui, Xiang Ren 0001, Aditya G. Parameswaran, Jiawei Han 0001 |
CIKM | 2 |
| 2016 | AFET: Automatic Fine-Grained Entity Typing by Hierarchical Partial-Label EmbeddingabstractDistant supervision has been widely used in current systems of fine-grained entity typing to automatically assign categories (entity types) to entity mentions.However, the types so obtained from knowledge bases are often incorrect for the entity mention's local context.This paper proposes a novel embedding method to separately model "clean" and "noisy" mentions, and incorporates the given type hierarchy to induce loss functions.We formulate a joint optimization problem to learn embeddings for mentions and typepaths, and develop an iterative algorithm to solve the problem.Experiments on three public datasets demonstrate the effectiveness and robustness of the proposed method, with an average 15% improvement in accuracy over the next best compared method 1 . * Equal contribution.1 Codes and datasets used in this paper can be downloaded at https://github.com/shanzhenren/AFET. Xiang Ren 0001, Wenqi He, Meng Qu, Lifu Huang, Heng Ji 0001, Jiawei Han 0001 |
EMNLP | 1 |
| 2016 | Label Noise Reduction in Entity Typing by Heterogeneous Partial-Label EmbeddingabstractCurrent systems of fine-grained entity typing use distant supervision in conjunction with existing knowledge bases to assign categories (type labels) to entity mentions. However, the type labels so obtained from knowledge bases are often noisy (i.e., incorrect for the entity mention's local context). We define a new task, Label Noise Reduction in Entity Typing (LNR), to be the automatic identification of correct type labels (type-paths) for training examples, given the set of candidate type labels obtained by distant supervision with a given type hierarchy. The unknown type labels for individual entity mentions and the semantic similarity between entity types pose unique challenges for solving the LNR task. We propose a general framework, called PLE, to jointly embed entity mentions, text features and entity types into the same low-dimensional space where, in that space, objects whose types are semantically close have similar representations. Then we estimate the type-path for each training example in a top-down manner using the learned embeddings. We formulate a global objective for learning the embeddings from text corpora and knowledge bases, which adopts a novel margin-based loss that is robust to noisy labels and faithfully models type correlation derived from knowledge bases. Our experiments on three public typing datasets demonstrate the effectiveness and robustness of PLE, with an average of 25% improvement in accuracy compared to next best method. Xiang Ren 0001, Wenqi He, Meng Qu, Clare R. Voss, Heng Ji 0001, Jiawei Han 0001 |
KDD | 1 |
| 2016 | Automatic Entity Recognition and Typing in Massive Text DataabstractIn today's computerized and information-based society, individuals are constantly presented with vast amounts of text data, ranging from news articles, scientific publications, product reviews, to a wide range of textual information from social media. To extract value from these large, multi-domain pools of text, it is of great importance to gain an understanding of entities and their relationships. In this tutorial, we introduce data-driven methods to recognize typed entities of interest in massive, domain-specific text corpora. These methods can automatically identify token spans as entity mentions in documents and label their fine-grained types (e.g., people, product and food) in a scalable way. Since these methods do not rely on annotated data, predefined typing schema or hand-crafted features, they can be quickly adapted to a new domain, genre and language. We demonstrate on real datasets including various genres (e.g., news articles, discussion forum posts, and tweets), domains (general vs. bio-medical domains) and languages (e.g., English, Chinese, Arabic, and even low-resource languages like Hausa and Yoruba) how these typed entities aid in knowledge discovery and management. Xiang Ren 0001, Ahmed El-Kishky, Heng Ji 0001, Jiawei Han 0001 |
SIGMOD Conference | 1 |
| 2016 | Representing Documents via Latent Keyphrase InferenceabstractMany text mining approaches adopt bag-of-words or $n$-grams models to represent documents. Looking beyond just the words, fiie, the explicit surface forms, in a document can improve a computer's understanding of text. Being aware of this, researchers have proposed concept-based models that rely on a human-curated knowledge base to incorporate other related concepts in the document representation. But these methods are not desirable when applied to vertical domains (eg, literature, enterprise, etc) due to low coverage of in-domain concepts in the general knowledge base and interference from out-of-domain concepts. In this paper, we propose a data-driven model named Latent Keyphrase Inference LAKI) that represents documents with a vector of closely related domain keyphrases instead of single words or existing concepts in the knowledge base. We show that given a corpus of in-domain documents, topical content units can be learned for each domain keyphrase, which enables a computer to do smart inference to discover latent document keyphrases, going beyond just explicit mentions. Compared with the state-of-art document representation approaches, LAKI fills the gap between bag-of-words and concept-based models by using domain keyphrases as the basic representation unit. It removes dependency on a knowledge base while providing, with keyphrases, readily interpretable representations. When evaluated against 8 other methods on two text mining tasks over two corpora, LAKI outperformed all. Xiang Ren 0001, Jingbo Shang, Taylor Cassidy, Clare R. Voss, Jiawei Han 0001 |
WWW | 2 |
| 2016 | Texture Repairing by Unified Low Rank Optimization
Xiao Liang 0002, Xiang Ren 0001, Zhengdong Zhang 0001, Yi Ma 0001 |
J. Comput. Sci. Technol. | 2 |
| 2015 | Automatic Entity Recognition and Typing from Massive Text Corpora: A Phrase and Network Mining ApproachabstractIn today's computerized and information-based society, we are soaked with vast amounts of text data, ranging from news articles, scientific publications, product reviews, to a wide range of textual information from social media. To unlock the value of these unstructured text data from various domains, it is of great importance to gain an understanding of entities and their relationships. Xiang Ren 0001, Ahmed El-Kishky, Chi Wang 0001, Jiawei Han 0001 |
KDD | 1 |
| 2015 | ClusType: Effective Entity Recognition and Typing by Relation Phrase-Based ClusteringabstractEntity recognition is an important but challenging research problem. In reality, many text collections are from specific, dynamic, or emerging domains, which poses significant new challenges for entity recognition with increase in name ambiguity and context sparsity, requiring entity detection without domain restriction. In this paper, we investigate entity recognition (ER) with distant-supervision and propose a novel relation phrase-based ER framework, called ClusType, that runs data-driven phrase mining to generate entity mention candidates and relation phrases, and enforces the principle that relation phrases should be softly clustered when propagating type information between their argument entities. Then we predict the type of each entity mention based on the type signatures of its co-occurring relation phrases and the type indicators of its surface name, as computed over the corpus. Specifically, we formulate a joint optimization problem for two tasks, type propagation with relation phrases and multi-view relation phrase clustering. Our experiments on multiple genres---news, Yelp reviews and tweets---demonstrate the effectiveness and robustness of ClusType, with an average of 37% improvement in F1 score over the best compared method. Xiang Ren 0001, Ahmed El-Kishky, Chi Wang 0001, Fangbo Tao, Clare R. Voss, Jiawei Han 0001 |
KDD | 1 |
| 2015 | Mining Quality Phrases from Massive Text CorporaabstractText data are ubiquitous and play an essential role in big data applications. However, text data are mostly unstructured. Transforming unstructured text into structured units (e.g., semantically meaningful phrases) will substantially reduce semantic ambiguity and enhance the power and efficiency at manipulating such data using database technology. Thus mining quality phrases is a critical research problem in the field of databases. In this paper, we propose a new framework that extracts quality phrases from text corpora integrated with phrasal segmentation. The framework requires only limited training but the quality of phrases so generated is close to human judgment. Moreover, the method is scalable: both computation time and required space grow linearly as corpus size increases. Our experiments on large text corpora demonstrate the quality and efficiency of the new method. Jingbo Shang, Chi Wang 0001, Xiang Ren 0001, Jiawei Han 0001 |
SIGMOD Conference | 4 |
| 2014 | ClusCite: effective citation recommendation by information network-based clusteringabstractCitation recommendation is an interesting but challenging research problem. Most existing studies assume that all papers adopt the same criterion and follow the same behavioral pattern in deciding relevance and authority of a paper. However, in reality, papers have distinct citation behavioral patterns when looking for different references, depending on paper content, authors and target venues. In this study, we investigate the problem in the context of heterogeneous bibliographic networks and propose a novel cluster-based citation recommendation framework, called ClusCite, which explores the principle that citations tend to be softly clustered into interest groups based on multiple types of relationships in the network. Therefore, we predict each query's citations based on related interest groups, each having its own model for paper authority and relevance. Specifically, we learn group memberships for objects and the significance of relevance features for each interest group, while also propagating relative authority between objects, by solving a joint optimization problem. Experiments on both DBLP and PubMed datasets demonstrate the power of the proposed approach, with 17.68% improvement in Recall@50 and 9.57% growth in MRR over the best performing baseline. Xiang Ren 0001, Xiao Yu 0007, Urvashi Khandelwal, Quanquan Gu, Jiawei Han 0001 |
KDD | 1 |
| 2014 | Automatic Construction and Ranking of Topical Keyphrases on Collections of Short DocumentsabstractWe introduce a framework for topical keyphrase generation and ranking, based on the output of a topic model run on a collection of short documents. By shifting from the unigramcentric traditional methods of keyphrase extraction and ranking to a phrase-centric approach, we are able to directly compare and rank phrases of different lengths. Our method defines a function to rank topical keyphrases so that more highly ranked keyphrases are considered to be more representative phrases for that topic. We study the performance of our framework on multiple real world document collections, and also show that it is more scalable than comparable phrase-generating models. Marina Danilevsky, Chi Wang 0001, Nihit Desai, Xiang Ren 0001, Jingyi Guo, Jiawei Han 0001 |
SDM | 4 |
| 2014 | NewsNetExplorer: automatic construction and exploration of news information networksabstractNews data is one of the most abundant and familiar data sources. News data can be systematically utilized and ex- plored by database, data mining, NLP and information re- trieval researchers to demonstrate to the general public the power of advanced information technology. In our view, news data contains rich, inter-related and multi-typed data objects, forming one or a set of gigantic, interconnected, het- erogeneous information networks. Much knowledge can be derived and explored with such an information network if we systematically develop effective and scalable data-intensive information network analysis technologies. By further developing a set of information extraction, in- formation network construction, and information network mining methods, we extract types, topical hierarchies and other semantic structures from news data, construct a semi- structured news information network NewsNet. Further, we develop a set of news information network exploration and mining mechanisms that explore news in multi-dimensional space, which include (i) OLAP-based operations on the hierarchical dimensional and topical structures and rich-text, such as cell summary, single dimension analysis, and promo- tion analysis, (ii) a set of network-based operations, such as similarity search and ranking-based clustering, and (iii) a set of hybrid operations or network-OLAP operations, such as entity ranking at different granularity levels. These form the basis of our proposed NewsNetExplorer system. Although some of these functions have been studied in recent research, effective and scalable realization of such functions in large networks still poses multiple challenging research problems. Moreover, some functions are our on-going research tasks. By integrating these functions, NewsNetExplorer not only provides with us insightful recommendations in NewsNet exploration system but also helps us gain insight on how to perform effective information extraction, integration and mining in large unstructured datasets. Fangbo Tao, George Brova, Jiawei Han 0001, Heng Ji 0001, Chi Wang 0001, Brandon Norick, Ahmed El-Kishky, Xiang Ren 0001, Yizhou Sun |
SIGMOD Conference | 9 |
| 2014 | Heterogeneous graph-based intent learning with queries, web pages and Wikipedia conceptsabstractThe problem of learning user search intents has attracted intensive attention from both industry and academia. However, state-of-the-art intent learning algorithms suffer from different drawbacks when only using a single type of data source. For example, query text has difficulty in distinguishing ambiguous queries; search log is bias to the order of search results and users' noisy click behaviors. In this work, we for the first time leverage three types of objects, namely queries, web pages and Wikipedia concepts collaboratively for learning generic search intents and construct a heterogeneous graph to represent multiple types of relationships between them. A novel unsupervised method called heterogeneous graph-based soft-clustering is developed to derive an intent indicator for each object based on the constructed heterogeneous graph. With the proposed co-clustering method, one can enhance the quality of intent understanding by taking advantage of different types of data, which complement each other, and make the implicit intents easier to interpret with explicit knowledge from Wikipedia concepts. Experiments on two real-world datasets demonstrate the power of the proposed method where it achieves a 9.25% improvement in terms of NDCG on search ranking task and a 4.67% enhancement in terms of Rand index on object co-clustering task compared to the best state-of-the-art method. Xiang Ren 0001, Xiao Yu 0007, Jun Yan 0001, Zheng Chen 0001, Jiawei Han 0001 |
WSDM | 1 |
| 2014 | Personalized entity recommendation: a heterogeneous information network approachabstractAmong different hybrid recommendation techniques, network-based entity recommendation methods, which utilize user or item relationship information, are beginning to attract increasing attention recently. Most of the previous studies in this category only consider a single relationship type, such as friendships in a social network. In many scenarios, the entity recommendation problem exists in a heterogeneous information network environment. Different types of relationships can be potentially used to improve the recommendation quality. In this paper, we study the entity recommendation problem in heterogeneous information networks. Specifically, we propose to combine heterogeneous relationship information for each user differently and aim to provide high-quality personalized recommendation results using user implicit feedback data and personalized recommendation models. Xiao Yu 0007, Xiang Ren 0001, Yizhou Sun, Quanquan Gu, Bradley Sturt, Urvashi Khandelwal, Brandon Norick, Jiawei Han 0001 |
WSDM | 2 |
| 2013 | Semantic Frame-Based Document Representation for Comparable CorporaabstractDocument representation is a fundamental problem for text mining. Many efforts have been done to generate concise yet semantic representation, such as bag-of-words, phrase, sentence and topic-level descriptions. Nevertheless, most existing techniques counter difficulties in handling monolingual comparable corpus, which is a collection of monolingual documents conveying the same topic. In this paper, we propose the use of frame, a high-level semantic unit, and construct frame-based representations to semantically describe documents by bags of frames, using an information network approach. One major challenge in this representation is that semantically similar frames may be of different forms. For example, "radiation leaked" in one news article can appear as "the level of radiation increased" in another article. To tackle the problem, a text-based information network is constructed among frames and words, and a link-based similarity measure called SynRank is proposed to calculate similarity between frames. As a result, different variations of the semantically similar frames are merged into a single descriptive frame using clustering, and a document can then be represented as a bag of representative frames. It turns out that frame-based document representation not only is more interpretable, but also can facilitate other text analysis tasks such as event tracking effectively. We conduct both qualitative and quantitative experiments on three comparable news corpora, to study the effectiveness of frame-based document representation and the similarity measure SynRank, respectively, and demonstrate that the superior performance of frame-based document representation on different real-world applications. Hyungsul Kim, Xiang Ren 0001, Yizhou Sun, Chi Wang 0001, Jiawei Han 0001 |
ICDM | 2 |
| 2013 | Recommendation in heterogeneous information networks with implicit user feedbackabstractRecent studies suggest that by using additional user or item relationship information when building hybrid recommender systems, the recommendation quality can be largely improved. However, most such studies only consider a single type of relationship, e.g., social network. Notice that in many applications, the recommendation problem exists in an attribute-rich heterogeneous information network environment. In this paper, we study the entity recommendation problem in heterogeneous information networks. We propose to combine various relationship information from the network with user feedback to provide high quality recommendation results. Xiao Yu 0007, Xiang Ren 0001, Yizhou Sun, Bradley Sturt, Urvashi Khandelwal, Quanquan Gu, Brandon Norick, Jiawei Han 0001 |
RecSys | 2 |
| 2013 | Linearized Alternating Direction Method with Adaptive Penalty and Warm Starts for Fast Solving Transform Invariant Low-Rank Textures
Xiang Ren 0001, Zhouchen Lin |
Int. J. Comput. Vis. | 1 |
| 2012 | Repairing Sparse Low-Rank Texture
Xiao Liang 0002, Xiang Ren 0001, Zhengdong Zhang 0001, Yi Ma 0001 |
ECCV (5) | 2 |