VLDB 2026 Research / reviewers in the wild / expert
Muhao Chen 0001
dblp:173/2608
· DBLP profile ↗
101ranked-venue papers
13as first author
70since 2021 · last 2026
0000-0003-0118-3147ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 83 · 9 first-author · 63 since 2021Databases, data management, data science and information retrieval · 20 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Computer networks · 2 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GTA: Generating Long-horizon Tasks for Web Agents at ScaleabstractTenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou, Muhao Chen, Jonathan May, Chien-Sheng Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou, Muhao Chen 0001, Jonathan May, Chien-Sheng Wu |
ACL (1) | 5 |
| 2026 | RedCoder: Automated Multi-Turn Red Teaming for Code LLMsabstractWenjie Jacky Mo, Qin Liu, Xiaofei Wen, Dongwon Jung, Hadi Askari, Wenxuan Zhou, Zhe Zhao, Muhao Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Wenjie Mo 0001, Qin Liu 0010, Xiaofei Wen, Dongwon Jung, Hadi Askari, Wenxuan Zhou 0002, Muhao Chen 0001 |
ACL (1) | 8 |
| 2025 | SudoLM: Learning Access Control of Parametric Knowledge with Authorization AlignmentabstractExisting preference alignment is a one-size-fitsall alignment mechanism, where the part of the large language model (LLM) parametric knowledge with non-preferred features is uniformly blocked to all the users.However, this part of knowledge can be useful to advanced users whose expertise qualifies them to handle these information.The one-size-fits-all alignment mechanism undermines LLM's utility for these qualified users.To address this problem, we propose SUDOLM, a framework that lets LLMs learn access control over specific parametric knowledge for users with different credentials via authorization alignment.SUDOLM allows authorized users to unlock their access to all the parametric knowledge with an assigned SUDO key while blocking access to non-qualified users.Experiments on two application scenarios demonstrate that SUDOLM effectively controls the user's access to the parametric knowledge and maintains its general utility. Qin Liu 0010, Fei Wang 0060, Chaowei Xiao, Muhao Chen 0001 |
ACL (1) | 4 |
| 2025 | R2D2: Remembering, Replaying and Dynamic Decision Making with a Reflective Agentic MemoryabstractTenghao Huang, Kinjal Basu, Ibrahim Abdelaziz, Pavan Kapanipathi, Jonathan May, Muhao Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Tenghao Huang, Kinjal Basu 0002, Ibrahim Abdelaziz, Pavan Kapanipathi, Jonathan May, Muhao Chen 0001 |
ACL (1) | 6 |
| 2025 | AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety DetectionabstractWeidi Luo, Shenghong Dai, Xiaogeng Liu, Suman Banerjee, Huan Sun, Muhao Chen, Chaowei Xiao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Weidi Luo, Shenghong Dai, Xiaogeng Liu, Suman Banerjee 0001, Huan Sun 0001, Muhao Chen 0001, Chaowei Xiao |
ACL (1) | 6 |
| 2025 | SleeperMark: Towards Robust Watermark against Fine-Tuning Text-to-image Diffusion ModelsabstractRecent advances in large-scale text-to-image (T2I) diffusion models have enabled a variety of downstream applications. As T2I models require extensive resources for training, they constitute highly valued intellectual property (IP) for their legitimate owners, yet making them incentive targets for unauthorized fine-tuning by adversaries seeking to leverage these models for customized, usually profitable applications. Existing IP protection methods for diffusion models generally involve embedding watermark patterns and then verifying ownership through generated outputs examination, or inspecting the model’s feature space. However, these techniques are inherently ineffective in practical scenarios when the watermarked model undergoes fine-tuning, and the feature space is inaccessible during verification (i.e., black-box setting). The model is prone to forgetting the previously learned watermark knowledge when it adapts to a new task. To address this challenge, we propose SleeperMark, a novel framework designed to embed resilient watermarks into T2I diffusion models. SleeperMark explicitly guides the model to disentangle the watermark information from the semantic concepts it learns, allowing the model to retain the embedded watermark while continuing to be adapted to new downstream tasks. Our extensive experiments demonstrate the effectiveness of SleeperMark across various types of diffusion models, including latent diffusion models (e.g., Stable Diffusion) and pixel diffusion models (e.g., DeepFloyd-IF), showing robustness against downstream fine-tuning and various attacks at both the image and model levels, with minimal impact on the model’s generative capability. The code is available at https://github.com/taco-group/SleeperMark. Zilan Wang, Yiming Li 0004, Heng Huang 0001, Muhao Chen 0001, Zhengzhong Tu |
CVPR | 6 |
| 2025 | Code Execution as Grounded Supervision for LLM ReasoningabstractTraining large language models (LLMs) with chain-of-thought (CoT) supervision has proven effective for enhancing their reasoning abilities.However, obtaining reliable and accurate reasoning supervision remains a significant challenge.We propose a scalable method for generating a high-quality CoT supervision dataset by leveraging the determinism of program execution.Unlike existing reasoning dataset generation methods that rely on costly human annotations or error-prone LLM-generated CoT, our approach extracts verifiable, step-by-step reasoning traces from code execution and transforms them into a natural language CoT reasoning.Experiments on reasoning benchmarks across various domains show that our method effectively equips LLMs with transferable reasoning abilities across diverse tasks.Furthermore, the ablation studies validate that our method produces highly accurate reasoning data and reduces overall token length during inference by reducing meaningless repetition and overthinking.1 Dongwon Jung, Wenxuan Zhou 0002, Muhao Chen 0001 |
EMNLP | 3 |
| 2025 | Active Layer-Contrastive Decoding Reduces Hallucination in Large Language Model GenerationabstractRecent decoding methods improve the factuality of large language models (LLMs) by refining how the next token is selected during generation.These methods typically operate at the token level, leveraging internal representations to suppress superficial patterns.Nevertheless, LLMs remain prone to hallucinations, especially over longer contexts.In this paper, we propose Active Layer-Contrastive Decoding (ActLCD), a novel decoding strategy that actively decides when to apply contrasting layers during generation.By casting decoding as a sequential decision-making problem, ActLCD employs a reinforcement learning policy guided by a reward-aware classifier to optimize factuality beyond the token level.Our experiments demonstrate that ActLCD surpasses state-of-the-art methods across five benchmarks, showcasing its effectiveness in mitigating hallucinations in diverse generation scenarios. Hongxiang Zhang, Hao Chen 0003, Muhao Chen 0001, Tianyi Zhang 0001 |
EMNLP | 3 |
| 2025 | Benchmarking Geospatial Question Answering with MapQAabstractGeospatial question answering (QA) is a fundamental task in navigation and point of interest (POI) searches, yet existing datasets are limited in scale, diversity, and they rely on text-only descriptions without incorporating geometries. We introduce MapQA, a dataset that couples question-answer pairs with geo-entity geometries from OpenStreetMap (OSM) across two regions (Southern California and Illinois). MapQA contains 3,154 QA pairs covering nine geospatial reasoning types, including neighborhood inference and type identification, expanding both the quantity and variety of existing resources. To evaluate methods, we compare (1) a retrieval-based model that ranks geo-entities by embedding similarity and (2) large language models (LLMs) that translate questions into SQL queries executed on OSM. Retrieval-based models capture spatial relations like closeness and direction but fail on explicit distance computations, while LLMs excel at one-hop reasoning yet struggle with multi-hop tasks, revealing a key challenge for future systems. MapQA is publicly available at https://github.com/knowledge-computing/MapQA-dataset. Zekun Li 0007, Malcolm Grossman, Ehsan Qasemi, Mihir Kulkarni, Muhao Chen 0001, Yao-Yi Chiang |
SIGSPATIAL/GIS | 5 |
| 2025 | Benchmarking Vision Language Model Unlearning via Fictitious Facial Identity DatasetabstractMachine unlearning has emerged as an effective strategy for forgetting specific information in the training data. However, with the increasing integration of visual data, privacy concerns in Vision Language Models (VLMs) remain underexplored. To address this, we introduce Facial Identity Unlearning Benchmark (FIUBench), a novel VLM unlearning benchmark designed to robustly evaluate the effectiveness of unlearning algorithms under the Right to be Forgotten setting. Specifically, we formulate the VLM unlearning task via constructing the Fictitious Facial Identity VQA dataset and apply a two-stage evaluation pipeline that is designed to precisely control the sources of information and their exposure levels. In terms of evaluation, since VLM supports various forms of ways to ask questions with the same semantic meaning, we also provide robust evaluation metrics including membership inference attacks and carefully designed adversarial privacy attacks to evaluate the performance of algorithms. Through the evaluation of four baseline VLM unlearning algorithms within FIUBench, we find that all methods remain limited in their unlearning performance, with significant trade-offs between model utility and forget quality. Furthermore, our findings also highlight the importance of privacy attacks for robust evaluations. We hope FIUBench will drive progress in developing more effective VLM unlearning algorithms. Yingzi Ma, Jiongxiao Wang, Fei Wang 0060, Jiazhao Li, Jinsheng Pan, Xiujun Li, Furong Huang, Lichao Sun 0001, Bo Li 0026, Yejin Choi 0001, Muhao Chen 0001, Chaowei Xiao |
ICLR | 12 |
| 2025 | BadJudge: Backdoor Vulnerabilities of LLM-As-A-JudgeabstractThis paper proposes a novel backdoor threat attacking the LLM-as-a-Judge evaluation regime, where the adversary controls both the candidate and evaluator model. The backdoored evaluator victimizes benign users by unfairly assigning inflated scores to adversary. A trivial single token backdoor poisoning 1% of the evaluator training data triples the adversary's score with respect to their legitimate score. We systematically categorize levels of data access corresponding to three real-world settings, (1) web poisoning, (2) malicious annotator, and (3) weight poisoning. These regimes reflect a weak to strong escalation of data access that highly correlates with attack severity. Under the weakest assumptions - web poisoning (1), the adversary still induces a 20% score inflation. Likewise, in the (3) weight poisoning regime, the stronger assumptions enable the adversary to inflate their scores from 1.5/5 to 4.9/5. The backdoor threat generalizes across different evaluator architectures, trigger designs, evaluation tasks, and poisoning rates. By poisoning 10% of the evaluator training data, we control toxicity judges (Guardrails) to misclassify toxic prompts as non-toxic 89% of the time, and document reranker judges in RAG to rank the poisoned document first 97% of the time. LLM-as-a-Judge is uniquely positioned at the intersection of ethics and technology, where social implications of mislead model selection and evaluation constrain the available defensive tools. Amidst these challenges, model merging emerges as a principled tool to offset the backdoor, reducing ASR to near 0% whilst maintaining SOTA performance. Model merging's low computational cost and convenient integration into the current LLM Judge training pipeline position it as a promising avenue for backdoor mitigation in the LLM-as-a-Judge setting. Terry Tong, Fei Wang 0060, Muhao Chen 0001 |
ICLR | 4 |
| 2025 | MuirBench: A Comprehensive Benchmark for Robust Multi-image UnderstandingabstractWe introduce MuirBench, a comprehensive benchmark that focuses on robust multi-image understanding capabilities of multimodal LLMs. MuirBench consists of 12 diverse multi-image tasks (e.g., scene understanding, ordering) that involve 10 categories of multi-image relations (e.g., multiview, temporal relations). Comprising 11,264 images and 2,600 multiple-choice questions, MuirBench is created in a pairwise manner, where each standard instance is paired with an unanswerable variant that has minimal semantic differences, in order for a reliable assessment. Evaluated upon 20 recent multi-modal LLMs, our results reveal that even the best-performing models like GPT-4o and Gemini Pro find it challenging to solve MuirBench, achieving 68.0% and 49.3% in accuracy. Open-source multimodal LLMs trained on single images can hardly generalize to multi-image questions, hovering below 33.3% in accuracy. These results highlight the importance of MuirBench in encouraging the community to develop multimodal LLMs that can look beyond a single image, suggesting potential pathways for future improvements. Fei Wang 0060, James Y. Huang, Zekun Li 0007, Qin Liu 0010, Xiaogeng Liu, Mingyu Derek Ma, Nan Xu 0014, Wenxuan Zhou 0002, Kai Zhang 0008, Tianyi Lorena Yan, Wenjie Mo 0001, Hsiang-Hui Liu, Pan Lu, Chunyuan Li, Chaowei Xiao, Kai-Wei Chang 0001, Dan Roth 0001, Sheng Zhang 0012, Hoifung Poon, Muhao Chen 0001 |
ICLR | 21 |
| 2025 | MetamatBench: Integrating Heterogeneous Data, Computational Tools, and Visual Interface for Metamaterial DiscoveryabstractMetamaterials, engineered materials with architected structures across multiple length scales, offer unprecedented and tunable mechanical properties that surpass those of conventional materials. However, leveraging advanced machine learning (ML) for metamaterial discovery is hindered by three fundamental challenges: (C1) Data Heterogeneity Challenge arises from heterogeneous data sources, heterogeneous composition scales, and heterogeneous structure categories; (C2) Model Complexity Challenge stems from the intricate geometric constraints of ML models, which complicate their adaptation to metamaterial structures; and (C3) Human-AI Collaboration Challenge comes from the ''dual black-box'' nature of sophisticated ML models and the need for intuitive user interfaces. To tackle these challenges, we introduce a unified framework, named MetamatBench, that operates on three levels. (1) At the data level, we integrate and standardize 5 heterogeneous, multi-modal metamaterial datasets. (2) The ML level provides a comprehensive toolkit that adapts 17 state-of-the-art ML methods for metamaterial discovery. It also includes a comprehensive evaluation suite with 12 novel performance metrics plus a finite element-based assessment to ensure accurate and reliable model validation. (3) The user level features a visual-interactive interface that bridges the gap between complex ML techniques and non-ML researchers, advancing property prediction and inverse design of metamaterials for research and applications. MetamatBench offers a unified platform that enables machine learning researchers and practitioners to develop and evaluate new methodologies in metamaterial discovery. For accessibility and reproducibility, we open-source our benchmark and the codebase at https://github.com/cjpcool/Metamaterial-Benchmark. Jianpeng Chen, Wangzhi Zhan, Haohui Wang, Zian Jia, Jingru Gan, Jingyuan Qi, Lifu Huang, Muhao Chen 0001, Wei Wang 0010, Dawei Zhou 0003 |
KDD (2) | 10 |
| 2025 | FoodPuzzle: Toward Developing Large Language Model Agents as Autonomous Flavor ScientistsabstractFlavor development in the food industry is increasingly challenged by the need for rapid innovation and precise flavor profile creation. Traditional flavor research methods typically rely on iterative, subjective testing, which lacks the efficiency and scalability required for modern demands. This paper presents three contributions to address these challenges. Firstly, we define a new problem domain for scientific agents in flavor science, conceptualized as the generation of hypotheses for flavor profile sourcing and understanding. By leveraging their capacity to identify relevant evidence and reason within large context spaces, language model-backed agents can perform the labor-intensive tasks of flavor sourcing and understanding with enhanced efficiency and precision. To facilitate research in this area, we introduce the FoodPuzzle dataset, a challenging benchmark consisting of 978 food items and 1,766 flavor molecule profiles. We propose a novel Scientific Agent approach, integrating in-context learning and retrieval augmented techniques to generate grounded hypotheses in the domain of food science. Experimental results indicate that our model significantly surpasses traditional methods in flavor profile prediction tasks, demonstrating its potential to transform flavor development practices. Tenghao Huang, John Sweeney, Jiatong Shi, Emily Steliotes, Matthew Lange, Jonathan May, Muhao Chen 0001 |
KDD (2) | 8 |
| 2025 | From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context LearningabstractNan Xu, Fei Wang, Sheng Zhang, Hoifung Poon, Muhao Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Nan Xu 0014, Fei Wang 0060, Sheng Zhang 0012, Hoifung Poon, Muhao Chen 0001 |
NAACL (Long Papers) | 5 |
| 2025 | LayerIF: Estimating Layer Quality for Large Language Models using Influence FunctionsabstractPretrained Large Language Models (LLMs) achieve strong performance across a wide range of tasks, yet exhibit substantial variability in the various layers' training quality with respect to specific downstream applications, limiting their downstream performance. It is therefore critical to estimate layer-wise training quality in a manner that accounts for both model architecture and training data. However, existing approaches predominantly rely on model-centric heuristics (such as spectral statistics, outlier detection, or uniform allocation) while overlooking the influence of data. To address these limitations, we propose **LayerIF**, a data-driven framework that leverages *Influence Functions* to quantify the training quality of individual layers in a principled and task-sensitive manner. By isolating each layer's gradients and measuring the sensitivity of the validation loss to training examples by computing layer-wise influences, we derive data-driven estimates of layer importance. Notably, our method produces *task-specific* layer importance estimates for the *same* LLM, revealing how layers specialize for different test-time evaluation tasks. We demonstrate the utility of our scores by leveraging them for two downstream applications: (a) expert allocation in LoRA-MoE architectures and (b) layer-wise sparsity distribution for LLM pruning. Experiments across multiple LLM architectures demonstrate that our model-agnostic, influence-guided allocation leads to consistent gains in task performance. Hadi Askari, Shivanshu Gupta, Fei Wang 0060, Anshuman Chhabra, Muhao Chen 0001 |
NeurIPS | 5 |
| 2024 | RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language ModelsabstractReinforcement Learning with Human Feedback (RLHF) is a methodology designed to align Large Language Models (LLMs) with human preferences, playing an important role in LLMs alignment.Despite its advantages, RLHF relies on human annotators to rank the text, which can introduce potential security vulnerabilities if any adversarial annotator (i.e., attackers) manipulates the ranking score by upranking any malicious text to steer the LLM adversarially.To assess the red-teaming of RLHF against human preference data poisoning, we propose RankPoison, a poisoning attack method on candidates' selection of preference rank flipping to reach certain malicious behaviors (e.g., generating longer sequences, which can increase the computational cost).With poisoned dataset generated by RankPoison, we can perform poisoning attacks on LLMs to generate longer tokens without hurting the original safety alignment performance.Moreover, applying RankPoison, we also successfully implement a backdoor attack where LLMs can generate longer answers under questions with the trigger word.Our findings highlight critical security challenges in RLHF, underscoring the necessity for more robust alignment methods for LLMs. Jiongxiao Wang, Junlin Wu 0001, Muhao Chen 0001, Yevgeniy Vorobeychik, Chaowei Xiao |
ACL (1) | 3 |
| 2024 | AdaShield : Safeguarding Multimodal Large Language Models from Structure-Based Attack via Adaptive Shield Prompting
Yu Wang 0027, Xiaogeng Liu, Yu Li 0003, Muhao Chen 0001, Chaowei Xiao |
ECCV (20) | 4 |
| 2024 | Are Large Language Models Capable of Generating Human-Level Narratives?abstractYufei Tian, Tenghao Huang, Miri Liu, Derek Jiang, Alexander Spangher, Muhao Chen, Jonathan May, Nanyun Peng. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yufei Tian, Tenghao Huang, Miri Liu, Derek Jiang, Alexander Spangher, Muhao Chen 0001, Jonathan May, Nanyun Peng 0001 |
EMNLP | 6 |
| 2024 | mDPO: Conditional Preference Optimization for Multimodal Large Language ModelsabstractDirect preference optimization (DPO) has shown to be an effective method for large language model (LLM) alignment.Recent works have attempted to apply DPO to multimodal scenarios but have found it challenging to achieve consistent improvement.Through a comparative experiment, we identify the unconditional preference problem in multimodal preference optimization, where the model overlooks the image condition.To address this problem, we propose MDPO, a multimodal DPO objective that prevents the over-prioritization of language-only preferences by also optimizing image preference.Moreover, we introduce a reward anchor that forces the reward to be positive for chosen responses, thereby avoiding the decrease in their likelihood-an intrinsic problem of relative preference optimization.Experiments on two multimodal LLMs of different sizes and three widely used benchmarks demonstrate that MDPO effectively addresses the unconditional preference problem in multimodal preference optimization and significantly improves model performance, particularly in reducing hallucination. Fei Wang 0060, Wenxuan Zhou 0002, James Y. Huang, Nan Xu 0014, Sheng Zhang 0012, Hoifung Poon, Muhao Chen 0001 |
EMNLP | 7 |
| 2024 | Red Teaming Language Models for Processing Contradictory DialoguesabstractMost language models currently available are prone to self-contradiction during dialogues.To mitigate this issue, this study explores a novel contradictory dialogue processing task that aims to detect and modify contradictory statements in a conversation.This task is inspired by research on context faithfulness and dialogue comprehension, which have demonstrated that the detection and understanding of contradictions often necessitate detailed explanations.We develop a dataset comprising contradictory dialogues, in which one side of the conversation contradicts itself.Each dialogue is accompanied by an explanatory label that highlights the location and details of the contradiction.With this dataset, we present a Red Teaming framework for contradictory dialogue processing.The framework detects and attempts to explain the dialogue, then modifies the existing contradictory content using the explanation.Our experiments demonstrate that the framework improves the ability to detect contradictory dialogues and provides valid explanations.Additionally, it showcases distinct capabilities for modifying such dialogues.Our study highlights the importance of the logical inconsistency problem in conversational AI 1 Prompts Instructions Explanation Xiaofei Wen, Bangzheng Li, Tenghao Huang, Muhao Chen 0001 |
EMNLP | 4 |
| 2024 | AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsabstractThe aligned Large Language Models (LLMs) are powerful language understanding and decision-making tools that are created through extensive alignment with human feedback. However, these large models remain susceptible to jailbreak attacks, where adversaries manipulate prompts to elicit malicious outputs that should not be given by aligned LLMs. Investigating jailbreak prompts can lead us to delve into the limitations of LLMs and further guide us to secure them. Unfortunately, existing jailbreak techniques suffer from either (1) scalability issues, where attacks heavily rely on manual crafting of prompts, or (2) stealthiness problems, as attacks depend on token-based algorithms to generate prompts that are often semantically meaningless, making them susceptible to detection through basic perplexity testing. In light of these challenges, we intend to answer this question: Can we develop an approach that can automatically generate stealthy jailbreak prompts? In this paper, we introduce AutoDAN, a novel jailbreak attack against aligned LLMs. AutoDAN can automatically generate stealthy jailbreak prompts by the carefully designed hierarchical genetic algorithm. Extensive evaluations demonstrate that AutoDAN not only automates the process while preserving semantic meaningfulness, but also demonstrates superior attack strength in cross-model transferability, and cross-sample universality compared with the baseline. Moreover, we also compare AutoDAN with perplexity-based defense methods and show that AutoDAN can bypass them effectively. Code is available at https://github.com/SheltonLiu-N/AutoDAN. Xiaogeng Liu, Nan Xu 0014, Muhao Chen 0001, Chaowei Xiao |
ICLR | 3 |
| 2024 | UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity RecognitionabstractLarge language models (LLMs) have demonstrated remarkable generalizability, such as understanding arbitrary entities and relations. Instruction tuning has proven effective for distilling LLMs into more cost-efficient models such as Alpaca and Vicuna. Yet such student models still trail the original LLMs by large margins in downstream applications. In this paper, we explore targeted distillation with mission-focused instruction tuning to train student models that can excel in a broad application class such as open information extraction. Using named entity recognition (NER) for case study, we show how ChatGPT can be distilled into much smaller UniversalNER models for open NER. For evaluation, we assemble the largest NER benchmark to date, comprising 43 datasets across 9 diverse domains such as biomedicine, programming, social media, law, finance. Without using any direct supervision, UniversalNER attains remarkable NER accuracy across tens of thousands of entity types, outperforming general instruction-tuned models such as Alpaca and Vicuna by over 30 absolute F1 points in average. With a tiny fraction of parameters, UniversalNER not only acquires ChatGPT's capability in recognizing arbitrary entity types, but also outperforms its NER accuracy by 7-9 absolute F1 points in average. Remarkably, UniversalNER even outperforms by a large margin state-of-the-art multi-task instruction-tuned systems such as InstructUIE, which uses supervised NER examples. We also conduct thorough ablation studies to assess the impact of various components in our distillation approach. We release the distillation recipe, data, and UniversalNER models to facilitate future research on targeted distillation. Wenxuan Zhou 0002, Sheng Zhang 0012, Yu Gu 0017, Muhao Chen 0001, Hoifung Poon |
ICLR | 4 |
| 2024 | Rethinking Tabular Data Understanding with Large Language ModelsabstractTianyang Liu, Fei Wang, Muhao Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Tianyang Liu 0003, Fei Wang 0060, Muhao Chen 0001 |
NAACL-HLT | 3 |
| 2024 | Two Heads are Better than One: Nested PoE for Robust Defense Against Multi-BackdoorsabstractVictoria Graf, Qin Liu, Muhao Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Victoria Graf, Qin Liu 0010, Muhao Chen 0001 |
NAACL-HLT | 3 |
| 2024 | Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination?abstractBangzheng Li, Ben Zhou, Fei Wang, Xingyu Fu, Dan Roth, Muhao Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Bangzheng Li, Ben Zhou, Fei Wang 0060, Dan Roth 0001, Muhao Chen 0001 |
NAACL-HLT | 6 |
| 2024 | From Shortcuts to Triggers: Backdoor Defense with Denoised PoEabstractQin Liu, Fei Wang, Chaowei Xiao, Muhao Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Qin Liu 0010, Fei Wang 0060, Chaowei Xiao, Muhao Chen 0001 |
NAACL-HLT | 4 |
| 2024 | How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their VulnerabilitiesabstractLingbo Mo, Boshi Wang, Muhao Chen, Huan Sun. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Lingbo Mo, Boshi Wang, Muhao Chen 0001, Huan Sun 0001 |
NAACL-HLT | 3 |
| 2024 | Instructional Fingerprinting of Large Language ModelsabstractJiashu Xu, Fei Wang, Mingyu Ma, Pang Wei Koh, Chaowei Xiao, Muhao Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Fei Wang 0060, Mingyu Derek Ma, Pang Wei Koh, Chaowei Xiao, Muhao Chen 0001 |
NAACL-HLT | 6 |
| 2024 | Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language ModelsabstractJiashu Xu, Mingyu Ma, Fei Wang, Chaowei Xiao, Muhao Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Mingyu Derek Ma, Fei Wang 0060, Chaowei Xiao, Muhao Chen 0001 |
NAACL-HLT | 5 |
| 2024 | BackdoorAlign: Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety AlignmentabstractDespite the general capabilities of Large Language Models (LLMs) like GPT-4, these models still request fine-tuning or adaptation with customized data when meeting the specific business demands and intricacies of tailored use cases. However, this process inevitably introduces new safety threats, particularly against the Fine-tuning based Jailbreak Attack (FJAttack) under the setting of Language-Model-as-a-Service (LMaaS), where the model's safety has been significantly compromised by fine-tuning on users' uploaded examples that contain just a few harmful examples. Though potential defenses have been proposed that the service providers of LMaaS can integrate safety examples into the fine-tuning dataset to reduce safety issues, such approaches require incorporating a substantial amount of data, making it inefficient. To effectively defend against the FJAttack with limited safety examples under LMaaS, we propose the Backdoor Enhanced Safety Alignment method inspired by an analogy with the concept of backdoor attacks. In particular, service providers will construct prefixed safety examples with a secret prompt, acting as a "backdoor trigger". By integrating prefixed safety examples into the fine-tuning dataset, the subsequent fine-tuning process effectively acts as the "backdoor attack", establishing a strong correlation between the secret prompt and safety generations. Consequently, safe responses are ensured once service providers prepend this secret prompt ahead of any user input during inference. Our comprehensive experiments demonstrate that through the Backdoor Enhanced Safety Alignment with adding as few as 11 prefixed safety examples, the maliciously fine-tuned LLMs will achieve similar safety performance as the original aligned models without harming the benign performance. Furthermore, we also present the effectiveness of our method in a more practical setting where the fine-tuning data consists of both FJAttack examples and the fine-tuning task data. Jiongxiao Wang, Jiazhao Li, Yiquan Li, Xiangyu Qi, Junjie Hu 0001, Yixuan Li 0001, Patrick McDaniel, Muhao Chen 0001, Bo Li 0026, Chaowei Xiao |
NeurIPS | 8 |
| 2023 | Can NLI Provide Proper Indirect Supervision for Low-resource Biomedical Relation Extraction?abstractTwo key obstacles in biomedical relation extraction (RE) are the scarcity of annotations and the prevalence of instances without explicitly pre-defined labels due to low annotation coverage.Existing approaches, which treat biomedical RE as a multi-class classification task, often result in poor generalization in low-resource settings and do not have the ability to make selective predictions on unknown cases but give a guess from seen relations, hindering the applicability of those approaches.We present NBR, which converts biomedical RE as a natural language inference formulation to provide indirect supervision.By converting relations to natural language hypotheses, NBR is capable of exploiting semantic cues to alleviate annotation scarcity.By incorporating a ranking-based loss that implicitly calibrates abstinent instances, NBR learns a clearer decision boundary and is instructed to abstain on uncertain instances.Extensive experiments on three widely-used biomedical RE benchmarks, namely ChemProt, DDI, and GAD, verify the effectiveness of NBR in both full-shot and low-resource regimes.Our analysis demonstrates that indirect supervision benefits biomedical RE even when a domain gap exists, and combining NLI knowledge with biomedical knowledge leads to the best performance gains. 1 Mingyu Derek Ma, Muhao Chen 0001 |
ACL (1) | 3 |
| 2023 | Continual Contrastive Finetuning Improves Low-Resource Relation ExtractionabstractRelation extraction (RE), which has relied on structurally annotated corpora for model training, has been particularly challenging in lowresource scenarios and domains.Recent literature has tackled low-resource RE by selfsupervised learning, where the solution involves pretraining the entity pair embedding by RE-based objective and finetuning on labeled data by classification-based objective.However, a critical challenge to this approach is the gap in objectives, which prevents the RE model from fully utilizing the knowledge in pretrained representations.In this paper, we aim at bridging the gap and propose to pretrain and finetune the RE model using consistent objectives of contrastive learning.Since in this kind of representation learning paradigm, one relation may easily form multiple clusters in the representation space, we further propose a multi-center contrastive loss that allows one relation to form multiple clusters to better align with pretraining.Experiments on two document-level RE datasets, BioRED and Re-DocRED, demonstrate the effectiveness of our method.Particularly, when using 1% end-task training data, our method outperforms PLMbased RE classifier by 10.5% and 6.1% on the two datasets, respectively. Wenxuan Zhou 0002, Sheng Zhang 0012, Tristan Naumann, Muhao Chen 0001, Hoifung Poon |
ACL (1) | 4 |
| 2023 | Detecting Semantic Errors in Tables using Textual EvidenceabstractTables can contain various types of errors, including both syntactic and semantic errors. Semantic errors relate to the meaning of the data and can be detrimental for downstream applications. The existing approaches for semantic error detection use structured knowledge sources such as Wikidata and DBpedia, but the coverage of such sources is quite limited. There is much more information available in free text to validate the contents of tables. In this paper, we present a novel semantic-error-detection approach that exploits open-domain textual data to verify the semantic correctness of tables. Our approach leverages contrastive learning, table linearization, and pre-trained language models to implement the error detection process. We implement our approach in a system called SEED and show in the evaluation that it significantly outperforms the other competing approaches. Minh Pham 0004, Craig A. Knoblock, Muhao Chen 0001 |
IEEE Big Data | 3 |
| 2023 | How Fragile is Relation Extraction under Entity Replacements?abstractYiwei Wang, Bryan Hooi, Fei Wang, Yujun Cai, Yuxuan Liang, Wenxuan Zhou, Jing Tang, Manjuan Duan, Muhao Chen. Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL). 2023. Yiwei Wang 0001, Bryan Hooi, Fei Wang 0060, Yujun Cai, Yuxuan Liang 0002, Wenxuan Zhou 0002, Jing Tang 0004, Manjuan Duan, Muhao Chen 0001 |
CoNLL | 9 |
| 2023 | Extracting or Guessing? Improving Faithfulness of Event Temporal Relation ExtractionabstractIn this paper, we seek to improve the faithfulness of TEMPREL extraction models from two perspectives.The first perspective is to extract genuinely based on contextual description.To achieve this, we propose to conduct counterfactual analysis to attenuate the effects of two significant types of training biases: the event trigger bias and the frequent label bias.We also add tense information into event representations to explicitly place an emphasis on the contextual description.The second perspective is to provide proper uncertainty estimation and abstain from extraction when no relation is described in the text.By parameterization of Dirichlet Prior over the model-predicted categorical distribution, we improve the model estimates of the correctness likelihood and make TEMPREL predictions more selective.We also employ temperature scaling to recalibrate the model confidence measure after bias mitigation.Through experimental analysis on MATRES, MATRES-DS, and TDDiscourse, we demonstrate that our model extracts TEMPREL and timelines more faithfully compared to SOTA methods, especially under distribution shifts. Haoyu Wang 0005, Hongming Zhang 0009, Yuqian Deng, Jacob R. Gardner, Dan Roth 0001, Muhao Chen 0001 |
EACL | 6 |
| 2023 | Parameter-Efficient Tuning with Special Token AdaptationabstractParameter-efficient tuning aims at updating only a small subset of parameters when adapting a pretrained model to downstream tasks.In this work, we introduce PASTA, in which we only modify the special token representations (e.g., [SEP] and [CLS] in BERT) before the self-attention module at each layer in Transformer-based models.PASTA achieves comparable performance to full finetuning in natural language understanding tasks including text classification and NER with up to only 0.029% of total parameters trained.Our work not only provides a simple yet effective way of parameter-efficient tuning, which has a wide range of practical applications when deploying finetuned models for multiple tasks, but also demonstrates the pivotal role of special tokens in pretrained language models. 1 Xiaocong Yang, James Y. Huang, Wenxuan Zhou 0002, Muhao Chen 0001 |
EACL | 4 |
| 2023 | Are All Steps Equally Important? Benchmarking Essentiality Detection in Event ProcessesabstractNatural language expresses events with varying granularities, where coarse-grained events (goals) can be broken down into finer-grained event sequences (steps).A critical yet overlooked aspect of understanding event processes is recognizing that not all step events hold equal importance toward the completion of a goal.In this paper, we address this gap by examining the extent to which current models comprehend the essentiality of step events in relation to a goal event.Cognitive studies suggest that such capability enables machines to emulate human commonsense reasoning about preconditions and necessary efforts of everyday tasks.We contribute a high-quality corpus of (goal, step) pairs gathered from the community guideline website WikiHow, with steps manually annotated for their essentiality concerning the goal by experts.The high inter-annotator agreement demonstrates that humans possess a consistent understanding of event essentiality.However, after evaluating multiple statistical and largescale pre-trained language models, we find that existing approaches considerably underperform compared to humans.This observation highlights the need for further exploration into this critical and challenging task 1 . Haoyu Wang 0005, Hongming Zhang 0009, Yueguan Wang, Yuqian Deng, Muhao Chen 0001, Dan Roth 0001 |
EMNLP | 5 |
| 2023 | GeoLM: Empowering Language Models for Geospatially Grounded Language UnderstandingabstractHumans subconsciously engage in geospatial reasoning when reading articles.We recognize place names and their spatial relations in text and mentally associate them with their physical locations on Earth.Although pretrained language models can mimic this cognitive process using linguistic context, they do not utilize valuable geospatial information in large, widely available geographical databases, e.g., OpenStreetMap.This paper introduces GEOLM ( ), a geospatially grounded language model that enhances the understanding of geo-entities in natural language.GEOLM leverages geo-entity mentions as anchors to connect linguistic information in text corpora with geospatial information extracted from geographical databases.GEOLM connects the two types of context through contrastive learning and masked language modeling.It also incorporates a spatial coordinate embedding mechanism to encode distance and direction relations to capture geospatial context.In the experiment, we demonstrate that GEOLM exhibits promising capabilities in supporting toponym recognition, toponym linking, relation extraction, and geo-entity typing, which bridge the gap between natural language processing and geospatial sciences.The code is publicly available at https://github.com/ knowledge-computing/geolm. Zekun Li 0007, Wenxuan Zhou 0002, Yao-Yi Chiang, Muhao Chen 0001 |
EMNLP | 4 |
| 2023 | Bridging Continuous and Discrete Spaces: Interpretable Sentence Representation Learning via Compositional OperationsabstractTraditional sentence embedding models encode sentences into vector representations to capture useful properties such as the semantic similarity between sentences.However, in addition to similarity, sentence semantics can also be interpreted via compositional operations such as sentence fusion or difference.It is unclear whether the compositional semantics of sentences can be directly reflected as compositional operations in the embedding space.To more effectively bridge the continuous embedding and discrete text spaces, we explore the plausibility of incorporating various compositional properties into the sentence embedding space that allows us to interpret embedding transformations as compositional sentence operations.We propose INTERSENT, an end-toend framework for learning interpretable sentence embeddings that supports compositional sentence operations in the embedding space.Our method optimizes operator networks and a bottleneck encoder-decoder model to produce meaningful and interpretable sentence embeddings.Experimental results demonstrate that our method significantly improves the interpretability of sentence embeddings on four textual generation tasks over existing approaches while maintaining strong performance on traditional semantic similarity tasks. 1 . James Y. Huang, Wenlin Yao, Kaiqiang Song, Hongming Zhang 0009, Muhao Chen 0001, Dong Yu 0001 |
EMNLP | 5 |
| 2023 | Primacy Effect of ChatGPTabstractInstruction-tuned large language models (LLMs), such as ChatGPT, have led to promising zero-shot performance in discriminative natural language understanding (NLU) tasks.This involves querying the LLM using a prompt containing the question, and the candidate labels to choose from.The question-answering capabilities of ChatGPT arise from its pre-training on large amounts of human-written text, as well as its subsequent fine-tuning on human preferences, which motivates us to ask: Does ChatGPT also inherit humans' cognitive biases?In this paper, we study the primacy effect of ChatGPT: the tendency of selecting the labels at earlier positions as the answer.We have two main findings: i) ChatGPT's decision is sensitive to the order of labels in the prompt; ii) ChatGPT has a clearly higher chance to select the labels at earlier positions as the answer.We hope that our experiments and analyses provide additional insights into building more reliable ChatGPT-based solutions.We release the source code at https: //github.com/wangywUST/PrimacyEffectGPT. Yiwei Wang 0001, Yujun Cai, Muhao Chen 0001, Yuxuan Liang 0002, Bryan Hooi |
EMNLP | 3 |
| 2023 | PINTO: Faithful Language Reasoning Using Prompt-Generated Rationales
Peifeng Wang, Aaron Chan, Filip Ilievski, Muhao Chen 0001, Xiang Ren 0001 |
ICLR | 4 |
| 2023 | Automated Summarization of Stack Overflow PostsabstractSoftware developers often resort to Stack Overflow (SO) to fill their programming needs. Given the abundance of relevant posts, navigating them and comparing different solutions is tedious and time-consuming. Recent work has proposed to automatically summarize SO posts to concise text to facilitate the navigation of SO posts. However, these techniques rely only on information retrieval methods or heuristics for text summarization, which is insufficient to handle the ambiguity and sophistication of natural language. This paper presents a deep learning based framework called Assortfor SO post summarization. Assortincludes two complementary learning methods,$\mathbf{Assort}_{S}$and$\mathbf{Assort}_{IS}$, to address the lack of labeled training data for SO post summarization.$\mathbf{Assort}_{S}$is designed to directly train a novel ensemble learning model with BERT embeddings and domain-specific features to account for the unique characteristics of SO posts. By contrast,$\mathbf{Assort}_{IS}$is designed to reuse pre-trained models while addressing the domain shift challenge when no training data is present (i.e., zero-shot learning). Both$\mathbf{Assort}_{S}$and$\mathbf{Assort}_{IS}$outperform six existing techniques by at least 13% and 7% respectively in terms of the F1 score. Furthermore, a human study shows that participants significantly preferred summaries generated by$\mathbf{Assort}_{S}$and$\mathbf{Assort}_{IS}$over the best baseline, while the preference difference between$\mathbf{Assort}_{S}$and$\mathbf{Assort}_{IS}$was small. Bonan Kou, Muhao Chen 0001, Tianyi Zhang 0001 |
ICSE | 2 |
| 2023 | Software Entity Recognition with Noise-Robust LearningabstractRecognizing software entities such as library names from free-form text is essential to enable many software engineering (SE) technologies, such as traceability link recovery, automated documentation, and API recommendation. While many approaches have been proposed to address this problem, they suffer from small entity vocabularies or noisy training data, hindering their ability to recognize software entities mentioned in sophisticated narratives. To address this challenge, we leverage the Wikipedia taxonomy to develop a comprehensive entity lexicon with 79K unique software entities in 12 fine-grained types, as well as a large labeled dataset of over 1.7M sentences. Then, we propose self-regularization, a noise-robust learning approach, to the training of our software entity recognition (SER) model by accounting for many dropouts. Results show that models trained with self-regularization outperform both their vanilla counterparts and state-of-the-art approaches on our Wikipedia benchmark and two Stack Overflow benchmarks. We release our models11https://huggingface.co/taidng/wikiser-bert-base; https.//huggingface.co/taidng/wikiser-bert-large., data, and code for future research.22https://github.com/taidnguyen/software_entity_recognition Tai Nguyen 0005, Yifeng Di, Joohan Lee, Muhao Chen 0001, Tianyi Zhang 0001 |
ASE | 4 |
| 2023 | Knowledge-Based Version Incompatibility Detection for Deep LearningabstractVersion incompatibility issues are rampant when reusing or reproducing deep learning models and applications. Existing techniques are limited to library dependency specifications declared in PyPI. Therefore, these techniques cannot detect version issues due to undocumented version constraints or issues involving hardware drivers or OS. To address this challenge, we propose to leverage the abundant discussions of DL version issues from Stack Overflow to facilitate version incompatibility detection. We reformulate the problem of knowledge extraction as a Question-Answering (QA) problem and use a pre-trained QA model to extract version compatibility knowledge from online discussions. The extracted knowledge is further consolidated into a weighted knowledge graph to detect potential version incompatibilities when reusing a DL project. Our evaluation results show that (1) our approach can accurately extract version knowledge with 84% accuracy, and (2) our approach can accurately identify 65% of known version issues in 10 popular DL projects with a high precision (92%), while two state-of-the-art approaches can only detect 29% and 6% of these issues with 33% and 17% precision respectively. Bonan Kou, Mohamed Yilmaz Ibrahim, Muhao Chen 0001, Tianyi Zhang 0001 |
ESEC/SIGSOFT FSE | 4 |
| 2022 | Prix-LM: Pretraining for Multilingual Knowledge Base ConstructionabstractKnowledge bases (KBs) contain plenty of structured world and commonsense knowledge.As such, they often complement distributional text-based information and facilitate various downstream tasks.Since their manual construction is resource-and timeintensive, recent efforts have tried leveraging large pretrained language models (PLMs) to generate additional monolingual knowledge facts for KBs.However, such methods have not been attempted for building and enriching multilingual KBs.Besides wider application, such multilingual KBs can provide richer combined knowledge than monolingual (e.g., English) KBs.Knowledge expressed in different languages may be complementary and unequally distributed: this implies that the knowledge available in high-resource languages can be transferred to low-resource ones.To achieve this, it is crucial to represent multilingual knowledge in a shared/unified space.To this end, we propose a unified representation model, Prix-LM , for multilingual KB construction and completion.We leverage two types of knowledge, monolingual triples and cross-lingual links, extracted from existing multilingual KBs, and tune a multilingual language encoder XLM-R via a causal language modeling objective.Prix-LM integrates useful multilingual and KB-based factual knowledge into a single model.Experiments on standard entity-related tasks, such as link prediction in multiple languages, cross-lingual entity linking and bilingual lexicon induction, demonstrate its effectiveness, with gains reported over strong task-specialised baselines. Wenxuan Zhou 0002, Fangyu Liu 0001, Ivan Vulic, Nigel Collier, Muhao Chen 0001 |
ACL (1) | 5 |
| 2022 | Salience Allocation as Guidance for Abstractive SummarizationabstractFei Wang, Kaiqiang Song, Hongming Zhang, Lifeng Jin, Sangwoo Cho, Wenlin Yao, Xiaoyang Wang, Muhao Chen, Dong Yu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Fei Wang 0060, Kaiqiang Song, Hongming Zhang 0009, Lifeng Jin, Sangwoo Cho, Wenlin Yao, Xiaoyang Wang 0001, Muhao Chen 0001, Dong Yu 0001 |
EMNLP | 8 |
| 2022 | Does Your Model Classify Entities Reasonably? Diagnosing and Mitigating Spurious Correlations in Entity TypingabstractEntity typing aims at predicting one or more words that describe the type(s) of a specific mention in a sentence.Due to shortcuts from surface patterns to annotated entity labels and biased training, existing entity typing models are subject to the problem of spurious correlations.To comprehensively investigate the faithfulness and reliability of entity typing methods, we first systematically define distinct kinds of model biases that are reflected mainly from spurious correlations.Particularly, we identify six types of existing model biases, including mention-context bias, lexical overlapping bias, named entity bias, pronoun bias, dependency bias, and overgeneralization bias.To mitigate model biases, we then introduce a counterfactual data augmentation method.By augmenting the original training set with their debiased counterparts, models are forced to fully comprehend sentences and discover the fundamental cues for entity typing, rather than relying on spurious correlations for shortcuts.Experimental results on the UFET dataset show our counterfactual data augmentation approach helps improve generalization of different entity typing models with consistently better performance on both the original and debiased test sets 1 .Input: Last week I stayed in Treasure Island for two nights when visiting Las Vegas.Gold labels: hotel, resort, location, place Pred labels: island, land, location, place Input: Next day (-> Next twenty-four hour period), after the Slovaks captured Rymanow Zdroj and the the Germans seized Krosno, the Brigade was ordered to withdraw to Sanok and leave Dukla.Gold labels: day, time, event, date, year Pred labels: day, time, date -> time, hour period Input: Kevin Donovan (-> Brennan), after seeing Michael Caine movie about the Zulu uprising, decided to form the Universal Zulu Nation, an organization based on merits derived from art and achievements Nan Xu 0014, Fei Wang 0060, Bangzheng Li, Mingtao Dong, Muhao Chen 0001 |
EMNLP | 5 |
| 2022 | Contextualized Scene Imagination for Generative Commonsense Reasoning
Peifeng Wang, Jonathan Zamora, Filip Ilievski, Muhao Chen 0001, Xiang Ren 0001 |
ICLR | 5 |
| 2022 | SOSum: A Dataset of Stack Overflow Post SummariesabstractStack Overflow (SO) is becoming an indispensable part of modern software development workflow. However, given the limited time, attention, and memory capacity of programmers, navigating SO posts and comparing different solutions is time-consuming and cumbersome. Recent research has proposed to summarize SO posts to concise text to help programmers quickly assess the relevance and quality of SO posts. Yet there is no large dataset of high-quality SO post summaries, hindering the development and evaluation of post summarization techniques. We present SOSum, a dataset of 2,278 popular SO answer posts with manually labeled summative sentences. Questions in SOSum cover 669 tags with a median view count of 253K and a median post score of 17. This dataset will foster research on sentence-level summarization of SO posts and has the potential to facilitate text summarization research on other types of textual software artifacts such as programming tutorials. Bonan Kou, Yifeng Di, Muhao Chen 0001, Tianyi Zhang 0001 |
MSR | 3 |
| 2022 | Unified Semantic Typing with Meaningful Label InferenceabstractJames Y. Huang, Bangzheng Li, Jiashu Xu, Muhao Chen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. James Y. Huang, Bangzheng Li, Muhao Chen 0001 |
NAACL-HLT | 4 |
| 2022 | Should We Rely on Entity Mentions for Relation Extraction? Debiasing Relation Extraction with Counterfactual AnalysisabstractYiwei Wang, Muhao Chen, Wenxuan Zhou, Yujun Cai, Yuxuan Liang, Dayiheng Liu, Baosong Yang, Juncheng Liu, Bryan Hooi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yiwei Wang 0001, Muhao Chen 0001, Wenxuan Zhou 0002, Yujun Cai, Yuxuan Liang 0002, Dayiheng Liu, Baosong Yang, Bryan Hooi |
NAACL-HLT | 2 |
| 2022 | Robust (Controlled) Table-to-Text Generation with Structure-Aware Equivariance LearningabstractControlled table-to-text generation seeks to generate natural language descriptions for highlighted subparts of a table.Previous SOTA systems still employ a sequence-to-sequence generation method, which merely captures the table as a linear structure and is brittle when table layouts change.We seek to go beyond this paradigm by (1) effectively expressing the relations of content pieces in the table, and (2) making our model robust to content-invariant structural transformations.Accordingly, we propose an equivariance learning framework, LATTICE ( ), which encodes tables with a structure-aware self-attention mechanism.This prunes the full self-attention structure into an order-invariant graph attention that captures the connected graph structure of cells belonging to the same row or column, and it differentiates between relevant cells and irrelevant cells from the structural perspective.Our framework also modifies the positional encoding mechanism to preserve the relative position of tokens in the same cell but enforce position invariance among different cells.Our technology is free to be plugged into existing table-to-text generation models, and has improved T5-based models to offer better performance on ToTTo and HiTab.Moreover, on a harder version of ToTTo, we preserve promising performance, while previous SOTA systems, even with transformationbased data augmentation, have seen significant performance drops. 1 Fei Wang 0060, Zhewei Xu, Pedro A. Szekely, Muhao Chen 0001 |
NAACL-HLT | 4 |
| 2022 | Answer Consolidation: Formulation and BenchmarkingabstractWenxuan Zhou, Qiang Ning, Heba Elfardy, Kevin Small, Muhao Chen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Wenxuan Zhou 0002, Qiang Ning, Heba Elfardy, Kevin Small, Muhao Chen 0001 |
NAACL-HLT | 5 |
| 2022 | AdaDiag: Adversarial Domain Adaptation of Diagnostic Prediction with Clinical Event SequencesabstractEarly detection of heart failure (HF) can provide patients with the opportunity for more timely intervention and better disease management, as well as efficient use of healthcare resources. Recent machine learning (ML) methods have shown promising performance on diagnostic prediction using temporal sequences from electronic health records (EHRs). In practice, however, these models may not generalize to other populations due to dataset shift. Shifts in datasets can be attributed to a range of factors such as variations in demographics, data management methods, and healthcare delivery patterns. In this paper, we use unsupervised adversarial domain adaptation methods to adaptively reduce the impact of dataset shift on cross-institutional transfer performance. The proposed framework is validated on a next-visit HF onset prediction task using a BERT-style Transformer-based language model pre-trained with a masked language modeling (MLM) task. Our model empirically demonstrates superior prediction performance relative to non-adversarial baselines in both transfer directions on two different clinical event sequence data sources. Muhao Chen 0001, Alex Bui |
J. Biomed. Informatics | 2 |
| 2022 | Ultra-fine Entity Typing with Indirect Supervision from Natural Language InferenceabstractAbstract The task of ultra-fine entity typing (UFET) seeks to predict diverse and free-form words or phrases that describe the appropriate types of entities mentioned in sentences. A key challenge for this task lies in the large number of types and the scarcity of annotated data per type. Existing systems formulate the task as a multi-way classification problem and train directly or distantly supervised classifiers. This causes two issues: (i) the classifiers do not capture the type semantics because types are often converted into indices; (ii) systems developed in this way are limited to predicting within a pre-defined type set, and often fall short of generalizing to types that are rarely seen or unseen in training. This work presents LITE🍻, a new approach that formulates entity typing as a natural language inference (NLI) problem, making use of (i) the indirect supervision from NLI to infer type information meaningfully represented as textual hypotheses and alleviate the data scarcity issue, as well as (ii) a learning-to-rank objective to avoid the pre-defining of a type set. Experiments show that, with limited training data, LITE obtains state-of-the-art performance on the UFET task. In addition, LITE demonstrates its strong generalizability by not only yielding best results on other fine-grained entity typing benchmarks, more importantly, a pre-trained LITE system works well on new data containing unseen types.1 Bangzheng Li, Wenpeng Yin 0001, Muhao Chen 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2021 | Visual Pivoting for (Unsupervised) Entity AlignmentabstractThis work studies the use of visual semantic representations to align entities in heterogeneous knowledge graphs (KGs). Images are natural components of many existing KGs. By combining visual knowledge with other auxiliary information, we show that the proposed new approach, EVA, creates a holistic entity representation that provides strong signals for cross-graph entity alignment. Besides, previous entity alignment methods require human labelled seed alignment, restricting availability. EVA provides a completely unsupervised solution by leveraging the visual similarity of entities to create an initial seed dictionary (visual pivots). Experiments on benchmark data sets DBP15k and DWY15k show that EVA offers state-of-the-art performance on both monolingual and cross-lingual entity alignment tasks. Furthermore, we discover that images are particularly useful to align long-tail KG entities, which inherently lack the structural contexts necessary for capturing the correspondences. Code release: https://github.com/cambridgeltl/eva; project page: http://cogcomp.org/page/publication view/927. Fangyu Liu 0001, Muhao Chen 0001, Dan Roth 0001, Nigel Collier |
AAAI | 2 |
| 2021 | Learning from History: Modeling Temporal Knowledge Graphs with Sequential Copy-Generation NetworksabstractLarge knowledge graphs often grow to store temporal facts that model the dynamic relations or interactions of entities along the timeline. Since such temporal knowledge graphs often suffer from incompleteness, it is important to develop time-aware representation learning models that help to infer the missing temporal facts. While the temporal facts are typically evolving, it is observed that many facts often show a repeated pattern along the timeline, such as economic crises and diplomatic activities. This observation indicates that a model could potentially learn much from the known facts appeared in history. To this end, we propose a new representation learning model for temporal knowledge graphs, namely CyGNet, based on a novel time-aware copy-generation mechanism. CyGNet is not only able to predict future facts from the whole entity vocabulary, but also capable of identifying facts with repetition and accordingly predicting such future facts with reference to the known facts in the past. We evaluate the proposed method on the knowledge graph completion task using five benchmark datasets. Extensive experiments demonstrate the effectiveness of CyGNet for predicting future facts with repetition as well as de novo fact prediction. Cunchao Zhu, Muhao Chen 0001, Changjun Fan, Guangquan Cheng, Yan Zhang 0080 |
AAAI | 2 |
| 2021 | Knowing the No-match: Entity Alignment with Dangling CasesabstractZequn Sun, Muhao Chen, Wei Hu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zequn Sun 0001, Muhao Chen 0001, Wei Hu 0007 |
ACL/IJCNLP (1) | 2 |
| 2021 | Tabular Functional Block Detection with Embedding-based Agglomerative Cell ClusteringabstractTables are a widely-used format for data curation. The diversity of domains, layouts, and content of tables makes knowledge extraction challenging. Understanding table layouts is an important step for automatically harvesting knowledge from tabular data. Since table cells are spatially organized into regions, correctly identifying such regions and inferring their functional roles, referred to as functional block detection, is a critical part of understanding table layouts. Earlier functional block detection approaches fail to leverage spatial relationships and higher-level structure, either depending on cell-level predictions or relying on data types as signals for identifying blocks. In this paper, we introduce a flexible functional block detection method by applying agglomerative clustering techniques which merge smaller blocks into larger blocks using two merging strategies. Our proposed method uses cell embeddings with a customized dissimilarity function which utilizes local and margin distances, as well as block coherence metrics to capture cell, block, and table scoped features. Given the diversity of tables in real-world corpora, we also introduce a sampling-based approach for automatically tuning distance thresholds for each table. Experimental results show that our method improves over the earlier state-of-the-art method in terms of several evaluation metrics. Kexuan Sun 0002, Fei Wang 0060, Muhao Chen 0001, Jay Pujara |
CIKM | 3 |
| 2021 | Cross-lingual Entity Alignment with Incidental SupervisionabstractMuch research effort has been put to multilingual knowledge graph (KG) embedding methods to address the entity alignment task, which seeks to match entities in different languagespecific KGs that refer to the same real-world object.Such methods are often hindered by the insufficiency of seed alignment provided between KGs.Therefore, we propose an incidentally supervised model, JEANS , which jointly represents multilingual KGs and text corpora in a shared embedding scheme, and seeks to improve entity alignment with incidental supervision signals from text.JEANS first deploys an entity grounding process to combine each KG with the monolingual text corpus.Then, two learning processes are conducted: (i) an embedding learning process to encode the KG and text of each language in one embedding space, and (ii) a selflearning based alignment learning process to iteratively induce the matching of entities and that of lexemes between embeddings.Experiments on benchmark datasets show that JEANS leads to promising improvement on entity alignment with incidental supervision, and significantly outperforms state-of-the-art methods that solely rely on internal information of KGs. 1 * Indicating equal contributions. Muhao Chen 0001, Ben Zhou, Dan Roth 0001 |
EACL | 1 |
| 2021 | Learning Constraints and Descriptive Segmentation for Subevent DetectionabstractEvent mentions in text correspond to realworld events of varying degrees of granularity.The task of subevent detection aims to resolve this granularity issue, recognizing the membership of multi-granular events in event complexes.Since knowing the span of descriptive contexts of event complexes helps infer the membership of events, we propose the task of event-based text segmentation (EVENTSEG) as an auxiliary task to improve the learning for subevent detection.To bridge the two tasks together, we propose an approach to learning and enforcing constraints that capture dependencies between subevent detection and EVENTSEG prediction, as well as guiding the model to make globally consistent inference.Specifically, we adopt Rectifier Networks for constraint learning and then convert the learned constraints to a regularization term in the loss function of the neural model.Experimental results show that the proposed method outperforms baseline methods by 2.3% and 2.5% on benchmark datasets for subevent detection, HiEve and IC, respectively, while achieving a decent performance on EVENTSEG prediction 1 . Haoyu Wang 0005, Hongming Zhang 0009, Muhao Chen 0001, Dan Roth 0001 |
EMNLP (1) | 3 |
| 2021 | Salience-Aware Event Chain Modeling for Narrative UnderstandingabstractStorytelling, whether via fables, news reports, documentaries, or memoirs, can be thought of as the communication of interesting and related events that, taken together, form a concrete process.It is desirable to extract the event chains that represent such processes.However, this extraction remains a challenging problem.We posit that this is due to the nature of the texts from which chains are discovered.Natural language text interleaves a narrative of concrete, salient events with background information, contextualization, opinion, and other elements that are important for a variety of necessary discourse and pragmatics acts but are not part of the principal chain of events being communicated.We introduce methods for extracting this principal chain from natural language text, by filtering away non-salient events and supportive sentences.We demonstrate the effectiveness of our methods at isolating critical event chains by comparing their effect on downstream tasks.We show that by pre-training large language models on our extracted chains, we obtain improvements in two tasks that benefit from a clear understanding of event chains: narrative prediction and event-based temporal question answering.The demonstrated improvements and ablative studies confirm that our extraction method isolates critical event chains. 1 Xiyang Zhang 0004, Muhao Chen 0001, Jonathan May |
EMNLP (1) | 2 |
| 2021 | Contrastive Out-of-Distribution Detection for Pretrained TransformersabstractPretrained Transformers achieve remarkable performance when training and test data are from the same distribution.However, in realworld scenarios, the model often faces out-ofdistribution (OOD) instances that can cause severe semantic shift problems at inference time.Therefore, in practice, a reliable model should identify such instances, and then either reject them during inference or pass them over to models that handle another distribution.In this paper, we develop an unsupervised OOD detection method, in which only the indistribution (ID) data are used in training.We propose to fine-tune the Transformers with a contrastive loss, which improves the compactness of representations, such that OOD instances can be better differentiated from ID ones.These OOD instances can then be accurately detected using the Mahalanobis distance in the model's penultimate layer.We experiment with comprehensive settings and achieve near-perfect OOD detection performance, outperforming baselines drastically.We further investigate the rationales behind the improvement, finding that more compact representations through margin-based contrastive learning bring the improvement.We release our code to the community for future research 1 . Wenxuan Zhou 0002, Fangyu Liu 0001, Muhao Chen 0001 |
EMNLP (1) | 3 |
| 2021 | Learning from Noisy Labels for Entity-Centric Information ExtractionabstractRecent information extraction approaches have relied on training deep neural models.However, such models can easily overfit noisy labels and suffer from performance degradation.While it is very costly to filter noisy labels in large learning resources, recent studies show that such labels take more training steps to be memorized and are more frequently forgotten than clean labels, therefore are identifiable in training.Motivated by such properties, we propose a simple co-regularization framework for entity-centric information extraction, which consists of several neural models with identical structures but different parameter initialization.These models are jointly optimized with the task-specific losses and are regularized to generate similar predictions based on an agreement loss, which prevents overfitting on noisy labels.Extensive experiments on two widely used but noisy benchmarks for information extraction, TACRED and CoNLL03, demonstrate the effectiveness of our framework.We release our code to the community for future research 1 . Wenxuan Zhou 0002, Muhao Chen 0001 |
EMNLP (1) | 2 |
| 2021 | SPADE: A Semi-supervised Probabilistic Approach for Detecting Errors in TablesabstractError detection is one of the most important steps in data cleaning and usually requires extensive human interaction to ensure quality. Existing supervised methods in error detection require a significant amount of training data while unsupervised methods rely on fixed inductive biases, which are usually hard to generalize, to solve the problem. In this paper, we present SPADE, a novel semi-supervised probabilistic approach for error detection. SPADE introduces a novel probabilistic active learning model, where the system suggests examples to be labeled based on the agreements between user labels and indicative signals, which are designed to capture potential errors. SPADE uses a two-phase data augmentation process to enrich a dataset before training a deep learning classifier to detect unlabeled errors. In our evaluation, SPADE achieves an average F1-score of 0.91 over five datasets and yields a 10% improvement compared with the state-of-the-art systems. Minh Pham 0004, Craig A. Knoblock, Muhao Chen 0001, Jay Pujara |
IJCAI | 3 |
| 2021 | From Tables to Knowledge: Recent Advances in Table UnderstandingabstractA wealth of human knowledge is expressed in structured tables, across web pages, scientific articles, spreadsheets, and databases. This wealth of knowledge is mirrored by diversity in the vast number of layout structures, content types, formats, and surface forms used to express tables. Recent advances in representation learning and knowledge representation have made progress in exploiting structural regularities in tabular data to unlock this knowledge. In this tutorial, we provide a survey of these advances for a host of table understanding tasks, including table segmentation, semantic typing of cells, transforming tables to knowledge graphs, entity linking, and table retrieval tasks for question answering. Jay Pujara, Pedro A. Szekely, Huan Sun 0001, Muhao Chen 0001 |
KDD | 4 |
| 2021 | Probabilistic Box Embeddings for Uncertain Knowledge Graph ReasoningabstractXuelu Chen, Michael Boratko, Muhao Chen, Shib Sankar Dasgupta, Xiang Lorraine Li, Andrew McCallum. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Xuelu Chen, Michael Boratko, Muhao Chen 0001, Shib Sankar Dasgupta, Xiang Li 0069, Andrew McCallum |
NAACL-HLT | 3 |
| 2021 | Retrieving Complex Tables with Multi-Granular Graph Representation LearningabstractThe task of natural language table retrieval (NLTR) seeks to retrieve semantically relevant tables based on natural language queries. Existing learning systems for this task often treat tables as plain text based on the assumption that tables are structured as dataframes. However, tables can have complex layouts which indicate diverse dependencies between subtable structures, such as nested headers. As a result, queries may refer to different spans of relevant content that is distributed across these structures. Moreover, such systems fail to generalize to novel scenarios beyond those seen in the training set. Prior methods are still distant from a generalizable solution to the NLTR problem, as they fall short in handling complex table layouts or queries over multiple granularities. To address these issues, we propose Graph-based Table Retrieval (GTR), a generalizable NLTR framework with multi-granular graph representation learning. In our framework, a table is first converted into a tabular graph, with cell nodes, row nodes and column nodes to capture content at different granularities. Then the tabular graph is input to a Graph Transformer model that can capture both table cell content and the layout structures. To enhance the robustness and generalizability of the model, we further incorporate a self-supervised pre-training task based on graph-context matching. Experimental results on two benchmarks show that our method leads to significant improvements over the current state-of-the-art systems. Further experiments demonstrate promising performance of our method on cross-dataset generalization, and enhanced capability of handling complex tables and fulfilling diverse query intents. Fei Wang 0060, Kexuan Sun 0002, Muhao Chen 0001, Jay Pujara, Pedro A. Szekely |
SIGIR | 3 |
| 2021 | JEDI: circular RNA prediction based on junction encoders and deep interaction among splice sitesabstractMOTIVATION: Circular RNA (circRNA) is a novel class of long non-coding RNAs that have been broadly discovered in the eukaryotic transcriptome. The circular structure arises from a non-canonical splicing process, where the donor site backspliced to an upstream acceptor site. These circRNA sequences are conserved across species. More importantly, rising evidence suggests their vital roles in gene regulation and association with diseases. As the fundamental effort toward elucidating their functions and mechanisms, several computational methods have been proposed to predict the circular structure from the primary sequence. Recently, advanced computational methods leverage deep learning to capture the relevant patterns from RNA sequences and model their interactions to facilitate the prediction. However, these methods fail to fully explore positional information of splice junctions and their deep interaction. RESULTS: We present a robust end-to-end framework, Junction Encoder with Deep Interaction (JEDI), for circRNA prediction using only nucleotide sequences. JEDI first leverages the attention mechanism to encode each junction site based on deep bidirectional recurrent neural networks and then presents the novel cross-attention layer to model deep interaction among these sites for backsplicing. Finally, JEDI can not only predict circRNAs but also interpret relationships among splice sites to discover backsplicing hotspots within a gene region. Experiments demonstrate JEDI significantly outperforms state-of-the-art approaches in circRNA prediction on both isoform level and gene level. Moreover, JEDI also shows promising results on zero-shot backsplicing discovery, where none of the existing approaches can achieve. AVAILABILITY AND IMPLEMENTATION: The implementation of our framework is available at https://github.com/hallogameboy/JEDI. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jyun-Yu Jiang, Chelsea J.-T. Ju, Junheng Hao, Muhao Chen 0001, Wei Wang 0010 |
Bioinform. | 4 |
| 2020 | Knowledge Graph Alignment Network with Gated Multi-Hop Neighborhood AggregationabstractGraph neural networks (GNNs) have emerged as a powerful paradigm for embedding-based entity alignment due to their capability of identifying isomorphic subgraphs. However, in real knowledge graphs (KGs), the counterpart entities usually have non-isomorphic neighborhood structures, which easily causes GNNs to yield different representations for them. To tackle this problem, we propose a new KG alignment network, namely AliNet, aiming at mitigating the non-isomorphism of neighborhood structures in an end-to-end manner. As the direct neighbors of counterpart entities are usually dissimilar due to the schema heterogeneity, AliNet introduces distant neighbors to expand the overlap between their neighborhood structures. It employs an attention mechanism to highlight helpful distant neighbors and reduce noises. Then, it controls the aggregation of both direct and distant neighborhood information using a gating mechanism. We further propose a relation loss to refine entity representations. We perform thorough experiments with detailed ablation studies and analyses on five entity alignment datasets, demonstrating the effectiveness of AliNet. Zequn Sun 0001, Wei Hu 0007, Muhao Chen 0001, Yuzhong Qu |
AAAI | 4 |
| 2020 | Diagnostic Prediction with Sequence-of-sets Representation Learning for Clinical Events
Muhao Chen 0001, Alex Bui |
AIME | 2 |
| 2020 | What Are You Trying to Do? Semantic Typing of Event ProcessesabstractThis paper studies a new cognitively motivated semantic typing task, multi-axis event process typing, that, given an event process, attempts to infer free-form type labels describing (i) the type of action made by the process and (ii) the type of object the process seeks to affect.This task is inspired by computational and cognitive studies of event understanding, which suggest that understanding processes of events is often directed by recognizing the goals, plans or intentions of the protagonist(s).We develop a large dataset containing over 60k event processes, featuring ultra fine-grained typing on both the action and object type axes with very large (10 3 ∼ 10 4 ) label vocabularies.We then propose a hybrid learning framework, P2GT, which addresses the challenging typing problem with indirect supervision from glosses 1 and a joint learning-to-rank framework.As our experiments indicate, P2GT supports identifying the intent of processes, as well as the fine semantic type of the affected object.It also demonstrates the capability of handling fewshot cases, and strong generalizability on outof-domain processes.2 * This work was done when the author was visiting the University of Pennsylvania.1 A gloss provides a sense definition for a lexeme. 2 The contributed learning resources, software and a system demonstration are available at http://cogcomp.org/page/publication_view/915. Muhao Chen 0001, Hongming Zhang 0009, Haoyu Wang 0005, Dan Roth 0001 |
CoNLL | 1 |
| 2020 | ReadNet: A Hierarchical Transformer Framework for Web Article Readability Analysis
Changping Meng, Muhao Chen 0001, Jie Mao, Jennifer Neville |
ECIR (1) | 2 |
| 2020 | Knowledge Association with Hyperbolic Knowledge Graph EmbeddingsabstractCapturing associations for knowledge graphs (KGs) through entity alignment, entity type inference and other related tasks benefits NLP applications with comprehensive knowledge representations.Recent related methods built on Euclidean embeddings are challenged by the hierarchical structures and different scales of KGs.They also depend on high embedding dimensions to realize enough expressiveness.Differently, we explore with low-dimensional hyperbolic embeddings for knowledge association.We propose a hyperbolic relational graph neural network for KG embedding and capture knowledge associations with a hyperbolic transformation.Extensive experiments on entity alignment and type inference demonstrate the effectiveness and efficiency of our method. Zequn Sun 0001, Muhao Chen 0001, Wei Hu 0007 |
EMNLP (1) | 2 |
| 2020 | Joint Constrained Learning for Event-Event Relation ExtractionabstractUnderstanding natural language involves recognizing how multiple event mentions structurally and temporally interact with each other.In this process, one can induce event complexes that organize multi-granular events with temporal order and membership relations interweaving among them.Due to the lack of jointly labeled data for these relational phenomena and the restriction on the structures they articulate, we propose a joint constrained learning framework for modeling event-event relations.Specifically, the framework enforces logical constraints within and across multiple temporal and subevent relations by converting these constraints into differentiable learning objectives.We show that our joint constrained learning approach effectively compensates for the lack of jointly labeled data, and outperforms SOTA methods on benchmarks for both temporal relation extraction and event hierarchy construction, replacing a commonly used but more expensive global inference process.We also present a promising case study showing the effectiveness of our approach in inducing event complexes on an external corpus. 1 Haoyu Wang 0005, Muhao Chen 0001, Hongming Zhang 0009, Dan Roth 0001 |
EMNLP (1) | 2 |
| 2020 | Analogous Process Structure Induction for Sub-event Sequence PredictionabstractComputational and cognitive studies of event understanding suggest that identifying, comprehending, and predicting events depend on having structured representations of a sequence of events and on conceptualizing (abstracting) its components into (soft) event categories.Thus, knowledge about a known process such as "buying a car" can be used in the context of a new but analogous process such as "buying a house".Nevertheless, most event understanding work in NLP is still at the ground level and does not consider abstraction.In this paper, we propose an Analogous Process Structure Induction (APSI) framework, which leverages analogies among processes and conceptualization of sub-event instances to predict the whole sub-event sequence of previously unseen open-domain processes.As our experiments and analysis indicate, APSI 1 supports the generation of meaningful sub-event sequences for unseen processes and can help predict missing events. Hongming Zhang 0009, Muhao Chen 0001, Haoyu Wang 0005, Yangqiu Song, Dan Roth 0001 |
EMNLP (1) | 2 |
| 2020 | A Benchmarking Study of Embedding-based Entity Alignment for Knowledge Graphs
Zequn Sun 0001, Qingheng Zhang, Wei Hu 0007, Muhao Chen 0001, Farahnaz Akrami, Chengkai Li 0001 |
Proc. VLDB Endow. | 5 |
| 2019 | Embedding Uncertain Knowledge GraphsabstractEmbedding models for deterministic Knowledge Graphs (KG) have been extensively studied, with the purpose of capturing latent semantic relations between entities and incorporating the structured knowledge they contain into machine learning. However, there are many KGs that model uncertain knowledge, which typically model the inherent uncertainty of relations facts with a confidence score, and embedding such uncertain knowledge represents an unresolved challenge. The capturing of uncertain knowledge will benefit many knowledge-driven applications such as question answering and semantic search by providing more natural characterization of the knowledge. In this paper, we propose a novel uncertain KG embedding model UKGE, which aims to preserve both structural and uncertainty information of relation facts in the embedding space. Unlike previous models that characterize relation facts with binary classification techniques, UKGE learns embeddings according to the confidence scores of uncertain relation facts. To further enhance the precision of UKGE, we also introduce probabilistic soft logic to infer confidence scores for unseen relation facts during training. We propose and evaluate two variants of UKGE based on different confidence score modeling strategies. Experiments are conducted on three real-world uncertain KGs via three tasks, i.e. confidence prediction, relation fact ranking, and relation fact classification. UKGE shows effectiveness in capturing uncertain knowledge by achieving promising results, and it consistently outperforms baselines on these tasks. Xuelu Chen, Muhao Chen 0001, Yizhou Sun, Carlo Zaniolo |
AAAI | 2 |
| 2019 | Learning to Differentiate Between Main-articles and Sub-articles in WikipediaabstractCurrent Wikipedia editing approaches typically summarize a named entity by one main-article supplemented by multiple sub-articles describing various aspects and subtopics of the entity. Such separation of articles aims at improving the curation of content-rich Wikipedia entities. However, a wide range of Wikipedia-based technologies critically rely on the article-as-concept assumption, which requires a one-to-one mapping between entities (or concepts) and the articles that describe these entities. Thus, the current editing approaches sow confusion and ambiguity to knowledge representation, and cause problems to a wide-range of downstream technologies. In this paper, we present an approach that resolves these problems by differentiating the main-article from the sub-articles that are not at the core of entity representations. We propose a hybrid neural article model that learns on two facets of a Wikipedia article: (i) Two neural document encoders capture the latent semantic features from the article title and text contents. (ii) A set of explicit features measure and characterize the symbolic and structural aspects of each article. In this study, we use crowdsourcing to create a large annotated dataset for feature extraction, and for evaluating a variety of encoding techniques and learning structures. The optimized model so derived identifies main articles with near-perfect precision and recall, and outperforms various baselines on the contributed dataset. Muhao Chen 0001, Changping Meng, Carlo Zaniolo |
IEEE BigData | 1 |
| 2019 | Fast and Accurate Network Embeddings via Very Sparse Random ProjectionabstractWe present FastRP, a scalable and performant algorithm for learning distributed node representations in a graph. FastRP is over 4,000 times faster than state-of-the-art methods such as DeepWalk and node2vec, while achieving comparable or even better performance as evaluated on several real-world networks on various downstream tasks. We observe that most network embedding methods consist of two components: construct a node similarity matrix and then apply dimension reduction techniques to this matrix. We show that the success of these methods should be attributed to the proper construction of this similarity matrix, rather than the dimension reduction method employed. FastRP is proposed as a scalable algorithm for network embeddings. Two key features of FastRP are: 1) it explicitly constructs a node similarity matrix that captures transitive relationships in a graph and normalizes matrix entries based on node degrees; 2) it utilizes very sparse random projection, which is a scalable optimization-free method for dimension reduction. An extra benefit from combining these two design choices is that it allows the iterative computation of node embeddings so that the similarity matrix need not be explicitly constructed, which further speeds up FastRP. FastRP is also advantageous for its ease of implementation, parallelization and hyperparameter tuning. The source code is available at https://github.com/GTmac/FastRP. Haochen Chen, Syed Fahad Sultan, Yingtao Tian, Muhao Chen 0001, Steven Skiena |
CIKM | 4 |
| 2019 | Learning to Identify High Betweenness Centrality Nodes from Scratch: A Novel Graph Neural Network ApproachabstractBetweenness centrality (BC) is a widely used centrality measures for network analysis, which seeks to describe the importance of nodes in a network in terms of the fraction of shortest paths that pass through them. It is key to many valuable applications, including community detection and network dismantling. Computing BC scores on large networks is computationally challenging due to its high time complexity. Many sampling-based approximation algorithms have been proposed to speed up the estimation of BC. However, these methods still need considerable long running time on large-scale networks, and their results are sensitive to even small perturbation to the networks. In this paper, we focus on the efficient identification of top-k nodes with highest BC in a graph, which is an essential task to many network applications. Different from previous heuristic methods, we turn this task into a learning problem and design an encoder-decoder based framework as a solution. Specifically, the encoder leverages the network structure to represent each node as an embedding vector, which captures the important structural information of the node. The decoder transforms each embedding vector into a scalar, which identifies the relative rank of a node in terms of its BC. We use the pairwise ranking loss to train the model to identify the orders of nodes regarding their BC. By training on small-scale networks, the model is capable of assigning relative BC scores to nodes for much larger networks, and thus identifying the highly-ranked nodes. Experiments on both synthetic and real-world networks demonstrate that, compared to existing baselines, our model drastically speeds up the prediction without noticeable sacrifice in accuracy, and even outperforms the state-of-the-arts in terms of accuracy on several large real-world networks. Changjun Fan, Yuhui Ding, Muhao Chen 0001, Yizhou Sun, Zhong Liu 0002 |
CIKM | 4 |
| 2019 | Learning to Represent Bilingual DictionariesabstractBilingual word embeddings have been widely used to capture the correspondence of lexical semantics in different human languages.However, the cross-lingual correspondence between sentences and words is less studied, despite that this correspondence can significantly benefit many applications such as crosslingual semantic search and textual inference.To bridge this gap, we propose a neural embedding model that leverages bilingual dictionaries 1 .The proposed model is trained to map the lexical definitions to the cross-lingual target words, for which we explore with different sentence encoding techniques.To enhance the learning process on limited resources, our model adopts several critical learning strategies, including multi-task learning on different bridges of languages, and joint learning of the dictionary model with a bilingual word embedding model.We conduct experiments on two new tasks.In the cross-lingual reverse dictionary retrieval task, we demonstrate that our model is capable of comprehending bilingual concepts based on descriptions, and the proposed learning strategies are effective.In the bilingual paraphrase identification task, we show that our model effectively associates sentences in different languages via a shared embedding space, and outperforms existing approaches in identifying bilingual paraphrases. Muhao Chen 0001, Yingtao Tian, Haochen Chen, Kai-Wei Chang 0001, Steven Skiena, Carlo Zaniolo |
CoNLL | 1 |
| 2019 | Social Relation Inference via Label Propagation
Yingtao Tian, Haochen Chen, Bryan Perozzi, Muhao Chen 0001, Steven Skiena |
ECIR (1) | 4 |
| 2019 | Retrofitting Contextualized Word Embeddings with ParaphrasesabstractWeijia Shi, Muhao Chen, Pei Zhou, Kai-Wei Chang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Muhao Chen 0001, Kai-Wei Chang 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Examining Gender Bias in Languages with Grammatical GenderabstractPei Zhou, Weijia Shi, Jieyu Zhao, Kuan-Hao Huang, Muhao Chen, Ryan Cotterell, Kai-Wei Chang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jieyu Zhao 0001, Kuan-Hao Huang, Muhao Chen 0001, Ryan Cotterell, Kai-Wei Chang 0001 |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Multi-view Knowledge Graph Embedding for Entity AlignmentabstractWe study the problem of embedding-based entity alignment between knowledge graphs (KGs). Previous works mainly focus on the relational structure of entities. Some further incorporate another type of features, such as attributes, for refinement. However, a vast of entity features are still unexplored or not equally treated together, which impairs the accuracy and robustness of embedding-based entity alignment. In this paper, we propose a novel framework that unifies multiple views of entities to learn embeddings for entity alignment. Specifically, we embed entities based on the views of entity names, relations and attributes, with several combination strategies. Furthermore, we design some cross-KG inference methods to enhance the alignment between two KGs. Our experiments on real-world datasets show that the proposed framework significantly outperforms the state-of-the-art embedding-based entity alignment methods. The selected views, cross-KG inference and combination strategies all contribute to the performance improvement. Qingheng Zhang, Zequn Sun 0001, Wei Hu 0007, Muhao Chen 0001, Lingbing Guo, Yuzhong Qu |
IJCAI | 4 |
| 2019 | Universal Representation Learning of Knowledge Bases by Jointly Embedding Instances and Ontological ConceptsabstractMany large-scale knowledge bases simultaneously represent two views of knowledge graphs (KGs): an ontology view for abstract and commonsense concepts, and an instance view for specific entities that are instantiated from ontological concepts. Existing KG embedding models, however, merely focus on representing one of the two views alone. In this paper, we propose a novel two-view KG embedding model, JOIE, with the goal to produce better knowledge embedding and enable new applications that rely on multi-view knowledge. JOIE employs both cross-view and intra-view modeling that learn on multiple facets of the knowledge base. The cross-view association model is learned to bridge the embeddings of ontological concepts and their corresponding instance-view entities. The intra-view models are trained to capture the structured knowledge of instance and ontology views in separate embedding spaces, with a hierarchy-aware encoding technique enabled for ontologies with hierarchies. We explore multiple representation techniques for the two model components and investigate with nine variants of JOIE. Our model is trained on large-scale knowledge bases that consist of massive instances and their corresponding ontological concepts connected via a (small) set of cross-view links. Experimental results on public datasets show that the best variant of JOIE significantly outperforms previous models on instance-view triple prediction task as well as ontology population on ontology-view KG. In addition, our model successfully extends the use of KG embeddings to entity typing with promising performance. Junheng Hao, Muhao Chen 0001, Wenchao Yu, Yizhou Sun, Wei Wang 0010 |
KDD | 2 |
| 2019 | TransEdge: Translating Relation-Contextualized Embeddings for Knowledge Graphs
Zequn Sun 0001, Jiacheng Huang 0001, Wei Hu 0007, Muhao Chen 0001, Lingbing Guo, Yuzhong Qu |
ISWC (1) | 4 |
| 2019 | Embedding Edge-attributed Relational HierarchiesabstractRelational embedding methods encode objects and their relations as low-dimensional vectors. While achieving competitive performance on a variety of relational inference tasks, these methods fall short of preserving the hierarchies that are often formed in existing graph data, and ignore the rich edge attributes that describe the relation facts. In this paper, we propose a novel embedding method that simultaneously preserve the hierarchical property and the edge information in the edge-attributed relational hierarchies. The proposed method preserves the hierarchical relations by leveraging the non-linearity of hyperbolic vector translations, for which the edge attributes are exploited to capture the importance of each relation fact. Our experiment is conducted on the well-known Enron organizational chart, where the supervision relations between employees of the Enron company are accompanied with email-based attributes. We show that our method produces relational embeddings of higher quality than state-of-the-art methods, and outperforms a variety of strong baselines in reconstructing the organizational chart. Muhao Chen 0001, Chris Quirk |
SIGIR | 1 |
| 2019 | Multifaceted protein-protein interaction prediction based on Siamese residual RCNNabstractMOTIVATION: Sequence-based protein-protein interaction (PPI) prediction represents a fundamental computational biology problem. To address this problem, extensive research efforts have been made to extract predefined features from the sequences. Based on these features, statistical algorithms are learned to classify the PPIs. However, such explicit features are usually costly to extract, and typically have limited coverage on the PPI information. RESULTS: We present an end-to-end framework, PIPR (Protein-Protein Interaction Prediction Based on Siamese Residual RCNN), for PPI predictions using only the protein sequences. PIPR incorporates a deep residual recurrent convolutional neural network in the Siamese architecture, which leverages both robust local features and contextualized information, which are significant for capturing the mutual influence of proteins sequences. PIPR relieves the data pre-processing efforts that are required by other systems, and generalizes well to different application scenarios. Experimental evaluations show that PIPR outperforms various state-of-the-art systems on the binary PPI prediction problem. Moreover, it shows a promising performance on more challenging problems of interaction type prediction and binding affinity estimation, where existing approaches fall short. AVAILABILITY AND IMPLEMENTATION: The implementation is available at https://github.com/muhaochen/seq_ppi.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Muhao Chen 0001, Chelsea J.-T. Ju, Xuelu Chen, Kai-Wei Chang 0001, Carlo Zaniolo, Wei Wang 0010 |
Bioinform. | 1 |
| 2018 | Enhanced Network Embeddings via Exploiting Edge LabelsabstractNetwork embedding methods aim at learning low-dimensional latent representation of nodes in a network. While achieving competitive performance on a variety of network inference tasks such as node classification and link prediction, these methods treat the relations between nodes as a binary variable and ignore the rich semantics of edges. In this work, we attempt to learn network embeddings which simultaneously preserve network structure and relations between nodes. Experiments on several real-world networks illustrate that by considering different relations between different node pairs, our method is capable of producing node embeddings of higher quality than a number of state-of-the-art network embedding methods, as evaluated on a challenging multi-label node classification task. Haochen Chen, Yingtao Tian, Bryan Perozzi, Muhao Chen 0001, Steven Skiena |
CIKM | 5 |
| 2018 | Co-training Embeddings of Knowledge Graphs and Entity Descriptions for Cross-lingual Entity AlignmentabstractMultilingual knowledge graph (KG) embeddings provide latent semantic representations of entities and structured knowledge with cross-lingual inferences, which benefit various knowledge-driven cross-lingual NLP tasks. However, precisely learning such cross-lingual inferences is usually hindered by the low coverage of entity alignment in many KGs. Since many multilingual KGs also provide literal descriptions of entities, in this paper, we introduce an embedding-based approach which leverages a weakly aligned multilingual KG for semi-supervised cross-lingual learning using entity descriptions. Our approach performs co-training of two embedding models, i.e. a multilingual KG embedding model and a multilingual literal description embedding model. The models are trained on a large Wikipedia-based trilingual dataset where most entity alignment is unknown to training. Experimental results show that the performance of the proposed approach on the entity alignment task improves at each iteration of co-training, and eventually reaches a stage at which it significantly surpasses previous approaches. We also show that our approach has promising abilities for zero-shot entity alignment, and cross-lingual KG completion. Muhao Chen 0001, Yingtao Tian, Kai-Wei Chang 0001, Steven Skiena, Carlo Zaniolo |
IJCAI | 1 |
| 2018 | Demand-driven Cache Allocation Based on Context-aware Collaborative FilteringabstractMany recent advances of network caching focus on i) more effectively modeling the preferences of a regional user group to different web contents, and ii) reducing the cost of content delivery by storing the most popular contents in regional caches. However, the context under which the users interact with the network system usually causes tremendous variations in a user group's preferences on the contents. To effectively leverage such contextual information for more efficient network caching, we propose a novel mechanism to incorporate context-aware collaborative filtering into demand-driven caching. By differentiating the characterization of user interests based on a priori contexts, our approach seeks to enhance the cache performance with a more dynamic and fine-grained cache allocation process. In particular, our approach is general and adapts to various types of context information. Our evaluation shows that this new approach significantly outperforms previous non-demand-driven caching strategies by offering much higher cached content rate, especially when utilizing the contextual information. Muhao Chen 0001, Qi Zhao 0002, Pengyuan Du, Carlo Zaniolo, Mario Gerla |
MobiHoc | 1 |
| 2018 | Towards Opportunistic Resource Sharing in Mobile Social Networks: an Evolutionary Game Theoretic ApproachabstractIn mobile social networks, the success of resource sharing depends on a high level of cooperations. The motivation of this work is to seek conditions under which cooperation prevails without additional incentive mechanisms such as credit and reputation-based schemes. We apply the Evolutionary Game Theory framework to investigate the formation of cooperation in opportunistic resource sharing. First, we extend the existing Small World In Motion mobility model to preserve real-world localized mobility patterns. On top of the mobility model, a game theoretic model tailored for resource sharing is developed. Preliminary simulation results show that high user cooperation rate emerges when the cost of resource sharing is sufficiently small, even if the Nash Equilibrium of the resource sharing game is non-cooperation. Moreover, we discovered that heterogeneous user mobility patterns promote cooperation. Pengyuan Du, Seunghyun Yoo, Qi Zhao 0002, Muhao Chen 0001, Mario Gerla |
MobiHoc | 4 |
| 2018 | Neural Article Pair Modeling for Wikipedia Sub-article Matching
Muhao Chen 0001, Changping Meng, Carlo Zaniolo |
ECML/PKDD (3) | 1 |
| 2018 | On2Vec: Embedding-based Relation Prediction for Ontology PopulationabstractPopulating ontology graphs represents a long-standing problem for the Semantic Web community. Recent advances in translation-based graph embedding methods for populating instance-level knowledge graphs lead to promising new approaching for the ontology population problem. However, unlike instance-level graphs, the majority of relation facts in ontology graphs come with comprehensive semantic relations, which often include the properties of transitivity and symmetry, as well as hierarchical relations. These comprehensive relations are often too complex for existing graph embedding methods, and direct application of such methods is not feasible. Hence, we propose On2Vec, a novel translation-based graph embedding method for ontology population. On2Vec integrates two model components that effectively characterize comprehensive relation facts in ontology graphs. The first is the Component-specific Model that encodes concepts and relations into low-dimensional embedding spaces without a loss of relational properties; the second is the Hierarchy Model that performs focused learning of hierarchical relation facts. Experiments on several well-known ontology graphs demonstrate the promising capabilities of On2Vec in predicting and verifying new relation facts. These promising results also make possible significant improvements in related methods. Muhao Chen 0001, Yingtao Tian, Xuelu Chen, Zijun Xue, Carlo Zaniolo |
SDM | 1 |
| 2018 | User-friendly temporal queries on historical knowledge bases
Carlo Zaniolo, Shi Gao, Maurizio Atzori, Muhao Chen 0001, Jiaqi Gu 0001 |
Inf. Comput. | 4 |
| 2017 | Multilingual Knowledge Graph Embeddings for Cross-lingual Knowledge AlignmentabstractMany recent works have demonstrated the benefits of knowledge graph embeddings in completing monolingual knowledge graphs. Inasmuch as related knowledge bases are built in several different languages, achieving cross-lingual knowledge alignment will help people in constructing a coherent knowledge base, and assist machines in dealing with different expressions of entity relationships across diverse human languages. Unfortunately, achieving this highly desirable cross-lingual alignment by human labor is very costly and error-prone. Thus, we propose MTransE, a translation-based model for multilingual knowledge graph embeddings, to provide a simple and automated solution. By encoding entities and relations of each language in a separated embedding space, MTransE provides transitions for each embedding vector to its cross-lingual counterparts in other spaces, while preserving the functionalities of monolingual embeddings. We deploy three different techniques to represent cross-lingual transitions, namely axis calibration, translation vectors, and linear transformations, and derive five variants for MTransE using different loss functions. Our models can be trained on partially aligned graphs, where just a small portion of triples are aligned with their cross-lingual counterparts. The experiments on cross-lingual entity matching and triple-wise alignment verification show promising results, with some variants consistently outperforming others on different tasks. We also explore how MTransE preserves the key properties of its monolingual counterpart. Muhao Chen 0001, Yingtao Tian, Mohan Yang, Carlo Zaniolo |
IJCAI | 1 |
| 2017 | Learning Multi-faceted Knowledge Graph Embeddings for Natural Language ProcessingabstractKnowledge graphs have challenged the present embedding-based approaches for representing their multifacetedness. To address some of the issues, we have investigated some novel approaches that (i) captures multilingual transitions on different language-specific versions of knowledge, and (ii) encodes the commonly existing monolingual knowledge with important relational properties and hierarchies. In addition, we propose the use of our approaches in a wide spectrum of NLP tasks that have not been well explored by related works. Muhao Chen 0001, Carlo Zaniolo |
IJCAI | 1 |
| 2016 | Converting spatiotemporal data Among heterogeneous granularity systemsabstractSpatiotemporal data are often expressed in terms of granularities to indicate the measurement units of the data. A granularity system usually consists of a set of granularities that share a “common refined granularity” (CRG) to enable granular comparison and data conversion within the system. However, if data from multiple granularity systems needs to be used in a unified application, it is necessary to extend the data conversion and comparison within a granularity system to those for multiple granularity systems. This paper proposes a formal framework to enable such an extension. The framework involves essentially some preconditions and properties for verifying the existence of a CRG and unifying conversions of incongruous semantics, and supports the approach to integrate multiple systems into one so as to process granular interoperation across systems just like in a single system. Quantification of uncertainty in granularity conversion is also considered to improve the precision of granular comparison. Muhao Chen 0001, Shi Gao, Xiaoyang Sean Wang |
FUZZ-IEEE | 1 |