Kai-Wei Chang 0001

dblp:18/2428-1 · DBLP profile ↗
← Back
227ranked-venue papers
14as first author
162since 2021 · last 2026
0000-0001-5365-0072ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 211 · 14 first-author · 150 since 2021Graphics, computer vision, multimedia, augmented reality and games · 39 · 4 first-author · 33 since 2021Databases, data management, data science and information retrieval · 12 · 4 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Gold-Medal-Level Olympiad Geometry Solving with Efficient Heuristic Auxiliary Constructions
abstract
Boyan Duan, Xiao Liang, Shuai Lu, Yaoxiang Wang, Yelong Shen, Kai-Wei Chang, Ying Nian Wu, Mao Yang, Weizhu Chen, Yeyun Gong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Boyan Duan, Yaoxiang Wang, Yelong Shen, Kai-Wei Chang 0001, Ying Nian Wu, Mao Yang 0004, Weizhu Chen, Yeyun Gong
ACL (1)6
2026 LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMs
abstract
Pei-Fu Guo, Yun-Da Tsai, Chun-Chia Hsu, Kai-Xin Chen, Ya An Tsai, Kai-Wei Chang, Nanyun Peng, Mi-Yen Yeh, Shou-De Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Pei-Fu Guo, Yunda Tsai, Chun-Chia Hsu, Kai-Xin Chen, Ya-An Tsai, Kai-Wei Chang 0001, Nanyun Peng 0001, Mi-Yen Yeh, Shou-De Lin
ACL (1)6
2026 MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Knowledge Poisoning Attacks
abstract
Hyeonjeong Ha, Qiusi Zhan, Jeonghwan Kim, Dimitrios Bralios, Saikrishna Sanniboina, Nanyun Peng, Kai-Wei Chang, Daniel Kang, Heng Ji. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hyeonjeong Ha, Qiusi Zhan, Dimitrios Bralios, Saikrishna Sanniboina, Nanyun Peng 0001, Kai-Wei Chang 0001, Daniel Kang 0001, Heng Ji 0001
ACL (1)7
2026 Dynamic Generation of Multi LLM Agents Communication Topologies with Graph Diffusion Models
abstract
Eric Hanchen Jiang, Levina Li, Frank Wan, Xiao Liang, Sophia Yin, Yuchen Wu, Xinfeng Li, Yizhou Sun, Wei Wang, Kai-Wei Chang, Ying Nian Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Eric Hanchen Jiang, Levina Li, Frank Wan, Sophia Yin, Xinfeng Li, Yizhou Sun, Kai-Wei Chang 0001, Ying Nian Wu
ACL (1)10
2026 Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy
abstract
Eric Hanchen Jiang, Weixuan Ou, Run Liu, Shengyuan Pang, Guancheng Wan, Ranjie Duan, Wei Dong, Kai-Wei Chang, XiaoFeng Wang, Ying Nian Wu, Xinfeng Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Eric Hanchen Jiang, Weixuan Ou, Run Liu 0005, Shengyuan Pang, Guancheng Wan, Ranjie Duan, Wei Dong 0007, Kai-Wei Chang 0001, Xiaofeng Wang 0001, Ying Nian Wu, Xinfeng Li
ACL (1)8
2026 Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability
abstract
Xiao Liang, Zhong-Zhi Li, Zhenghao Lin, Eric Hanchen Jiang, Hengyuan Zhang, Yelong Shen, Kai-Wei Chang, Ying Nian Wu, Yeyun Gong, Weizhu Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhongzhi Li, Zhenghao Lin, Eric Hanchen Jiang, Yelong Shen, Kai-Wei Chang 0001, Ying Nian Wu, Yeyun Gong, Weizhu Chen
ACL (1)7
2026 ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System
abstract
Jiacheng Liang, Yao Ma, Tharindu Kumarage, Satyapriya Krishna, Rahul Gupta, Kai-Wei Chang, Aram Galstyan, Charith Peris. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiacheng Liang, Tharindu Kumarage, Satyapriya Krishna, Rahul Gupta 0001, Kai-Wei Chang 0001, Aram Galstyan, Charith Peris
ACL (1)6
2026 InsideOut: Measuring and Mitigating Insider-Outsider Bias in Interview Script Generation
abstract
Advancements in Large language models (LLMs) have enabled a variety of downstream applications like story and interview script generation.However, recent research raised concerns about culture-related fairness issues in LLM-generated content.In this work, we identify and systematically investigate LLMs' insider-outsider bias, a phenomenon where models position themselves as "insiders" of mainstream cultures during generation while externalizing less dominant cultures.We propose the INSIDEOUT benchmark with 4,000 generation prompts and three evaluation metrics to quantify this bias through a culturally situated interview script generation task, in which an LLM is positioned as a reporter interviewing local people across 10 diverse cultures.Empirical evaluation on 5 state-of-the-art LLMs reveals that while models adopt insider tones in over 88% US-contexted scripts on average, they disproportionately default to "outsider" stances for non-Western cultures.To mitigate these biases, we propose 2 inference-time methods: a baseline prompt-based Fairness Intervention Pillars (FIP) method, and a structured Mitigation via Fairness Agents (MFA) framework consisting of a Single-Agent (MFA-SA), a Hierarchical-Agent (MFA-HA), and an autonomous Agentic Planning (MFA-Plan) pipeline.Empirical results demonstrate that agent-based MFA methods achieve outstanding and robust performance in mitigating the insider-outsider bias: For instance, on the Cultural Alignment Gap (CAG) metric, MFA-SA reduces bias in Llama model by 89.70 % and MFA-HA mitigates bias in Qwen by 82.54%.These findings showcase the effectiveness of agent-based methods as a promising direction for mitigating biases in generative LLMs. * Equal contribution.
Yixin Wan, Xingrun Chen, Kai-Wei Chang 0001
ACL (1)3
2026 VisRet: Visualization Improves Knowledge-Intensive Text-to-Image Retrieval
abstract
Text-to-image retrieval (T2I retrieval) remains challenging because cross-modal embeddings often behave as bags of concepts, underrepresenting structured visual relationships such as pose and viewpoint. We proposeVisualize-then-Retrieve (VisRet), a retrieval paradigm that mitigates this limitation of cross-modal similarity alignment. VisRet first projects textual queries into the image modality via T2I generation, then performs retrieval within the image modality to bypass the weaknesses of cross-modal retrievers in recognizing subtle visual-spatial features. Across four benchmarks (Visual-RAG, INQUIRE-Rerank, Microsoft COCO, and our new Visual-RAG-ME featuring multi-entity comparisons), VisRet substantially outperforms cross-modal similarity matching and baselines that recast T2I retrieval as text-to-text similarity matching, improving nDCG@30 by 0.125 on average with CLIP as the retriever and by 0.121 with E5-V. For downstream question answering, VisRet increases accuracy on Visual-RAG and Visual-RAG-ME by 3.8% and 15.7% in top-1 retrieval, and by 3.9% and 11.1% in top-10 retrieval. Ablation studies show compatibility with different T2I instruction LLMs, T2I generation models, and downstream LLMs. VisRet provides a simple yet effective perspective for advancing in text-image retrieval. Our code and the new benchmark are publicly available at https://github.com/xiaowu0162/Visualize-then-Retrieve.
Di Wu 0054, Yixin Wan, Kai-Wei Chang 0001
ACL (1)3
2026 SWAN: Semantic Watermarking with Abstract Meaning Representation
abstract
Ziping Ye, Gourab Dey, Christos Christodoulopoulos, Charith Peris, Anil Ramakrishna, Weitong Ruan, Aram Galstyan, Kai-Wei Chang, Rahul Gupta, Ninareh Mehrabi. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ziping Ye, Gourab Dey, Christos Christodoulopoulos 0001, Charith Peris, Anil Ramakrishna, Weitong Ruan, Aram Galstyan, Kai-Wei Chang 0001, Rahul Gupta 0001, Ninareh Mehrabi
ACL (1)8
2026 Balancing Efficiency and Empathy: Healthcare Providers' Perspectives on AI-Supported Workflows for Serious Illness Conversations in the Emergency Department
abstract
Serious Illness Conversations (SICs)—discussions about values and care preferences for patients with life-threatening illness—rarely occur in Emergency Departments (EDs), despite evidence that early conversations improve care alignment and reduce unnecessary interventions. We interviewed 11 ED providers to identify challenges in SICs and opportunities for technology support, with a focus on AI. Our analysis revealed a four-stage SIC workflow (identification, preparation, conduction, documentation) and barriers at each stage, including fragmented patient information, limited time and space, lack of conversational guidance, and burdensome documentation. Providers expressed interest in AI systems for synthesizing information, supporting real-time conversations, and automating documentation, but emphasized concerns about preserving human connection and clinical autonomy. This tension highlights the need for technologies that enhance efficiency without undermining the interpersonal nature of SICs. We propose design guidelines for ambient and peripheral AI systems to support providers while preserving the essential humanity of these conversations.
Menglin Zhao, Zhuorui Yong, Ruijian Hannah Guan, Kai-Wei Chang 0001, Adrian Haimovich, Kei Ouchi, Timothy W. Bickmore, Zhan Zhang 0008, Bingsheng Yao, Dakuo Wang, Smit Desai
CHI4
2026 Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
abstract
Abstract The lack of reasoning capabilities in Vision-Language Models (VLMs) has remained at the forefront of research discourse. We posit that this behavior stems from a reporting bias in their training data. That is, how people communicate about visual content by default omits tacit information needed to supervise some types of reasoning; e.g., “at the game today!” is a more likely caption than “a photo of 37 people standing behind a field”. We investigate the data underlying the popular VLMs OpenCLIP, LLaVA-1.5 and Molmo through the lens of theories from pragmatics, and find that reporting bias results in insufficient representation of four reasoning skills (spatial, temporal, negation, and counting), despite the corpora being of web-scale, and/or synthetically generated. With a set of curated benchmarks, we demonstrate that: (i) VLMs perform poorly on the aforementioned types of reasoning suppressed in the training data by reporting bias; (ii) contrary to popular belief, scaling data size, model size, and to multiple languages does not result in emergence of these skills by default; but, promisingly, (iii) incorporating annotations specifically collected to obtain tacit information is effective. Our findings highlight the need for more intentional training data curation methods, rather than counting on scale for emergence of reasoning capabilities.
Amita Kamath, Jack Hessel, Khyathi Raghavi Chandu, Jena D. Hwang, Kai-Wei Chang 0001, Ranjay Krishna
Trans. Assoc. Comput. Linguistics5
2025 SYNTHIA: Novel Concept Design with Affordance Composition
abstract
Hyeonjeong Ha, Xiaomeng Jin, Jeonghwan Kim, Jiateng Liu, Zhenhailong Wang, Khanh Duy Nguyen, Ansel Blume, Nanyun Peng, Kai-Wei Chang, Heng Ji. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Hyeonjeong Ha, Xiaomeng Jin, Jiateng Liu, Zhenhailong Wang, Khanh Duy Nguyen, Ansel Blume, Nanyun Peng 0001, Kai-Wei Chang 0001, Heng Ji 0001
ACL (1)9
2025 METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling
abstract
Chart generation aims to generate code to produce charts satisfying the desired visual properties, e.g., texts, layout, color, and type.It has great potential to empower the automatic professional report generation in financial analysis, research presentation, education, and healthcare.In this work, we build a vision-language model (VLM) based multi-agent framework for effective automatic chart generation.Generating high-quality charts requires both strong visual design skills and precise coding capabilities that embed the desired visual properties into code.Such a complex multi-modal reasoning process is difficult for direct prompting of VLMs.To resolve these challenges, we propose METAL (Multi-agEnT frAmework with vision Language models for chart generation), a multi-agent framework that decomposes the task of chart generation into the iterative collaboration among specialized agents.METAL achieves a 5.2% improvement in the F1 score over the current best result in the chart generation task.Additionally, METAL improves chart generation performance by 11.33% over Direct Prompting with LLAMA 3.2-11B.Furthermore, the METAL framework exhibits the phenomenon of test-time scaling: its performance increases monotonically as the logarithm of computational budget grows from 2 9 to 2 13 tokens.
Yiwei Wang 0001, Jiuxiang Gu, Kai-Wei Chang 0001, Nanyun Peng 0001
ACL (1)4
2025 Vulnerability of LLMs to Vertically Aligned Text Manipulations
abstract
Zhecheng Li, Yiwei Wang, Bryan Hooi, Yujun Cai, Zhen Xiong, Nanyun Peng, Kai-Wei Chang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhecheng Li, Yiwei Wang 0001, Bryan Hooi, Yujun Cai, Zhen Xiong, Nanyun Peng 0001, Kai-Wei Chang 0001
ACL (1)7
2025 White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs
abstract
Social biases can manifest in language agency. However, very limited research has investigated such biases in Large Language Model (LLM)-generated content. In addition, previous works often rely on string-matching techniques to identify agentic and communal words within texts, falling short of accurately classifying language agency. We introduce the Language Agency Bias Evaluation (LABE) benchmark, which comprehensively evaluates biases in LLMs by analyzing agency levels attributed to different demographic groups in model generations. LABE tests for gender, racial, and intersectional language agency biases in LLMs on 3 text generation tasks: biographies, professor reviews, and reference letters. Using LABE, we unveil language agency social biases in 3 recent LLMs: ChatGPT, Llama3, and Mistral. We observe that: (1) LLM generations tend to demonstrate greater gender bias than human-written texts; (2) Models demonstrate remarkably higher levels of intersectional bias than the other bias aspects. (3) Prompt-based mitigation is unstable and frequently leads to bias exacerbation. Based on our observations, we propose Mitigation via Selective Rewrite (MSR), a novel bias mitigation strategy that leverages an agency classifier to identify and selectively revise parts of generated texts that demonstrate communal traits. Empirical results prove MSR to be more effective and reliable than prompt-based mitigation method, showing a promising research direction.
Yixin Wan, Kai-Wei Chang 0001
ACL (1)2
2025 The Male CEO and the Female Assistant: Evaluation and Mitigation of Gender Biases in Text-To-Image Generation of Dual Subjects
abstract
Recent large-scale T2I models like DALLE-3 have made progress in reducing gender stereotypes when generating single-person images. However, significant biases remain when generating images with more than one person. To systematically evaluate this, we propose the Paired Stereotype Test (PST) framework, which queries T2I models to depict two individuals assigned with male-stereotyped and female-stereotyped social identities, respectively (e.g. “a CEO” and “an Assistant”). This contrastive setting often triggers T2I models to generate gender-stereotyped images. Using PST, we evaluate two aspects of gender biases – the well-known bias in gendered occupation and a novel aspect: bias in organizational power. Experiments show that over 74% images generated by DALLE-3 display gender-occupational biases. Additionally, compared to single-person settings, DALLE-3 is more likely to perpetuate male-associated stereotypes under PST. We further propose FairCritic, a novel and interpretable framework that leverages an LLM-based critic model to i) detect bias in generated images, and ii) adaptively provide feedback to T2I models for improving fairness. FairCritic achieves near-perfect fairness on PST, overcoming the limitations of previous prompt-based intervention approaches.
Yixin Wan, Kai-Wei Chang 0001
ACL (1)2
2025 Magnet: Multi-turn Tool-use Data Synthesis and Distillation via Graph Translation
abstract
Fan Yin, Zifeng Wang, I-Hung Hsu, Jun Yan, Ke Jiang, Yanfei Chen, Jindong Gu, Long Le, Kai-Wei Chang, Chen-Yu Lee, Hamid Palangi, Tomas Pfister. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Fan Yin, Zifeng Wang 0002, I-Hung Hsu, Jun Yan 0001, Yanfei Chen, Jindong Gu, Long T. Le, Kai-Wei Chang 0001, Chen-Yu Lee, Hamid Palangi, Tomas Pfister
ACL (1)9
2025 Con-ReCall: Detecting Pre-training Data in LLMs via Contrastive Decoding
abstract
The training data in large language models is key to their success, but it also presents privacy and security risks, as it may contain sensitive information. Detecting pre-training data is crucial for mitigating these concerns. Existing methods typically analyze target text in isolation or solely with non-member contexts, overlooking potential insights from simultaneously considering both member and non-member contexts. While previous work suggested that member contexts provide little information due to the minor distributional shift they induce, our analysis reveals that these subtle shifts can be effectively leveraged when contrasted with non-member contexts. In this paper, we propose Con-ReCall, a novel approach that leverages the asymmetric distributional shifts induced by member and non-member contexts through contrastive decoding, amplifying subtle differences to enhance membership inference. Extensive empirical evaluations demonstrate that Con-ReCall achieves state-of-the-art performance on the WikiMIA benchmark and is robust against various text manipulation techniques.
Yiwei Wang 0001, Bryan Hooi, Yujun Cai, Nanyun Peng 0001, Kai-Wei Chang 0001
COLING6
2025 VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning
abstract
The ability of large vision-language models (LVLMs) to critique and correct their reasoning is an essential building block towards their self-improvement. However, a systematic analysis of such capabilities in LVLMs is still lacking. We propose VISCO, the first benchmark to extensively analyze the fine-grained critique and correction capabilities of LVLMs. Compared to existing work that uses a single scalar value to critique the entire reasoning [4], VISCO features dense and fine-grained critique, requiring LVLMs to evaluate the correctness of each step in the chain-of-thought and provide natural language explanations to support their judgments. Extensive evaluation of 24 LVLMs demonstrates that human-written critiques significantly enhance the performance after correction, showcasing the potential of the self-improvement strategy. However, the model-generated critiques are less helpful and sometimes detrimental to the performance, suggesting that critique is the crucial bottleneck. We identified three common patterns in critique failures: failure to critique visual perception, reluctance to "say no", and exaggerated assumption of error propagation. To address these issues, we propose an effective LookBack strategy that revisits the image to verify each piece of information in the initial reasoning. LookBack significantly improves critique and correction performance by up to 13.5%.
Xueqing Wu 0001, Yuheng Ding, Pan Lu, Da Yin, Kai-Wei Chang 0001, Nanyun Peng 0001
CVPR6
2025 Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models
abstract
Despite inheriting security measures from underlying language models, Vision-Language Models (VLMs) may still be vulnerable to safety alignment issues. Through empirical analysis, we uncover two critical findings: scenario- matched images can significantly amplify harmful outputs, and contrary to common assumptions in gradient-based attacks, minimal loss values do not guarantee optimal attack effectiveness. Building on these insights, we introduce MLAI (Multi-Loss Adversarial Images), a novel jailbreak framework that leverages scenario-aware image generation for semantic alignment, exploits flat minima theory for robust adversarial image selection, and employs multi- image collaborative attacks for enhanced effectiveness. Extensive experiments demonstrate MLAI’s significant impact, achieving attack success rates of 77.75% on MiniGPT-4 and 82.80% on LLaVA-2, substantially outperforming existing methods by margins of 34.37% and 12.77% respectively. Furthermore, MLAI shows considerable transferability to commercial black-box VLMs, achieving up to 60.11% success rate. Our work reveals fundamental visual vulnerabilities in current VLMs safety mechanisms and underscores the need for stronger defenses. Warning: This paper contains potentially harmful example text.
Shuyang Hao, Bryan Hooi, Jun Liu 0036, Kai-Wei Chang 0001, Zi Huang, Yujun Cai
CVPR4
2025 SNaRe: Domain-aware Data Generation for Low-Resource Event Detection
abstract
Event Detection (ED) -the task of identifying event mentions from natural language text -is critical for enabling reasoning in highly specialized domains such as biomedicine, law, and epidemiology.Data generation has proven to be effective in broadening its utility to wider applications without requiring expensive expert annotations.However, when existing generation approaches are applied to specialized domains, they struggle with label noise, where annotations are incorrect, and domain drift, characterized by a distributional mismatch between generated sentences and the target domain.To address these issues, we introduce SNARE, a domain-aware synthetic data generation framework composed of three components: Scout, Narrator, and Refiner.Scout extracts triggers from unlabeled target domain data and curates a high-quality domain-specific trigger list using corpus-level statistics to mitigate domain drift.Narrator, conditioned on these triggers, generates high-quality domainaligned sentences, and Refiner identifies additional event mentions, ensuring high annotation quality.Experimentation on three diverse domain ED datasets reveals how SNARE outperforms the best baseline, achieving average F1 gains of 3-7% in the zero-shot/few-shot settings and 4-20% F1 improvement for multilingual generation.Analyzing the generated trigger hit rate and human evaluation substantiates SNARE's stronger annotation quality and reduced domain drift.We will release our code at https://github.com/PlusLabNLP/SNaRe.
Tanmay Parekh, Lucas Bandarkar, Artin Kim, I-Hung Hsu, Kai-Wei Chang 0001, Nanyun Peng 0001
EMNLP6
2025 DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM Reasoning
abstract
Zero-shot Event Detection (ED), the task of identifying event mentions in natural language text without any training data, is critical for document understanding in specialized domains.Understanding the complex event ontology, extracting domain-specific triggers from the passage, and structuring them appropriately overloads and limits the utility of Large Language Models (LLMs) for zero-shot ED.To this end, we propose DICORE, a divergent-convergent reasoning framework that decouples the task of ED using Dreamer and Grounder.Dreamer encourages divergent reasoning through openended event discovery, which helps to boost event coverage.Conversely, Grounder introduces convergent reasoning to align the freeform predictions with the task-specific instructions using finite-state machine guided constrained decoding.Additionally, an LLM-Judge verifies the final outputs to ensure high precision.Through extensive experiments on six datasets across five domains and nine LLMs, we demonstrate how DICORE consistently outperforms prior zero-shot, transfer-learning, and reasoning baselines, achieving 4-7% average F1 gains over the best baseline -establishing DICORE as a strong zero-shot ED framework.
Tanmay Parekh, Kartik Mehta, Ninareh Mehrabi, Kai-Wei Chang 0001, Nanyun Peng 0001
EMNLP4
2025 STIV: Scalable Text and Image Conditioned Video Generation
abstract
The field of video generation has made remarkable advancements, yet there remains a pressing need for a clear, systematic recipe that can guide the development of robust and scalable models. In this work, we present a comprehensive study that systematically explores the interplay of model architectures, training recipes, and data curation strategies, culminating in a simple and scalable text-image-conditioned video generation method, named STIV. Our framework integrates image condition into a Diffusion Transformer (DiT) through frame replacement, while incorporating text conditioning via a joint image-text conditional classifier-free guidance. This design enables STIV to perform both text-to-video (T2V) and text-image-to-video (TI2V) tasks simultaneously. Additionally, STIV can be easily extended to various applications, such as video prediction, frame interpolation, multi-view generation, and long video generation, etc. With comprehensive ablation studies on T2I, T2V, and TI2V, STIV demonstrate strong performance, despite its simple design. An 8.7B model with 512 resolution achieves 83.1 on VBench T2V, surpassing both leading open and closed-source models like CogVideoX-5B, Pika, Kling, and Gen-3. The same-sized model also achieves a state-of-the-art result of 90.1 on VBench I2V task at 512 resolution. By providing a transparent and extensible recipe for building cutting-edge video generation models, we aim to empower future research and accelerate progress toward more versatile and reliable video generation solutions.
Zongyu Lin, Chen Chen 0005, Jiasen Lu, Wenze Hu, Tsu-Jui Fu, Jesse Allardice, Zhengfeng Lai, Liangchen Song, Bowen Zhang 0002, Cha Chen, Yiran Fei, Lezhi Li, Yinfei Yang, Yizhou Sun, Kai-Wei Chang 0001
ICCV16
2025 Verbalized Representation Learning for Interpretable Few-Shot Generalization
abstract
Humans recognize objects after observing only a few examples, a remarkable capability enabled by their inherent language understanding of the real-world environment. Developing verbalized and interpretable representation can significantly improve model generalization in low-data settings. In this work, we propose Verbalized Representation Learning (VRL), a novel approach for automatically extracting human-interpretable features for object recognition using few-shot data. Our method uniquely captures inter-class differences and intra-class commonalities in the form of natural language by employing a Vision-Language Model (VLM) to identify key discriminative features between different classes and shared characteristics within the same class. These verbalized features are then mapped to numeric vectors through the VLM. The resulting feature vectors can be further utilized to train and infer with downstream classifiers. Experimental results show that, at the same model scale, VRL achieves a 24% absolute improvement over prior state-of-the-art methods while using 95% less data and a smaller mode. Furthermore, compared to human-labeled attributes, the features learned by VRL exhibit a 20% absolute gain when used for downstream classification tasks. Code is available at: https://github.com/joeyy5588/VRL/tree/main.
Cheng-Fu Yang, Da Yin, Wenbo Hu 0006, Heng Ji 0001, Nanyun Peng 0001, Bolei Zhou, Kai-Wei Chang 0001
ICCV7
2025 MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models
abstract
Existing multimodal retrieval benchmarks primarily focus on evaluating whether models can retrieve and utilize external textual knowledge for question answering. However, there are scenarios where retrieving visual information is either more beneficial or easier to access than textual data. In this paper, we introduce a multimodal retrieval-augmented generation benchmark, MRAG-Bench, in which we systematically identify and categorize scenarios where visually augmented knowledge is better than textual knowledge, for instance, more images from varying viewpoints. MRAG-Bench consists of 16,130 images and 1,353 human-annotated multiple-choice questions across 9 distinct scenarios. With MRAG-Bench, we conduct an evaluation of 10 open-source and 4 proprietary large vision-language models (LVLMs). Our results show that all LVLMs exhibit greater improvements when augmented with images compared to textual knowledge, confirming that MRAG-Bench is vision-centric. Additionally, we conduct extensive analysis with MRAG-Bench, which offers valuable insights into retrieval-augmented LVLMs. Notably, the top-performing model, GPT-4o, faces challenges in effectively leveraging retrieved knowledge, achieving only a 5.82\% improvement with ground-truth information, in contrast to a 33.16\% improvement observed in human participants. These findings highlight the importance of MRAG-Bench in encouraging the community to enhance LVLMs' ability to utilize retrieved visual knowledge more effectively.
Wenbo Hu 0006, Jia-Chen Gu, Zi-Yi Dou, Mohsen Fayyaz, Pan Lu, Kai-Wei Chang 0001, Nanyun Peng 0001
ICLR6
2025 Controllable Generation via Locally Constrained Resampling
abstract
Autoregressive models have demonstrated an unprecedented ability at modeling the intricacies of natural language. However, they continue to struggle with generating complex outputs that adhere to logical constraints. Sampling from a fully-independent distribution subject to a constraint is hard. Sampling from an autoregressive distribution subject to a constraint is doubly hard: We have to contend not only with the hardness of the constraint but also the distribution's lack of structure. We propose a tractable probabilistic approach that performs Bayesian conditioning to draw samples subject to a constraint. By factoring in information about the entire sequence, our approach offers better contextual awareness during constrained generation compared to current greedy approaches. Starting from a model sample, we induce a local, factorized distribution which we can tractably condition on the constraint. To generate samples that satisfy the constraint, we sample from the conditional distribution, correct for biases in the sample weights, and resample. The resulting samples closely approximate the target distribution and are guaranteed to satisfy the constraints. We evaluate our approach on several tasks, including LLM detoxification and solving Sudoku puzzles. We show that by disallowing a list of toxic expressions our approach is able to steer the model's outputs away from toxic generations, outperforming similar approaches to detoxification. We also show that our approach achieves a perfect accuracy on Sudoku, compared to less than $50\%$ for GPT4-o and Gemini 1.5.
Kareem Ahmed, Kai-Wei Chang 0001, Guy Van den Broeck
ICLR2
2025 VideoPhy: Evaluating Physical Commonsense for Video Generation
abstract
Recent advances in internet-scale video data pretraining have led to the development of text-to-video generative models that can create high-quality videos across a broad range of visual concepts, synthesize realistic motions and render complex objects. Hence, these generative models have the potential to become general-purpose simulators of the physical world. However, it is unclear how far we are from this goal with the existing text-to-video generative models. To this end, we present VideoPhy, a benchmark designed to assess whether the generated videos follow physical commonsense for real-world activities (e.g. marbles will roll down when placed on a slanted surface). Specifically, we curate diverse prompts that involve interactions between various material types in the physical world (e.g., solid-solid, solid-fluid, fluid-fluid). We then generate videos conditioned on these captions from diverse state-of-the-art text-to-video generative models, including open models (e.g., CogVideoX) and closed models (e.g., Lumiere, Dream Machine). Our human evaluation reveals that the existing models severely lack the ability to generate videos adhering to the given text prompts, while also lack physical commonsense. Specifically, the best performing model, CogVideoX-5B, generates videos that adhere to the caption and physical laws for 39.6% of the instances. VideoPhy thus highlights that the video generative models are far from accurately simulating the physical world. Finally, we propose an auto-evaluator, VideoCon-Physics, to assess the performance reliably for the newly released models. The code is available here: https://github.com/Hritikbansal/videophy.
Hritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong, Michal Yarom, Yonatan Bitton, Chenfanfu Jiang, Yizhou Sun, Kai-Wei Chang 0001, Aditya Grover
ICLR9
2025 SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
abstract
Human beings are endowed with a complementary learning system, which bridges the slow learning of general world dynamics with fast storage of episodic memory from a new experience. Previous video generation models, however, primarily focus on slow learning by pre-training on vast amounts of data, overlooking the fast learning phase crucial for episodic memory storage. This oversight leads to inconsistencies across temporally distant frames when generating longer videos, as these frames fall beyond the model's context window. To this end, we introduce SlowFast-VGen, a novel dual-speed learning system for action-driven long video generation. Our approach incorporates a masked conditional video diffusion model for the slow learning of world dynamics, alongside an inference-time fast learning strategy based on a temporal LoRA module. Specifically, the fast learning process updates its temporal LoRA parameters based on local inputs and outputs, thereby efficiently storing episodic memory in its parameters. We further propose a slow-fast learning loop algorithm that seamlessly integrates the inner fast learning loop into the outer slow learning loop, enabling the recall of prior multi-episode experiences for context-aware skill learning. To facilitate the slow learning of an approximate world model, we collect a large-scale dataset of 200k videos with language action annotations, covering a wide range of scenarios. Extensive experiments show that SlowFast-VGen outperforms baselines across various metrics for action-driven video generation, achieving an FVD score of 514 compared to 782, and maintaining consistency in longer videos, with an average of 0.37 scene cuts versus 0.89. The slow-fast learning loop algorithm significantly enhances performances on long-horizon planning tasks as well.
Yining Hong, Beide Liu, Maxine Wu, Yuanhao Zhai 0001, Kai-Wei Chang 0001, Chung-Ching Lin, Zhengyuan Yang, Ying Nian Wu
ICLR5
2025 Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
abstract
Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spoken language model that comprehends a wide range of natural language instructions is critical for bridging communication gaps and facilitating more intuitive interactions. However, the absence of a comprehensive evaluation benchmark poses a significant challenge. We present Dynamic-SUPERB Phase-2, an open and evolving benchmark for the comprehensive evaluation of instruction-based universal speech models. Building upon the first generation, this second version incorporates 125 new tasks contributed collaboratively by the global research community, expanding the benchmark to a total of 180 tasks, making it the largest benchmark for speech and audio evaluation. While the first generation of Dynamic-SUPERB was limited to classification tasks, Dynamic-SUPERB Phase-2 broadens its evaluation capabilities by introducing a wide array of novel and diverse tasks, including regression and sequence generation, across speech, music, and environmental audio. Evaluation results show that no model performed well universally. SALMONN-13B excelled in English ASR and Qwen2-Audio-7B-Instruct showed high accuracy in emotion recognition, but current models still require further innovations to handle a broader range of tasks. We open-source all task data and the evaluation pipeline at https://github.com/dynamic-superb/dynamic-superb.
Chien-Yu Huang, Wei-Chih Chen, Shu-Wen Yang, Andy T. Liu, Chen-An Li, Yu-Xiang Lin, Wei-Cheng Tseng, Anuj Diwan, Yi-Jen Shih, Jiatong Shi, Chih-Kai Yang, Xuanjun Chen, Chi-Yuan Hsiao, Puyuan Peng, Shih-Heng Wang, Chun-Yi Kuan, Ke-Han Lu, Kai-Wei Chang 0001, Fabian Ritter Gutierrez
ICLR19
2025 MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
abstract
We introduce MuirBench, a comprehensive benchmark that focuses on robust multi-image understanding capabilities of multimodal LLMs. MuirBench consists of 12 diverse multi-image tasks (e.g., scene understanding, ordering) that involve 10 categories of multi-image relations (e.g., multiview, temporal relations). Comprising 11,264 images and 2,600 multiple-choice questions, MuirBench is created in a pairwise manner, where each standard instance is paired with an unanswerable variant that has minimal semantic differences, in order for a reliable assessment. Evaluated upon 20 recent multi-modal LLMs, our results reveal that even the best-performing models like GPT-4o and Gemini Pro find it challenging to solve MuirBench, achieving 68.0% and 49.3% in accuracy. Open-source multimodal LLMs trained on single images can hardly generalize to multi-image questions, hovering below 33.3% in accuracy. These results highlight the importance of MuirBench in encouraging the community to develop multimodal LLMs that can look beyond a single image, suggesting potential pathways for future improvements.
Fei Wang 0060, James Y. Huang, Zekun Li 0007, Qin Liu 0010, Xiaogeng Liu, Mingyu Derek Ma, Nan Xu 0014, Wenxuan Zhou 0002, Kai Zhang 0008, Tianyi Lorena Yan, Wenjie Mo 0001, Hsiang-Hui Liu, Pan Lu, Chunyuan Li, Chaowei Xiao, Kai-Wei Chang 0001, Dan Roth 0001, Sheng Zhang 0012, Hoifung Poon, Muhao Chen 0001
ICLR17
2025 LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
abstract
Recent large language model (LLM)-driven chat assistant systems have integrated memory components to track user-assistant chat histories, enabling more accurate and personalized responses. However, their long-term memory capabilities in sustained interactions remain underexplored. We introduce LongMemEval, a comprehensive benchmark designed to evaluate five core long-term memory abilities of chat assistants: information extraction, multi-session reasoning, temporal reasoning, knowledge updates, and abstention. With 500 meticulously curated questions embedded within freely scalable user-assistant chat histories, LongMemEval presents a significant challenge to existing long-term memory systems, with commercial chat assistants and long-context LLMs showing a 30% accuracy drop on memorizing information across sustained interactions. We then present a unified framework that breaks down the long-term memory design into three stages: indexing, retrieval, and reading. Built upon key experimental insights, we propose several memory design optimizations including session decomposition for value granularity, fact-augmented key expansion for indexing, and time-aware query expansion for refining the search scope. Extensive experiments show that these optimizations greatly improve both memory recall and downstream question answering on LongMemEval. Overall, our study provides valuable resources and guidance for advancing the long-term memory capabilities of LLM-based chat assistants, paving the way toward more personalized and reliable conversational AI. Our benchmark and code are publicly available at https://github.com/xiaowu0162/LongMemEval.
Di Wu 0054, Wenhao Yu 0002, Yuwei Zhang 0001, Kai-Wei Chang 0001, Dong Yu 0001
ICLR5
2025 MQuAKE-Remastered: Multi-Hop Knowledge Editing Can Only Be Advanced with Reliable Evaluations
abstract
Large language models (LLMs) can give out erroneous answers to factually rooted questions either as a result of undesired training outcomes or simply because the world has moved on after a certain knowledge cutoff date. Under such scenarios, *knowledge editing* often comes to the rescue by delivering efficient patches for such erroneous answers without significantly altering the rest, where many editing methods have seen reasonable success when the editing targets are simple and direct (e.g., *``what club does Lionel Messi currently play for?''*). However, knowledge fragments like this are often deeply intertwined in the real world, making effectively propagating the editing effect to non-directly related questions a practical challenge (to entertain an extreme example: [*"What car did the wife of the owner of the club that Messi currently plays for used to get to school in the 80s?"*](youtube.com/watch?v=DbwiHC1Fu-E\&t=132s)). Prior arts have coined this task as *multi-hop knowledge editing* with the most popular dataset being MQuAKE, serving as the sole evaluation benchmark for many later proposed editing methods due to the expensive nature of constructing knowledge editing datasets at scale. In this work, we reveal that **up to 33\% or 76\% of \mquake{}'s questions and ground truth labels are, in fact, corrupted in various fashions due to some unintentional clerical or procedural oversights**. Our work provides a detailed audit of MQuAKE's error pattern and a comprehensive fix without sacrificing its dataset capacity. Additionally, we benchmarked almost all proposed MQuAKE-evaluated editing methods on our post-fix dataset, **MQuAKE-Remastered**. We observe that many methods try to overfit the original MQuAKE by exploiting some dataset idiosyncrasies of MQuAKE. We provide a guideline on how to approach such datasets faithfully and show that a simple, minimally invasive approach — **GWalk** — can offer beyond SOTA editing performance without such exploitation. The MQuAKE-Remastered datasets and utilities are available at [huggingface.co/datasets/henryzhongsc/MQuAKE-Remastered](https://huggingface.co/datasets/henryzhongsc/MQuAKE-Remastered) and [github.com/henryzhongsc/MQuAKE-Remastered](https://github.com/henryzhongsc/MQuAKE-Remastered), respectively.
Shaochen Zhong, Lize Shao, Bhargav Bhushanam, Xiaocong Du, Yixin Wan, Daochen Zha, Yiwei Wang 0001, Ninghao Liu 0001, Kaixiong Zhou, Kai-Wei Chang 0001, Louis Feng, Vipin Chaudhary, Xia Ben Hu
ICLR13
2025 Contrastive Visual Data Augmentation
abstract
Large multimodal models (LMMs) often struggle to recognize novel concepts, as they rely on pre-trained knowledge and have limited ability to capture subtle visual details. Domain-specific knowledge gaps in training also make them prone to confusing visually similar, commonly misrepresented, or low-resource concepts. To help LMMs better align nuanced visual features with language, improving their ability to recognize and reason about novel or rare concepts, we propose a Contrastive visual Data Augmentation (CoDA) strategy. CoDA extracts key contrastive textual and visual features of target concepts against the known concepts they are misrecognized as, and then uses multimodal generative models to produce targeted synthetic data. Automatic filtering of extracted features and augmented images is implemented to guarantee their quality, as verified by human annotators. We show the effectiveness and efficiency of CoDA on low-resource concept and diverse scene recognition datasets including INaturalist and SUN. We additionally collect NovelSpecies, a benchmark dataset consisting of newly discovered animal species that are guaranteed to be unseen by LMMs. LLaVA-1.6 1-shot updating results on these three datasets show CoDA significantly improves SOTA visual data augmentation strategies by 12.3% (NovelSpecies), 5.1% (SUN), and 6.0% (iNat) absolute gains in accuracy.
Yu Zhou 0030, Mohan Tang, Xiaomeng Jin, Te-Lin Wu, Kuan-Hao Huang, Heng Ji 0001, Kai-Wei Chang 0001, Nanyun Peng 0001
ICML8
2025 QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search
abstract
Language agents have become a promising solution to complex interactive tasks. One of the key ingredients to the success of language agents is the reward model on the trajectory of the agentic workflow, which provides valuable guidance during training or inference. However, due to the lack of annotations of intermediate interactions, most existing works use an outcome reward model to optimize policies across entire trajectories. This may lead to sub-optimal policies and hinder the overall performance. To address this, we propose QLASS (Q-guided Language Agent Stepwise Search), to automatically generate annotations by estimating Q-values in a stepwise manner for open language agents. By introducing a reasoning tree and performing process reward modeling, QLASS provides effective intermediate guidance for each step. With the stepwise guidance, we propose a Q-guided generation strategy to enable language agents to better adapt to long-term value, resulting in significant performance improvement during model inference on complex interactive agent tasks. Notably, even with almost half the annotated data, QLASS retains strong performance, demonstrating its efficiency in handling limited supervision. We also empirically demonstrate that QLASS can lead to more effective decision making through qualitative analysis.
Zongyu Lin, Xingcheng Yao, Da Yin, Ziniu Hu, Yizhou Sun, Kai-Wei Chang 0001
ICML7
2025 Contradiction Retrieval via Contrastive Learning with Sparsity
abstract
Contradiction retrieval refers to identifying and extracting documents that explicitly disagree with or refute the content of a query, which is important to many downstream applications like fact checking and data cleaning. To retrieve contradiction argument to the query from large document corpora, existing methods such as similarity search and cross-encoder models exhibit different limitations. To address these challenges, we introduce a novel approach: SparseCL that leverages specially trained sentence embeddings designed to preserve subtle, contradictory nuances between sentences. Our method utilizes a combined metric of cosine similarity and a sparsity function to efficiently identify and retrieve documents that contradict a given query. This approach dramatically enhances the speed of contradiction detection by reducing the need for exhaustive document comparisons to simple vector calculations. We conduct contradiction retrieval experiments on Arguana, MSMARCO, and HotpotQA, where our method produces an average improvement of $11.0%$ across different models. We also validate our method on downstream tasks like natural language inference and cleaning corrupted corpora. This paper outlines a promising direction for non-similarity-based information retrieval which is currently underexplored.
Haike Xu, Zongyu Lin, Kai-Wei Chang 0001, Yizhou Sun, Piotr Indyk
ICML3
2025 Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
Chi-Yuan Hsiao, Ke-Han Lu, Kai-Wei Chang 0001, Chih-Kai Yang, Wei-Chih Chen, Hung-yi Lee
INTERSPEECH3
2025 Unlearning as multi-task optimization: A normalized gradient difference approach with an adaptive learning rate
abstract
Xiaomeng Jin, Zhiqi Bu, Bhanukiran Vinzamuri, Anil Ramakrishna, Kai-Wei Chang, Volkan Cevher, Mingyi Hong. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Xiaomeng Jin, Zhiqi Bu, Bhanukiran Vinzamuri, Anil Ramakrishna, Kai-Wei Chang 0001, Volkan Cevher, Mingyi Hong 0001
NAACL (Long Papers)5
2025 PARTONOMY: Large Multimodal Models with Part-Level Visual Understanding
abstract
Real-world objects are composed of distinctive, object-specific parts. Identifying these parts is key to performing fine-grained, compositional reasoning—yet, large multimodal models (LMMs) struggle to perform this seemingly straightforward task. In this work, we introduce PARTONOMY, an LMM benchmark designed for pixel-level part grounding. We construct PARTONOMY from existing part datasets and our own rigorously annotated set of images, encompassing 862 parts and 5346 objects for evaluation. Unlike existing datasets that simply ask models to identify generic parts, PARTONOMY utilizes highly technical concepts and challenges models to compare objects’ parts, consider part-whole relationships, and justify textual predictions with visual segmentations. Our experiments demonstrate significant limitations in state-of-the-art LMMs (e.g., LISA-13B achieves only 5.9% gIoU), highlighting a critical gap in their part grounding abilities. We note that existing segmentation-enabled LMMs (segmenting LMMs) have two key architectural shortcomings: they use special [SEG] tokens not seen during pretraining which induce distribution shift, and they discard predicted segmentations instead of using past predictions to guide future ones. To address these deficiencies, we train several part-centric LMMs and propose PLUM, a novel segmenting LMM that utilizes span tagging instead of segmentation tokens and that conditions on prior predictions in a feedback loop. We find that pretrained PLUM dominates existing segmenting LMMs on reasoning segmentation, VQA, and visual hallucination benchmarks. In addition, PLUM finetuned on our proposed Explanatory Part Segmentation task is competitive with segmenting LMMs trained on significantly more segmentation data. Our work opens up new avenues towards enabling fine-grained, grounded visual understanding in LMMs.
Ansel Blume, Hyeonjeong Ha, Elen Chatikyan, Xiaomeng Jin, Khanh Duy Nguyen, Nanyun Peng 0001, Kai-Wei Chang 0001, Derek Hoiem, Heng Ji 0001
NeurIPS8
2025 OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
abstract
We introduce *OpenVLThinker*, one of the first open-source large vision–language models (LVLMs) to exhibit sophisticated chain-of-thought reasoning, achieving notable performance gains on challenging visual reasoning tasks. While text-based reasoning models (e.g., Deepseek R1) show promising results in text-only tasks, distilling their reasoning into LVLMs via supervised fine-tuning (SFT) often results in performance degradation due to imprecise visual grounding. Conversely, purely reinforcement learning (RL)-based methods face a large search space, hindering the emergence of reflective behaviors in smaller models (e.g., 7B LVLMs). Surprisingly, alternating between SFT and RL ultimately results in significant performance improvements after a few iterations. Our analysis reveals that the base model rarely exhibits reasoning behaviors initially, but SFT effectively surfaces these latent actions and narrows the RL search space, accelerating the development of reasoning capabilities. Each subsequent RL stage further refines the model's reasoning skills, producing higher-quality SFT data for continued self-improvement. OpenVLThinker-7B consistently advances performance across six benchmarks demanding mathematical and general reasoning, notably improving MathVista by 3.2\%, EMMA by 1.4\%, and HallusionBench by 2.7\%. Beyond demonstrating the synergy between SFT and RL for complex reasoning tasks, our findings provide early evidence towards achieving R1-style reasoning in multimodal contexts.
Yihe Deng, Hritik Bansal, Fan Yin, Nanyun Peng 0001, Wei Wang 0010, Kai-Wei Chang 0001
NeurIPS6
2025 Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence
abstract
AI agents today are mostly siloed — they either retrieve and reason over vast amount of digital information and knowledge obtained online; or interact with the physical world through embodied perception, planning and action — but rarely both. This separation limits their ability to solve tasks that require integrated physical and digital intelligence, such as cooking from online recipes, navigating with dynamic map data, or interpreting real-world landmarks using web knowledge. We introduce \textsc{Embodied Web Agents}, a novel paradigm for AI agents that fluidly bridge embodiment and web-scale reasoning. To operationalize this concept, we first develop the \textsc{Embodied Web Agents} task environments, a unified simulation platform that integrates realistic 3D indoor and outdoor environments with functional web interfaces. Building upon this platform, we construct and release the \textsc{Embodied Web Agents} Benchmark, which encompasses a diverse suite of tasks including cooking, navigation, shopping, tourism, and geolocation — all requiring coordinated reasoning across physical and digital realms for systematic assessment of cross-domain intelligence. Experimental results reveal significant performance gaps between state-of-the-art AI systems and human capabilities, establishing both challenges and opportunities at the intersection of embodied cognition and web-scale knowledge access.
Yining Hong, Rui Sun 0011, Xingcheng Yao, Maxine Wu, Alexander Chien, Da Yin, Ying Nian Wu, Zhecan James Wang, Kai-Wei Chang 0001
NeurIPS10
2025 3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
abstract
Humans excel at performing complex tasks by leveraging long-term memory across temporal and spatial experiences. In contrast, current Large Language Models (LLMs) struggle to effectively plan and act in dynamic, multi-room 3D environments. We posit that part of this limitation is due to the lack of proper 3D spatial-temporal memory modeling in LLMs. To address this, we first introduce 3DMem-Bench, a comprehensive benchmark comprising over 26,000 trajectories and 2,892 embodied tasks, question-answering and captioning, designed to evaluate an agent's ability to reason over long-term memory in 3D environments. Second, we propose 3DLLM-Mem, a novel dynamic memory management and fusion model for embodied spatial-temporal reasoning and actions in LLMs. Our model uses working memory tokens, which represents current observations, as queries to selectively attend to and fuse the most useful spatial and temporal features from episodic memory, which stores past observations and interactions. Our approach allows the agent to focus on task-relevant information while maintaining memory efficiency in complex, long-horizon environments. Experimental results demonstrate that 3DLLM-Mem achieves state-of-the-art performance across various tasks, outperforming the strongest baselines by 16.5\% in success rate on 3DMem-Bench's most challenging in-the-wild embodied tasks.
Wenbo Hu 0006, Yining Hong, Leison Gao, Zibu Wei, Xingcheng Yao, Nanyun Peng 0001, Yonatan Bitton, Idan Szpektor, Kai-Wei Chang 0001
NeurIPS10
2025 LaViDa: A Large Diffusion Language Model for Multimodal Understanding
abstract
Modern Vision-Language Models (VLMs) can solve a wide range of tasks requiring visual reasoning. In real-world scenarios, desirable properties for VLMs include fast inference and controllable generation (e.g., constraining outputs to adhere to a desired format). However, existing autoregressive (AR) VLMs like LLaVA struggle in these aspects. Discrete diffusion models (DMs) offer a promising alternative, enabling parallel decoding for faster inference and bidirectional context for controllable generation through text-infilling. While effective in language-only settings, DMs' potential for multimodal tasks is underexplored. We introduce LaViDa, a family of VLMs built on DMs. We build LaViDa by equipping DMs with a vision encoder and jointly fine-tune the combined parts for multimodal instruction following. To address challenges encountered, LaViDa incorporates novel techniques such as complementary masking for effective training, prefix KV cache for efficient inference, and timestep shifting for high-quality sampling. Experiments show that LaViDa achieves competitive or superior performance to AR VLMs on multi-modal benchmarks such as MMMU, while offering unique advantages of DMs, including flexible speed-quality tradeoff, controllability, and bidirectional reasoning. On COCO captioning, LaViDa surpasses Open-LLaVa-Next-8B by +4.1 CIDEr with 1.92x speedup. On bidirectional tasks, it achieves +59% improvement on Constrained Poem Completion. These results demonstrate LaViDa as a strong alternative to AR VLMs. Code and models is available at https://github.com/jacklishufan/LaViDa
Konstantinos Kallidromitis, Hritik Bansal, Akash Gokul, Yusuke Kato, Kazuki Kozuka, Jason Kuen, Kai-Wei Chang 0001, Aditya Grover
NeurIPS9
2024 Tree-of-Traversals: A Zero-Shot Reasoning Algorithm for Augmenting Black-box Language Models with Knowledge Graphs
abstract
Elan Markowitz, Anil Ramakrishna, Jwala Dhamala, Ninareh Mehrabi, Charith Peris, Rahul Gupta, Kai-Wei Chang, Aram Galstyan. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Elan Markowitz, Anil Ramakrishna, Jwala Dhamala, Ninareh Mehrabi, Charith Peris, Rahul Gupta 0001, Kai-Wei Chang 0001, Aram Galstyan
ACL (1)7
2024 Agent Lumos: Unified and Modular Training for Open-Source Language Agents
abstract
Da Yin, Faeze Brahman, Abhilasha Ravichander, Khyathi Chandu, Kai-Wei Chang, Yejin Choi, Bill Yuchen Lin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Da Yin, Faeze Brahman, Abhilasha Ravichander, Khyathi Raghavi Chandu, Kai-Wei Chang 0001, Yejin Choi 0001, Bill Y. Lin
ACL (1)5
2024 On Leveraging Encoder-only Pre-trained Language Models for Effective Keyphrase Generation
abstract
This study addresses the application of encoder-only Pre-trained Language Models (PLMs) in keyphrase generation (KPG) amidst the broader availability of domain-tailored encoder-only models compared to encoder-decoder models. We investigate three core inquiries: (1) the efficacy of encoder-only PLMs in KPG, (2) optimal architectural decisions for employing encoder-only PLMs in KPG, and (3) a performance comparison between in-domain encoder-only and encoder-decoder PLMs across varied resource settings. Our findings, derived from extensive experimentation in two domains reveal that with encoder-only PLMs, although keyphrase extraction with Conditional Random Fields slightly excels in identifying present keyphrases, the KPG formulation renders a broader spectrum of keyphrase predictions. Additionally, prefix-LM fine-tuning of encoder-only PLMs emerges as a strong and data-efficient strategy for KPG, outperforming general-domain seq2seq PLMs. We also identify a favorable parameter allocation towards model depth rather than width when employing encoder-decoder architectures initialized with encoder-only PLMs. The study sheds light on the potential of utilizing encoder-only PLMs for advancing KPG systems and provides a groundwork for future KPG methods. Our code and pre-trained checkpoints are released at https://github.com/uclanlp/DeepKPG.
Di Wu 0054, Wasi Uddin Ahmad, Kai-Wei Chang 0001
LREC/COLING3
2024 Can Small Language Models Help Large Language Models Reason Better?: LM-Guided Chain-of-Thought
abstract
We introduce a novel framework, LM-Guided CoT, that leverages a lightweight (i.e., <1B) language model (LM) for guiding a black-box large (i.e., >10B) LM in reasoning tasks. Specifically, the lightweight LM first generates a rationale for each input instance. The Frozen large LM is then prompted to predict a task output based on the rationale generated by the lightweight LM. Our approach is resource-efficient in the sense that it only requires training the lightweight LM. We optimize the model through 1) knowledge distillation and 2) reinforcement learning from rationale-oriented and task-oriented reward signals. We assess our method with multi-hop extractive question answering (QA) benchmarks, HotpotQA, and 2WikiMultiHopQA. Experimental results show that our approach outperforms all baselines regarding answer prediction accuracy. We also find that reinforcement learning helps the model to produce higher-quality rationales with improved QA performance.
Fan Yang 0155, Emre Barut, Kai-Wei Chang 0001
LREC/COLING6
2024 Medical Vision-Language Pre-Training for Brain Abnormalities
abstract
Vision-language models have become increasingly powerful for tasks that require an understanding of both visual and linguistic elements, bridging the gap between these modalities. In the context of multimodal clinical AI, there is a growing need for models that possess domain-specific knowledge, as existing models often lack the expertise required for medical applications. In this paper, we take brain abnormalities as an example to demonstrate how to automatically collect medical image-text aligned data for pretraining from public resources such as PubMed. In particular, we present a pipeline that streamlines the pre-training process by initially collecting a large brain image-text dataset from case reports and published journals and subsequently constructing a high-performance vision-language model tailored to specific medical tasks. We also investigate the unique challenge of mapping subfigures to subcaptions in the medical domain. We evaluated the resulting model with quantitative and qualitative intrinsic evaluations. The resulting dataset will be released to the community.
Masoud Monajatipoor, Zi-Yi Dou, Aichi Chien, Nanyun Peng 0001, Kai-Wei Chang 0001
LREC/COLING5
2024 VideoCon: Robust Video-Language Alignment via Contrast Captions
abstract
Despite being (pre)trained on a massive amount of data, state-of-the-art video-language alignment models are not robust to semantically-plausible contrastive changes in the video captions. Our work addresses this by identifying a broad spectrum of contrast misalignments, such as re-placing entities, actions, and flipping event order, which alignment models should be robust against. To this end, we introduce the VideoCon, a video-language alignment dataset constructed by a large language model that gen-erates plausible contrast video captions and explanations for differences between original and contrast video captions. Then, a generative video-language model is fine-tuned with VideoCon to assess video-language entailment and generate explanations. Our VideoCon-based alignment model significantly outperforms current models. It exhibits a 12-point increase in AUC for the video-language alignment task on human-generated contrast captions. Finally, our model sets new state of the art zero-shot performance in temporally-extensive video-language tasks such as text-to-video retrieval (SSv2-Temporal) and video question answering (ATP-Hard). Moreover, our model shows superior performance on novel videos and human-crafted captions and explanations.
Hritik Bansal, Yonatan Bitton, Idan Szpektor, Kai-Wei Chang 0001, Aditya Grover
CVPR4
2024 The Hard Positive Truth About Vision-Language Compositionality
Amita Kamath, Cheng-Yu Hsieh, Kai-Wei Chang 0001, Ranjay Krishna
ECCV (14)3
2024 MATHVERSE: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
Renrui Zhang, Dongzhi Jiang, Haokun Lin, Pengshuo Qiu, Aojun Zhou, Pan Lu, Kai-Wei Chang 0001, Yu Qiao 0001, Peng Gao 0007, Hongsheng Li 0001
ECCV (8)9
2024 Re-ReST: Reflection-Reinforced Self-Training for Language Agents
abstract
Finetuning language agents with reasoningaction trajectories is effective, but obtaining these trajectories from human annotations or stronger models is costly and sometimes impractical.In this paper, we investigate the use of self-training in language agents, which can generate supervision from the agent itself, offering a promising alternative without relying on human or stronger model demonstrations.Self-training, however, requires high-quality model-generated samples, which are hard to obtain for challenging language agent tasks.To address this, we present Reflection-Reinforced Self-Training (Re-ReST), which uses a reflector to refine low-quality generated samples during self-training.The reflector takes the agent's output and feedback from an external environment (e.g., unit test results in code generation) to produce improved samples.This technique enhances the quality of inferior samples and efficiently enriches the self-training dataset with higher-quality samples.We conduct extensive experiments on open-source language agents across tasks, including multi-hop question answering, sequential decision-making, code generation, visual question answering, and text-toimage generation.The results demonstrate the effectiveness of self-training and Re-ReST in language agent tasks, with self-training improving baselines by 7.6% on HotpotQA and 28.4% on AlfWorld, and Re-ReST further boosting performance by 2.0% and 14.1%, respectively.Our studies also confirm the efficiency of using a reflector to generate high-quality samples for self-training.Moreover, we demonstrate a method to employ reflection during inference without ground-truth feedback, addressing the limitation of previous reflection work.
Zi-Yi Dou, Cheng-Fu Yang, Xueqing Wu 0001, Kai-Wei Chang 0001, Nanyun Peng 0001
EMNLP4
2024 Model Editing Harms General Abilities of Large Language Models: Regularization to the Rescue
abstract
Model editing is a technique that edits the large language models (LLMs) with updated knowledge to alleviate hallucinations without resource-intensive retraining.While current model editing methods can effectively modify a model's behavior within a specific area of interest, they often overlook the potential unintended side effects on the general abilities of LLMs such as reasoning, natural language inference, and question answering.In this paper, we raise concerns that model editing's improvements on factuality may come at the cost of a significant degradation of the model's general abilities.We systematically analyze the side effects by evaluating four popular editing methods on three LLMs across eight representative tasks.Our extensive empirical experiments show that it is challenging for current editing methods to simultaneously improve factuality of LLMs and maintain their general abilities.Our analysis reveals that the side effects are caused by model editing altering the original model weights excessively, leading to overfitting to the edited facts.To mitigate this, a method named RECT is proposed to regularize the edit update weights by imposing constraints on their complexity based on the RElative Change in weighT.Evaluation results show that RECT can significantly mitigate the side effects of editing while still maintaining over 94% editing performance 1 .
Jia-Chen Gu, Hao-Xiang Xu, Jun-Yu Ma, Pan Lu, Zhen-Hua Ling, Kai-Wei Chang 0001, Nanyun Peng 0001
EMNLP6
2024 Control Large Language Models via Divide and Conquer
abstract
This paper investigates controllable generation for large language models (LLMs) with prompt-based control, focusing on Lexically Constrained Generation (LCG).We systematically evaluate the performance of LLMs on satisfying lexical constraints with prompt-based control, as well as their efficacy in downstream applications.We conclude that LLMs face significant challenges in consistently satisfying lexical constraints with prompt-based control.We identified three key limitations of LLMs for LCG, including (1) position bias, where LLMs tend to satisfy constraints that appear in specific positions within the input; (2) low responsiveness to decoding parameters, which render minimal impact on control of LLMs; and (3) struggle with handling the inherent complexity of certain constraints (e.g., compound words).To address these issues, we introduce a Divide and Conquer Generation strategy, effective for both white-box and black-box LLMs, to enhance LLMs performance in LCG tasks, which demonstrates over 90% improvement on success rate in the most challenging LCG task.Our analysis provides valuable insights into the performance of LLMs in LCG with prompt-based control, and our proposed strategy offers a pathway to more sophisticated and customized text generation applications.
Yiwei Wang 0001, Kai-Wei Chang 0001, Nanyun Peng 0001
EMNLP4
2024 FLIRT: Feedback Loop In-context Red Teaming
abstract
Ninareh Mehrabi, Palash Goyal, Christophe Dupuy, Qian Hu, Shalini Ghosh, Richard Zemel, Kai-Wei Chang, Aram Galstyan, Rahul Gupta. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Ninareh Mehrabi, Palash Goyal, Christophe Dupuy, Shalini Ghosh, Richard S. Zemel, Kai-Wei Chang 0001, Aram Galstyan, Rahul Gupta 0001
EMNLP7
2024 SPEED++: A Multilingual Event Extraction Framework for Epidemic Prediction and Preparedness
abstract
Tanmay Parekh, Jeffrey Kwan, Jiarui Yu, Sparsh Johri, Hyosang Ahn, Sreya Muppalla, Kai-Wei Chang, Wei Wang, Nanyun Peng. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Tanmay Parekh, Jeffrey Kwan, Jiarui Yu, Sparsh Johri, Hyosang Ahn, Sreya Muppalla, Kai-Wei Chang 0001, Wei Wang 0010, Nanyun Peng 0001
EMNLP7
2024 QUDSELECT: Selective Decoding for Questions Under Discussion Parsing
abstract
Question Under Discussion (QUD) is a discourse framework that uses implicit questions to reveal discourse relationships between sentences.In QUD parsing, each sentence is viewed as an answer to a question triggered by an anchor sentence in prior context.The resulting QUD structure is required to conform to several theoretical criteria like answer compatibility (how well the question is answered), making QUD parsing a challenging task.Previous works construct QUD parsers in a pipelined manner (i.e.detect the trigger sentence in context and then generate the question).However, these parsers lack a holistic view of the task and can hardly satisfy all the criteria.In this work, we introduce QUDSELECT, a joint-training framework that selectively decodes the QUD dependency structures considering the QUD criteria.Using instruction-tuning, we train models to simultaneously predict the anchor sentence and generate the associated question.To explicitly incorporate the criteria, we adopt a selective decoding strategy of sampling multiple QUD candidates during inference, followed by selecting the best one with criteria scorers.Our method outperforms the state-of-the-art baseline models by 9% in human evaluation and 4% in automatic evaluation, demonstrating the effectiveness of our framework.Code and data are in https://github.com/ asuvarna31/qudselect.
Ashima Suvarna, Xiao Liu 0032, Tanmay Parekh, Kai-Wei Chang 0001, Nanyun Peng 0001
EMNLP4
2024 The Factuality Tax of Diversity-Intervened Text-to-Image Generation: Benchmark and Fact-Augmented Intervention
abstract
Prompt-based “diversity interventions” are commonly adopted to improve the diversity of Text-to-Image (T2I) models depicting individuals with various racial or gender traits. However, will this strategy result in nonfactual demographic distribution, especially when generating real historical figures? In this work, we propose DemOgraphic FActualIty Representation (DoFaiR), a benchmark to systematically quantify the trade-off between using diversity interventions and preserving demographic factuality in T2I models. DoFaiR consists of 756 meticulously fact-checked test instances to reveal the factuality tax of various diversity prompts through an automated evidence-supported evaluation pipeline. Experiments on DoFaiR unveil that diversity-oriented instructions increase the number of different gender and racial groups in DALLE-3’s generations at the cost of historically inaccurate demographic distributions. To resolve this issue, we propose Fact-Augmented Intervention (FAI), which instructs a Large Language Model (LLM) to reflect on verbalized or retrieved factual information about gender and racial compositions of generation subjects in history, and incorporate it into the generation context of T2I models. By orienting model generations using the reflected historical truths, FAI significantly improves the demographic factuality under diversity interventions while preserving diversity.
Yixin Wan, Di Wu 0054, Haoran Wang 0005, Kai-Wei Chang 0001
EMNLP4
2024 Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
abstract
Data is a crucial element in large language model (LLM) alignment.Recent studies have explored using LLMs for efficient data collection.However, LLM-generated data often suffers from quality issues, with underrepresented or absent aspects and low-quality datapoints.To address these problems, we propose DATA ADVISOR, an enhanced LLMbased method for generating data that takes into account the characteristics of the desired dataset.Starting from a set of pre-defined principles in hand, DATA ADVISOR monitors the status of the generated data, identifies weaknesses in the current dataset, and advises the next iteration of data generation accordingly.DATA ADVISOR can be easily integrated into existing data generation methods to enhance data quality and coverage.Experiments on safety alignment of three representative LLMs (i.e., Mistral, Llama2, and Falcon) demonstrate the effectiveness of DATA ADVISOR in enhancing model safety against various fine-grained safety issues without sacrificing model utility.Warning: this paper contains example data that may be offensive or harmful.
Fei Wang 0060, Ninareh Mehrabi, Palash Goyal, Rahul Gupta 0001, Kai-Wei Chang 0001, Aram Galstyan
EMNLP5
2024 Synchronous Faithfulness Monitoring for Trustworthy Retrieval-Augmented Generation
abstract
Retrieval-augmented language models (RALMs) have shown strong performance and wide applicability in knowledge-intensive tasks.However, there are significant trustworthiness concerns as RALMs are prone to generating unfaithful outputs, including baseless information or contradictions with the retrieved context.This paper proposes SYNCHECK, a lightweight monitor that leverages fine-grained decoding dynamics including sequence likelihood, uncertainty quantification, context influence, and semantic alignment to synchronously detect unfaithful sentences.By integrating efficiently measurable and complementary signals, SYNCHECK enables accurate and immediate feedback and intervention, achieving 0.85 AUROC in detecting faithfulness errors across six long-form retrieval-augmented generation tasks, improving prior best method by 4%.Leveraging SYNCHECK, we further introduce FOD, a faithfulness-oriented decoding algorithm guided by beam search for long-form retrieval-augmented generation.Empirical results demonstrate that FOD outperforms traditional strategies such as abstention, reranking, or contrastive decoding significantly in terms of faithfulness, achieving over 10% improvement across six datasets.
Di Wu 0054, Jia-Chen Gu, Fan Yin, Nanyun Peng 0001, Kai-Wei Chang 0001
EMNLP5
2024 Dynamic-Superb: Towards a Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark For Speech
abstract
Text language models have shown remarkable zero-shot capability in generalizing to unseen tasks when provided with well-formulated instructions. However, existing studies in speech processing primarily focus on limited or specific tasks. Moreover, the lack of standardized benchmarks hinders a fair comparison across different approaches. Thus, we present Dynamic-SUPERB, a benchmark designed for building universal speech models capable of leveraging instruction tuning to perform multiple tasks in a zero-shot fashion. To achieve comprehensive coverage of diverse speech tasks and harness instruction tuning, we invite the community to collaborate and contribute, facilitating the dynamic growth of the benchmark. To initiate, Dynamic-SUPERB features 55 evaluation instances by combining 33 tasks and 22 datasets. This spans a broad spectrum of dimensions, providing a comprehensive platform for evaluation. Additionally, we propose several approaches to establish benchmark baselines. These include the utilization of speech models, text language models, and the multimodal encoder. Evaluation results indicate that while these baselines perform reasonably on seen tasks, they struggle with unseen ones. We release all materials to the public and welcome researchers to collaborate on the project, advancing technologies in the field together1.
Chien-Yu Huang, Ke-Han Lu, Shih-Heng Wang, Chi-Yuan Hsiao, Chun-Yi Kuan, Siddhant Arora, Kai-Wei Chang 0001, Jiatong Shi, Yifan Peng 0003, Roshan S. Sharma, Shinji Watanabe 0001, Bhiksha Raj, Shady Shehata, Hung-yi Lee
ICASSP8
2024 MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
abstract
Large Language Models (LLMs) and Large Multimodal Models (LMMs) exhibit impressive problem-solving skills in many tasks and domains, but their ability in mathematical reasoning in visual contexts has not been systematically studied. To bridge this gap, we present MathVista, a benchmark designed to combine challenges from diverse mathematical and visual tasks. It consists of 6,141 examples, derived from 28 existing multimodal datasets involving mathematics and 3 newly created datasets (i.e., IQTest, FunctionQA, and PaperQA). Completing these tasks requires fine-grained, deep visual understanding and compositional reasoning, which all state-of-the-art foundation models find challenging. With MathVista, we have conducted a comprehensive, quantitative evaluation of 12 prominent foundation models. The best-performing GPT-4V model achieves an overall accuracy of 49.9%, substantially outperforming Bard, the second-best performer, by 15.1%. Our in-depth analysis reveals that the superiority of GPT-4V is mainly attributed to its enhanced visual perception and mathematical reasoning. However, GPT-4V still falls short of human performance by 10.4%, as it often struggles to understand complex figures and perform rigorous reasoning. This significant gap underscores the critical role that MathVista will play in the development of general-purpose AI agents capable of tackling mathematically intensive and visually rich real-world tasks. We further explore the new ability of self-verification, the application of self-consistency, and the interactive chatbot capabilities of GPT-4V, highlighting its promising potential for future research. The project is available at https://mathvista.github.io/.
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 0010, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng 0002, Kai-Wei Chang 0001, Michel Galley, Jianfeng Gao 0001
ICLR8
2024 CoBIT: A Contrastive Bi-directional Image-Text Generation Model
abstract
The field of Vision-and-Language (VL) has witnessed a proliferation of pretrained foundation models. Current techniques typically employ only one type of training objective, whether it's (1) contrastive objectives (like CLIP), (2) image-to-text generative objectives (like PaLI), or (3) text-to-image generative objectives (like Parti). However, all these three objectives are mutually relevant and are all based on image-text pairs. Intuitively, the first two objectives can be considered as complementary projections between two modalities, and contrastive learning can preserve global alignment and generations facilitate fine-grained understanding. Inspired by this, we present a Contrastive Bi-directional Image-Text generation model (CoBIT) to first time unify the three pre-training objectives in one framework. Specifically, CoBIT employs a novel unicoder-decoder structure consisting of an image unicoder, a text unicoder, and a cross-modal decoder. The image/text unicoders can switch between encoding and decoding in different tasks, enabling flexibility and shared knowledge that benefits both image-to-text and text-to-image generations. CoBIT achieves superior performance in image understanding, image-text understanding (Retrieval, Captioning, VQA, SNLI-VE), and text-based content creation, particularly in zero-shot scenarios.
Haoxuan You, Mandy Guo, Zhecan Wang, Kai-Wei Chang 0001, Jason Baldridge
ICLR4
2024 Position: TrustLLM: Trustworthiness in Large Language Models
abstract
Large language models (LLMs) have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs present many challenges, particularly in the realm of trustworthiness. This paper introduces TrustLLM, a comprehensive study of trustworthiness in LLMs, including principles for different dimensions of trustworthiness, established benchmark, evaluation, and analysis of trustworthiness for mainstream LLMs, and discussion of open challenges and future directions. Specifically, we first propose a set of principles for trustworthy LLMs that span eight different dimensions. Based on these principles, we further establish a benchmark across six dimensions including truthfulness, safety, fairness, robustness, privacy, and machine ethics. We then present a study evaluating 16 mainstream LLMs in TrustLLM, consisting of over 30 datasets. Our findings firstly show that in general trustworthiness and capability (i.e., functional effectiveness) are positively related. Secondly, our observations reveal that proprietary LLMs generally outperform most open-source counterparts in terms of trustworthiness, raising concerns about the potential risks of widely accessible open-source LLMs. However, a few open-source LLMs come very close to proprietary ones, suggesting that open-source models can achieve high levels of trustworthiness without additional mechanisms like moderator, offering valuable insights for developers in this field. Thirdly, it is important to note that some LLMs may be overly calibrated towards exhibiting trustworthiness, to the extent that they compromise their utility by mistakenly treating benign prompts as harmful and consequently not responding. Besides these observations, we’ve uncovered key insights into the multifaceted trustworthiness in LLMs. We emphasize the importance of ensuring transparency not only in the models themselves but also in the technologies that underpin trustworthiness. We advocate that the establishment of an AI alliance between industry, academia, the open-source community to foster collaboration is imperative to advance the trustworthiness of LLMs.
Yue Huang 0001, Lichao Sun 0001, Haoran Wang 0005, Siyuan Wu 0001, Qihui Zhang, Chujie Gao, Wenhan Lyu, Yixuan Zhang 0001, Xiner Li, Hanchi Sun, Zhengliang Liu, Yixin Liu 0002, Yijue Wang, Bertie Vidgen, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric P. Xing, Furong Huang, Heng Ji 0001, Hongyi Wang 0001, Huan Zhang 0001, Huaxiu Yao, Manolis Kellis, Marinka Zitnik, Meng Jiang 0001, Mohit Bansal, James Zou 0001, Jian Pei 0001, Jianfeng Gao 0001, Jiawei Han 0001, Jieyu Zhao 0001, Jiliang Tang, Jindong Wang 0001, Joaquin Vanschoren, John C. Mitchell, Kai Shu, Kaidi Xu, Kai-Wei Chang 0001, Lifang He 0001, Lifu Huang, Michael Backes 0001, Neil Zhenqiang Gong, Philip S. Yu, Quanquan Gu, Ran Xu 0001, Rex Ying, Shuiwang Ji, Suman Jana, Tianlong Chen 0001, Tianming Liu 0001, Tianyi Zhou 0001, William Yang Wang, Xiang Li 0001, Xiangliang Zhang 0001, Xiao Wang 0012, Xing Xie 0001, Xuyu Wang, Yan Liu 0002, Yanfang Ye 0001, Yinzhi Cao, Yong Chen 0016, Yue Zhao 0016
ICML45
2024 ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
abstract
Many real-world tasks require an agent to reason jointly over text and visual objects, (e.g., navigating in public spaces), which we refer to as context-sensitive text-rich visual reasoning. Specifically, these tasks require an understanding of the context in which the text interacts with visual elements within an image. However, there is a lack of existing datasets to benchmark the state-of-the-art multimodal models' capability on context-sensitive text-rich visual reasoning. In this paper, we introduce ConTextual, a novel dataset featuring human-crafted instructions that require context-sensitive reasoning for text-rich images. We conduct experiments to assess the performance of 14 foundation models (GPT-4V, Gemini-Pro-Vision, LLaVA-Next) and establish a human performance baseline. Further, we perform human evaluations of the model responses and observe a significant performance gap of 30.8% between GPT-4V (the current best-performing Large Multimodal Model) and human performance. Our fine-grained analysis reveals that GPT-4V encounters difficulties interpreting time-related data and infographics. However, it demonstrates proficiency in comprehending abstract visual contexts such as memes and quotes. Finally, our qualitative analysis uncovers various factors contributing to poor performance including lack of precise visual perception and hallucinations. Our dataset, code, and leaderboard can be found on the project page https://con-textual.github.io/.
Rohan Wadhawan, Hritik Bansal, Kai-Wei Chang 0001, Nanyun Peng 0001
ICML3
2024 Characterizing Truthfulness in Large Language Model Generations with Local Intrinsic Dimension
abstract
We study how to characterize and predict the truthfulness of texts generated from large language models (LLMs), which serves as a crucial step in building trust between humans and LLMs. Although several approaches based on entropy or verbalized uncertainty have been proposed to calibrate model predictions, these methods are often intractable, sensitive to hyperparameters, and less reliable when applied in generative tasks with LLMs. In this paper, we suggest investigating internal activations and quantifying LLM’s truthfulness using the local intrinsic dimension (LID) of model activations. Through experiments on four question answering (QA) datasets, we demonstrate the effectiveness of our proposed method. Additionally, we study intrinsic dimensions in LLMs and their relations with model layers, autoregressive language modeling, and the training of LLMs, revealing that intrinsic dimensions can be a powerful approach to understanding LLMs.
Fan Yin, Jayanth Srinivasa, Kai-Wei Chang 0001
ICML3
2024 On Prompt-Driven Safeguarding for Large Language Models
abstract
Prepending model inputs with safety prompts is a common practice for safeguarding large language models (LLMs) against queries with harmful intents. However, the underlying working mechanisms of safety prompts have not been unraveled yet, restricting the possibility of automatically optimizing them to improve LLM safety. In this work, we investigate how LLMs’ behavior (i.e., complying with or refusing user queries) is affected by safety prompts from the perspective of model representation. We find that in the representation space, the input queries are typically moved by safety prompts in a "higher-refusal" direction, in which models become more prone to refusing to provide assistance, even when the queries are harmless. On the other hand, LLMs are naturally capable of distinguishing harmful and harmless queries without safety prompts. Inspired by these findings, we propose a method for safety prompt optimization, namely DRO (Directed Representation Optimization). Treating a safety prompt as continuous, trainable embeddings, DRO learns to move the queries’ representations along or opposite the refusal direction, depending on their harmfulness. Experiments with eight LLMs on out-of-domain and jailbreak benchmarks demonstrate that DRO remarkably improves the safeguarding performance of human-crafted safety prompts, without compromising the models’ general performance.
Chujie Zheng, Fan Yin, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Kai-Wei Chang 0001, Minlie Huang, Nanyun Peng 0001
ICML6
2024 Exploring In-Context Learning of Textless Speech Language Model for Speech Classification Tasks
Kai-Wei Chang 0001, Ming-Hao Hsu, Shang-Wen Li 0001, Hung-yi Lee
INTERSPEECH1
2024 MIRACLE: An Online, Explainable Multimodal Interactive Concept Learning System
abstract
We present MIRACLE, a system for online, interpretable visual concept and video action recognition. Through a chat interface, users query the recognition system with an uploaded image or video. For images, MIRACLE returns concept predictions from its structured knowledge base, justifying its predictions with heatmaps and natural language-based attribute detections. For videos, MIRACLE predicts an action and justifies its prediction with time varying entity-entity relations. With its ability to learn new concepts in an online, few-shot manner and its support of dynamic changes to its knowledge base, MIRACLE represents a step forward in interpretable multimodal learning systems.
Ansel Blume, Khanh Duy Nguyen, Zhenhailong Wang, Yangyi Chen, Michal Shlapentokh-Rothman, Xiaomeng Jin, Zhen Zhu 0006, Jiateng Liu, Kuan-Hao Huang, Mankeerat Sidhu, Xuanming Zhang, Vivian Liu, Raunak Sinha, Te-Lin Wu, Abhaysinh Zala, Elias Stengel-Eskin, Da Yin, Utkarsh Mall, Zhou Yu 0005, Kai-Wei Chang 0001, Camille Cobb, Karrie Karahalios, Lydia B. Chilton, Mohit Bansal, Nanyun Peng 0001, Carl Vondrick, Derek Hoiem, Heng Ji 0001
ACM Multimedia22
2024 Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
Junzhang Liu, Zhecan Wang, Hammad A. Ayyubi, Haoxuan You, Christopher Thomas 0004, Rui Sun 0011, Shih-Fu Chang, Kai-Wei Chang 0001
ACM Multimedia8
2024 The steerability of large language models toward data-driven personas
abstract
Junyi Li, Charith Peris, Ninareh Mehrabi, Palash Goyal, Kai-Wei Chang, Aram Galstyan, Richard Zemel, Rahul Gupta. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Junyi Li 0002, Charith Peris, Ninareh Mehrabi, Palash Goyal, Kai-Wei Chang 0001, Aram Galstyan, Richard S. Zemel, Rahul Gupta 0001
NAACL-HLT5
2024 CASA: Causality-driven Argument Sufficiency Assessment
abstract
Xiao Liu, Yansong Feng, Kai-Wei Chang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Xiao Liu 0032, Yansong Feng 0002, Kai-Wei Chang 0001
NAACL-HLT3
2024 Mitigating Bias for Question Answering Models by Tracking Bias Influence
abstract
Mingyu Ma, Jiun-Yu Kao, Arpit Gupta, Yu-Hsiang Lin, Wenbo Zhao, Tagyoung Chung, Wei Wang, Kai-Wei Chang, Nanyun Peng. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Mingyu Derek Ma, Jiun-Yu Kao, Arpit Gupta, Yu-Hsiang Lin, Wenbo Zhao 0006, Tagyoung Chung, Wei Wang 0010, Kai-Wei Chang 0001, Nanyun Peng 0001
NAACL-HLT8
2024 Contextual Label Projection for Cross-Lingual Structured Prediction
abstract
Tanmay Parekh, I-Hung Hsu, Kuan-Hao Huang, Kai-Wei Chang, Nanyun Peng. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Tanmay Parekh, I-Hung Hsu, Kuan-Hao Huang, Kai-Wei Chang 0001, Nanyun Peng 0001
NAACL-HLT4
2024 Event Detection from Social Media for Epidemic Prediction
abstract
Tanmay Parekh, Anh Mac, Jiarui Yu, Yuxuan Dong, Syed Shahriar, Bonnie Liu, Eric Yang, Kuan-Hao Huang, Wei Wang, Nanyun Peng, Kai-Wei Chang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Tanmay Parekh, Anh Mac, Jiarui Yu, Syed Shahriar, Bonnie Liu, Eric Yang, Kuan-Hao Huang, Wei Wang 0010, Nanyun Peng 0001, Kai-Wei Chang 0001
NAACL-HLT11
2024 DACO: Towards Application-Driven and Comprehensive Data Analysis via Code Generation
abstract
Data analysis is a crucial analytical process essential for deriving insights from real-world databases. As shown in Figure 1, the need for data analysis typically arises from specific application scenarios, and requires diverse reasoning skills including mathematical reasoning, logical reasoning, and strategic reasoning. Existing work often focus on simple factual retrieval or arithmetic resolutions and thus are insufficient for addressing complex real-world queries. This work aims to propose new resources and benchmarks on this crucial yet challenging and under-explored task. Due to the prohibitively high cost of collecting expert annotations, we use large language models (LLMs) enhanced by code generation to automatically generate high-quality data analysis, which will later be refined by human annotators. We construct the DACO dataset, containing (1) 440 databases (of tabular data) collected from real-world scenarios, (2) ~2k automatically generated query-answer pairs that can serve as weak supervision for model training, and (3) a concentrated but high-quality test set with human refined annotations that serves as our main evaluation benchmark. Experiments show that while LLMs like GPT-4 exhibit promising data analysis capabilities, they are still evaluated as less helpful than human-written analysis on 58.1% cases. Leveraging our weak supervision data, we experiment with various fine-tuning methods, including supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF). Our trained model outperforms existing baselines for table question answering, and RLHF further boosts the helpfulness of generated analysis on 58.5% cases.Data and code are released at https://github.com/shirley-wu/daco.
Xueqing Wu 0001, Jingzhen Sha, Te-Lin Wu, Hanyu Zhou, Mohan Tang, Kai-Wei Chang 0001, Nanyun Peng 0001, Haoran Huang
NeurIPS7
2024 Matryoshka Query Transformer for Large Vision-Language Models
abstract
Large Vision-Language Models (LVLMs) typically encode an image into a fixed number of visual tokens (e.g., 576) and process these tokens with a language model. Despite their strong performance, LVLMs face challenges in adapting to varying computational constraints. This raises the question: can we achieve flexibility in the number of visual tokens to suit different tasks and computational resources? We answer this with an emphatic yes. Inspired by Matryoshka Representation Learning, we introduce the Matryoshka Query Transformer (MQT), capable of encoding an image into $m$ visual tokens during inference, where $m$ can be any number up to a predefined maximum. This is achieved by employing a query transformer with $M$ latent query tokens to compress the visual embeddings. During each training step, we randomly select $m \leq M$ latent query tokens and train the model using only these first $m$ tokens, discarding the rest. Combining MQT with LLaVA, we train a single model once, and flexibly and drastically reduce the number of inference-time visual tokens while maintaining similar or better performance compared to training independent models for each number of tokens. Our model, MQT-LLaVA, matches LLaVA-1.5 performance across 11 benchmarks using a maximum of 256 tokens instead of LLaVA’s fixed 576. Reducing to 16 tokens (8x less TFLOPs) only sacrifices the performance by 2.4 points on MMBench. On certain tasks such as ScienceQA and MMMU, we can even go down to only 2 visual tokens with performance drops of just 3\% and 6\% each. Our exploration of the trade-off between the accuracy and computational cost brought about by the number of visual tokens facilitates future research to achieve the best of both worlds.
Wenbo Hu 0006, Zi-Yi Dou, Liunian Harold Li, Amita Kamath, Nanyun Peng 0001, Kai-Wei Chang 0001
NeurIPS6
2024 Enhancing Large Vision Language Models with Self-Training on Image Comprehension
abstract
Large vision language models (LVLMs) integrate large language models (LLMs) with pre-trained vision encoders, thereby activating the perception capability of the model to understand image inputs for different queries and conduct subsequent reasoning. Improving this capability requires high-quality vision-language data, which is costly and labor-intensive to acquire. Self-training approaches have been effective in single-modal settings to alleviate the need for labeled data by leveraging model's own generation. However, effective self-training remains a challenge regarding the unique visual perception and reasoning capability of LVLMs. To address this, we introduce **S**elf-**T**raining on **I**mage **C**omprehension (**STIC**), which emphasizes a self-training approach specifically for image comprehension. First, the model self-constructs a preference dataset for image descriptions using unlabeled images. Preferred responses are generated through a step-by-step prompt, while dis-preferred responses are generated from either corrupted images or misleading prompts. To further self-improve reasoning on the extracted visual information, we let the model reuse a small portion of existing instruction-tuning data and append its self-generated image descriptions to the prompts. We validate the effectiveness of STIC across seven different benchmarks, demonstrating substantial performance gains of 4.0% on average while using 70% less supervised fine-tuning data than the current method. Further studies dive into various components of STIC and highlight its potential to leverage vast quantities of unlabeled images for self-training.
Yihe Deng, Pan Lu, Fan Yin, Ziniu Hu, Sheng Shen 0001, Quanquan Gu, James Zou 0001, Kai-Wei Chang 0001, Wei Wang 0010
NeurIPS8
2024 JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images
abstract
Existing vision-language understanding benchmarks largely consist of images of objects in their usual contexts.As a consequence, recent multimodal large language models can perform well with only a shallow visual understanding by relying on background language biases. Thus, strong performance on these benchmarks does not necessarily correlate with strong visual understanding. In this paper, we release JourneyBench, a comprehensive human-annotated benchmark of generated images designed to assess the model's fine-grained multimodal reasoning abilities across five tasks: complementary multimodal chain of thought, multi-image VQA, imaginary image captioning, VQA with hallucination triggers, and fine-grained retrieval with sample-specific distractors.Unlike existing benchmarks, JourneyBench explicitly requires fine-grained multimodal reasoning in unusual imaginary scenarios where language bias and holistic image gist are insufficient. We benchmark state-of-the-art models on JourneyBench and analyze performance along a number of fine-grained dimensions. Results across all five tasks show that JourneyBench is exceptionally challenging for even the best models, indicating that models' visual reasoning abilities are not as strong as they first appear. We discuss the implications of our findings and propose avenues for further research.
Zhecan Wang, Junzhang Liu, Chia-Wei Tang, Hani Alomari, Anushka Sivakumar, Rui Sun 0011, Md. Atabuzzaman, Hammad A. Ayyubi, Haoxuan You, Alvi Md. Ishmam, Kai-Wei Chang 0001, Shih-Fu Chang, Christopher Thomas 0004
NeurIPS12
2024 SafeWorld: Geo-Diverse Safety Alignment
abstract
In the rapidly evolving field of Large Language Models (LLMs), ensuring safety is a crucial and widely discussed topic. However, existing works often overlooks the geo-diversity of cultural and legal standards across the world. To reveal the chal5 lenges posed by geo-diverse safety standards, we introduce SafeWorld, a novel benchmark specifically designed to evaluate LLMs’ ability to generate responses that are not only helpful but also culturally sensitive and legally compliant across diverse global contexts. SafeWorld encompasses 2,775 test user queries, each grounded in high-quality, human-verified cultural norms and legal policies from 50 countries and 493 regions/races. On top of it, we propose a multi-dimensional automatic safety evaluation framework that assesses the contextual appropriateness, accuracy, and comprehensiveness of responses. Our evaluations reveal that current LLMs struggle to meet these criteria effectively. To enhance LLMs’ alignment with geo-diverse safety standards, we synthesize helpful preference pairs for Direct Preference Optimization (DPO) alignment. The preference pair construction aims to encourage LLMs to behave appropriately and provide precise references to relevant cultural norms and policies when necessary. Our trained SafeWorldLM outperforms all competing models, including GPT-4o on all the three evaluation dimensions by a large margin. Global human evaluators also note a nearly 20% higher winning rate in helpfulness and harmfulness evaluation.
Da Yin, Haoyi Qiu, Kung-Hsiang Huang, Kai-Wei Chang 0001, Nanyun Peng 0001
NeurIPS4
2024 Open-Emotion: A Reproducible EMO-Superb For Speech Emotion Recognition Systems
abstract
Speech emotion recognition (SER) is an essential technology for human-computer interaction systems. However, the previous study reveals that 80.77% of SER papers yield results that cannot be reproduced on the well-known IEMOCAP dataset. The main reason for reproducibility challenges is that the database did not provide standard data splits (e.g., train, development, and test sets). Prior papers could define its partition, but they did not provide details of the partition or source code for processing the partition. Therefore, this work aims to make SER open and reproducible to everyone. We develop the EMO-SUPERB, shorted for EMOtion Speech Universal PERformance Benchmark, including a user-friendly codebase to leverage 16 state-of-the-art (SOTA) speech self-supervised learning models for exhaustive evaluation plus one SOTA SER model across 6 open-source SER datasets in English and Chinese. We make all resources open-source to facilitate future developments in SER. Researchers can easily upload their systems or datasets to EMO-SUPERB, and we name the project “Open-Emotion”.
Huang-Cheng Chou, Kai-Wei Chang 0001, Lucas Goncalves, Jiawei Du 0003, Jyh-Shing Roger Jang, Chi-Chun Lee, Hung-yi Lee
SLT3
2024 Codec-Superb @ SLT 2024: A Lightweight Benchmark For Neural Audio Codec Models
abstract
Neural audio codec models are becoming increasingly important as they serve as tokenizers for audio, enabling efficient transmission or facilitating speech language modeling. The ideal neural audio codec should maintain content, paralinguistics, speaker characteristics, and audio information even at low bitrates. Recently, numerous advanced neural codec models have been proposed. However, codec models are often tested under varying experimental conditions. As a result, we introduce the Codec-SUPERB challenge at SLT 20241, designed to facilitate fair and lightweight comparisons among existing codec models and inspire advancements in the field. This challenge brings together representative speech applications and objective metrics, and carefully selects license-free datasets, sampling them into small sets to reduce evaluation computation costs. This paper presents the challenge’s rules, datasets, participant systems, results, and findings.1https://codecsuperb.github.io/
Xuanjun Chen, Yi-Cheng Lin, Kai-Wei Chang 0001, Jiawei Du 0003, Ke-Han Lu, Alexander H. Liu, Ho-Lam Chung, Yuan-Kuei Wu, Dongchao Yang, Songxiang Liu, Yi-Chiao Wu, Xu Tan 0003, James R. Glass, Shinji Watanabe 0001, Hung-yi Lee
SLT4
2024 Tree-managed network ensembles for video prediction
Everett Fall, Kai-Wei Chang 0001, Liang-Gee Chen
Mach. Vis. Appl.2
2024 Red Teaming Language Model Detectors with Language Models
abstract
Abstract The prevalence and strong capability of large language models (LLMs) present significant safety and ethical risks if exploited by malicious users. To prevent the potentially deceptive usage of LLMs, recent work has proposed algorithms to detect LLM-generated text and protect LLMs. In this paper, we investigate the robustness and reliability of these LLM detectors under adversarial attacks. We study two types of attack strategies: 1) replacing certain words in an LLM’s output with their synonyms given the context; 2) automatically searching for an instructional prompt to alter the writing style of the generation. In both strategies, we leverage an auxiliary LLM to generate the word replacements or the instructional prompt. Different from previous works, we consider a challenging setting where the auxiliary LLM can also be protected by a detector. Experiments reveal that our attacks effectively compromise the performance of all detectors in the study with plausible generations, underscoring the urgent need to improve the robustness of LLM-generated text detection systems. Code is available at https://github.com/shizhouxing/LLM-Detector-Robustness.
Zhouxing Shi, Fan Yin, Xiangning Chen, Kai-Wei Chang 0001, Cho-Jui Hsieh
Trans. Assoc. Comput. Linguistics5
2023 TAGPRIME: A Unified Framework for Relational Structure Extraction
abstract
I-Hung Hsu, Kuan-Hao Huang, Shuning Zhang, Wenxin Cheng, Prem Natarajan, Kai-Wei Chang, Nanyun Peng. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
I-Hung Hsu, Kuan-Hao Huang, Wenxin Cheng, Premkumar Natarajan, Kai-Wei Chang 0001, Nanyun Peng 0001
ACL (1)6
2023 ParaAMR: A Large-Scale Syntactically Diverse Paraphrase Dataset by AMR Back-Translation
abstract
Kuan-Hao Huang, Varun Iyer, I-Hung Hsu, Anoop Kumar, Kai-Wei Chang, Aram Galstyan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Kuan-Hao Huang, Varun Iyer, I-Hung Hsu, Kai-Wei Chang 0001, Aram Galstyan
ACL (1)5
2023 Symbolic Chain-of-Thought Distillation: Small Models Can Also "Think" Step-by-Step
abstract
Liunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren, Kai-Wei Chang, Yejin Choi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Liunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren 0001, Kai-Wei Chang 0001, Yejin Choi 0001
ACL (1)5
2023 A Survey of Deep Learning for Mathematical Reasoning
abstract
Mathematical reasoning is a fundamental aspect of human intelligence and is applicable in various fields, including science, engineering, finance, and everyday life.The development of artificial intelligence (AI) systems capable of solving math problems and proving theorems in language has garnered significant interest in the fields of machine learning and natural language processing.For example, mathematics serves as a testbed for aspects of reasoning that are challenging for powerful deep learning models, driving new algorithmic and modeling advances.On the other hand, recent advances in large-scale neural language models have opened up new benchmarks and opportunities to use deep learning for mathematical reasoning.In this survey paper, we review the key tasks, datasets, and methods at the intersection of mathematical reasoning and deep learning over the past decade.We also evaluate existing benchmarks and methods, and discuss future research directions in this domain.Question: Bod has 2 apples and David has 5 apples.How many apples do they have in total?Rationale: x = 2 + 5 Solution: 7
Pan Lu, Liang Qiu 0001, Wenhao Yu 0002, Sean Welleck, Kai-Wei Chang 0001
ACL (1)5
2023 Resolving Ambiguities in Text-to-Image Generative Models
abstract
Ninareh Mehrabi, Palash Goyal, Apurv Verma, Jwala Dhamala, Varun Kumar, Qian Hu, Kai-Wei Chang, Richard Zemel, Aram Galstyan, Rahul Gupta. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Ninareh Mehrabi, Palash Goyal, Apurv Verma, Jwala Dhamala, Kai-Wei Chang 0001, Richard S. Zemel, Aram Galstyan, Rahul Gupta 0001
ACL (1)7
2023 GENEVA: Benchmarking Generalizability for Event Argument Extraction with Hundreds of Event Types and Argument Roles
abstract
Recent works in Event Argument Extraction (EAE) have focused on improving model generalizability to cater to new events and domains.However, standard benchmarking datasets like ACE and ERE cover less than 40 event types and 25 entity-centric argument roles.Limited diversity and coverage hinder these datasets from adequately evaluating the generalizability of EAE models.In this paper, we first contribute by creating a large and diverse EAE ontology.This ontology is created by transforming FrameNet, a comprehensive semantic role labeling (SRL) dataset for EAE, by exploiting the similarity between these two tasks.Then, exhaustive human expert annotations are collected to build the ontology, concluding with 115 events and 220 argument roles, with a significant portion of roles not being entities.We utilize this ontology to further introduce GENEVA, a diverse generalizability benchmarking dataset comprising four test suites, aimed at evaluating models' ability to handle limited data and unseen event type generalization.We benchmark six EAE models from various families.The results show that owing to non-entity argument roles, even the best-performing model can only achieve 39% F1 score, indicating how GENEVA provides new challenges for generalization in EAE.Overall, our large and diverse EAE ontology can aid in creating more comprehensive future resources, while GENEVA is a challenging benchmarking dataset encouraging further research for improving generalizability in EAE.
Tanmay Parekh, I-Hung Hsu, Kuan-Hao Huang, Kai-Wei Chang 0001, Nanyun Peng 0001
ACL (1)4
2023 Efficient Shapley Values Estimation by Amortization for Text Classification
abstract
Despite the popularity of Shapley Values in explaining neural text classification models, computing them is prohibitive for large pretrained models due to a large number of model evaluations.In practice, Shapley Values are often estimated with a small number of stochastic model evaluations.However, we show that the estimated Shapley Values are sensitive to random seed choices -the top-ranked features often have little overlap across different seeds, especially on examples with longer input texts.This can only be mitigated by aggregating thousands of model evaluations, which on the other hand, induces substantial computational overheads.To mitigate the trade-off between stability and efficiency, we develop an amortized model that directly predicts each input feature's Shapley Value without additional model evaluations.It is trained on a set of examples whose Shapley Values are estimated from a large number of model evaluations to ensure stability.Experimental results on two text classification datasets demonstrate that our amortized model estimates Shapley Values accurately with up to 60 times speedup compared to traditional methods.Furthermore, the estimated values are stable as the inference is deterministic.We release our code at https://github.com/yangalan123/ Amortized-Interpretability.
Chenghao Yang 0001, Fan Yin, He He 0001, Kai-Wei Chang 0001, Xiaofei Ma 0001, Bing Xiang
ACL (1)4
2023 Factoring the Matrix of Domination: A Critical Review and Reimagination of Intersectionality in AI Fairness
abstract
Intersectionality is a critical framework that, through inquiry and praxis, allows us to examine how social inequalities persist through domains of structure and discipline. Given AI fairness’ raison d’être of “fairness,” we argue that adopting intersectionality as an analytical framework is pivotal to effectively operationalizing fairness. Through a critical review of how intersectionality is discussed in 30 papers from the AI fairness literature, we deductively and inductively: 1) map how intersectionality tenets operate within the AI fairness paradigm and 2) uncover gaps between the conceptualization and operationalization of intersectionality. We find that researchers overwhelmingly reduce intersectionality to optimizing for fairness metrics over demographic subgroups. They also fail to discuss their social context and when mentioning power, they mostly situate it only within the AI pipeline. We: 3) outline and assess the implications of these gaps for critical inquiry and praxis, and 4) provide actionable recommendations for AI fairness researchers to engage with intersectionality in their work by grounding it in AI epistemology.
Anaelia Ovalle, Arjun Subramonian, Vagrant Gautam, Gilbert Gee, Kai-Wei Chang 0001
AIES5
2023 Semantic Strengthening of Neuro-Symbolic Learning
abstract
Numerous neuro-symbolic approaches have recently been proposed typically with the goal of adding symbolic knowledge to the output layer of a neural network. Ideally, such losses maximize the probability that the neural network’s predictions satisfy the underlying domain. Unfortunately, this type of probabilistic inference is often computationally infeasible. Neuro-symbolic approaches therefore commonly resort to fuzzy approximations of this probabilistic objective, sacrificing sound probabilistic semantics, or to sampling which is very seldom feasible. We approach the problem by first assuming the constraint decomposes conditioned on the features learned by the network. We iteratively strengthen our approximation, restoring the dependence between the constraints most responsible for degrading the quality of the approximation. This corresponds to computing the mutual information between pairs of constraints conditioned on the network’s learned features, and may be construed as a measure of how well aligned the gradients of two distributions are. We show how to compute this efficiently for tractable circuits. We test our approach on three tasks: predicting a minimum-cost path in Warcraft, predicting a minimum-cost perfect matching, and solving Sudoku puzzles, observing that it improves upon the baselines while sidestepping intractability.
Kareem Ahmed, Kai-Wei Chang 0001, Guy Van den Broeck
AISTATS2
2023 Prompting and Adapter Tuning For Self-Supervised Encoder-Decoder Speech Model
abstract
Prompting and adapter tuning have emerged as efficient alternatives to fine-tuning (FT) methods. However, existing studies on speech prompting focused on classification tasks and failed on more complex sequence generation tasks. Besides, adapter tuning is primarily applied with a focus on encoder-only self-supervised models. Our experiments show that prompting on Wav2Seq, a self-supervised encoder-decoder model, surpasses previous works in sequence generation tasks. It achieves a remarkable 53% relative improvement in word error rate for ASR and a 27% in F1 score for slot filling. Additionally, prompting competes with the FT method in the low-resource scenario. Moreover, we show the transferability of prompting and adapter tuning on Wav2Seq in cross-lingual ASR. When limited trainable parameters are involved, prompting and adapter tuning consistently outperform conventional FT across 7 languages. Notably, in the low-resource scenario, prompting consistently outperforms adapter tuning.
Kai-Wei Chang 0001, Ming-Hsin Chen, Yun-Ping Lin, Jing Neng Hsu, Paul Kuo-Ming Huang, Chien-Yu Huang, Shang-Wen Li 0001, Hung-yi Lee
ASRU1
2023 Towards General-Purpose Text-Instruction-Guided Voice Conversion
abstract
This paper introduces a novel voice conversion (VC) model, guided by text instructions such as “articulate slowly with a deep tone“ or “speak in a cheerful boyish voice”. Unlike traditional methods that rely on reference utterances to determine the attributes of the converted speech, using text instruction adds versatility and specificity to voice conversion. The proposed VC model is a neural codec language model which processes a sequence of discrete codes, resulting in the code sequence of converted speech. It utilizes text instructions as style prompts to modify the prosody and emotional information of the given speech. In contrast to previous approaches, which often rely on employing separate encoders like prosody and content encoders to handle different aspects of the source speech, our model handles various information of speech in an end-to-end manner. Experiments have demonstrated the impressive capabilities of our model in comprehending instructions and delivering reasonable results1.
Chun-Yi Kuan, Chen-An Li, Tsu-Yuan Hsu, Tse-Yang Lin, Ho-Lam Chung, Kai-Wei Chang 0001, Shuo-Yiin Chang, Hung-yi Lee
ASRU6
2023 Minisuperb: Lightweight Benchmark for Self-Supervised Speech Models
abstract
SUPERB was proposed to evaluate the generalizability of self-supervised learning (SSL) speech models across various tasks. However, it incurs high computational costs due to the large datasets and diverse tasks. In this paper, we introduce MiniSUPERB, a lightweight benchmark that efficiently evaluates SSL speech models with comparable results to SUPERB but lower computational costs significantly. We carefully select representative tasks, sample datasets, and extract model representations offline. Our approach achieves a Spearman’s rank correlation of 0.954 and 0.982 with SUPERB Paper and SUPERB Challenge, respectively. Additionally, we reduce the computational cost by 97 % in terms of Multiply-ACcumulate operations (MACs). Furthermore, we evaluate SSL speech models in few-shot scenarios and observe significant variations in their performance. To our knowledge, this is the first study to examine both the computational cost of the model itself and the cost of evaluating it on a benchmark.11Our code is available at https://github.com/Comet0322/MiniSUPERB
Yu-Hsiang Wang, Huang-Yu Chen, Kai-Wei Chang 0001, Winston H. Hsu, Hung-yi Lee
ASRU3
2023 Reveal: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge Memory
abstract
In this paper, we propose an end-to-end Retrieval-Augmented Visual Language Model (REVEAL) that learns to encode world knowledge into a large-scale memory, and to retrieve from it to answer knowledge-intensive queries. Reveal consists of four key components: the memory, the encoder, the retriever and the generator. The large-scale memory encodes various sources of multimodal world knowledge (e.g. image-text pairs, question answering pairs, knowledge graph triplets, etc.) via a unified encoder. The retriever finds the most relevant knowledge entries in the memory, and the generator fuses the retrieved knowledge with the input query to produce the output. A key novelty in our approach is that the memory, encoder, retriever and generator are all pre-trained end-to-end on a massive amount of data. Furthermore, our approach can use a diverse set of multimodal knowledge sources, which is shown to result in significant gains. We show that Reveal achieves state-of-the-art results on visual question answering and image captioning. The project page of this work is reveal. github. io.
Ziniu Hu, Ahmet Iscen, Chen Sun 0002, Kai-Wei Chang 0001, Yizhou Sun, Cordelia Schmid, David A. Ross, Alireza Fathi
CVPR5
2023 GIVL: Improving Geographical Inclusivity of Vision-Language Models with Pre-Training Methods
abstract
A key goal for the advancement of AI is to develop technologies that serve the needs not just of one group but of all communities regardless of their geographical re-gion. In fact, a significant proportion of knowledge is locally shared by people from certain regions but may not apply equally in other regions because of cultural dif-ferences. If a model is unaware of regional character-istics, it may lead to performance disparity across re-gions and result in bias against underrepresented groups. We propose GIVL, a Geographically Inclusive Vision-and-Language Pre-trained model. There are two attributes of geo-diverse visual concepts which can help to learn geo-diverse knowledge: 1) concepts under similar categories have unique knowledge and visual characteristics, 2) concepts with similar visual features may fall in completely different categories. Motivated by the attributes, we de-sign new pre-training objectives Image-Knowledge Matching (IKM) and Image Edit Checking (IEC) to pre-train GIVL. Compared with similar-size models pre-trained with similar scale of data, GIVL achieves state-of-the-art (SOTA) and more balanced performance on geo-diverse V &L tasks. Code and data are released at https://github.com/WadeYin9712/GIVL.
Da Yin, Govind Thattai, Michael Johnston, Kai-Wei Chang 0001
CVPR5
2023 Summarize and Generate to Back-translate: Unsupervised Translation of Programming Languages
abstract
Back-translation is widely known for its effectiveness in neural machine translation when there is little to no parallel data.In this approach, a source-to-target model is coupled with a target-to-source model trained in parallel.The target-to-source model generates noisy sources, while the source-to-target model is trained to reconstruct the targets and vice versa.Recent developments of multilingual pre-trained sequence-to-sequence models for programming languages have been very effective for a broad spectrum of downstream software engineering tasks.Hence, training them to build programming language translation systems via back-translation is compelling.However, these models cannot be further trained via back-translation since they learn to output sequences in the same language as the inputs during pre-training.As an alternative, we propose performing back-translation via code summarization and generation.In code summarization, a model learns to generate natural language (NL) summaries given code snippets.In code generation, the model learns to do the opposite.Therefore, target-to-source generation in back-translation can be viewed as a target-to-NL-to-source generation.We show that our proposed approach performs competitively with state-of-the-art methods.We have made the code publicly available.1
Wasi Uddin Ahmad, Saikat Chakraborty 0001, Baishakhi Ray, Kai-Wei Chang 0001
EACL4
2023 Retrieval Enhanced Data Augmentation for Question Answering on Privacy Policies
abstract
Prior studies in privacy policies frame the question answering (QA) task as identifying the most relevant text segment or a list of sentences from a policy document given a user query.Existing labeled datasets are heavily imbalanced (only a few relevant segments), limiting the QA performance in this domain.In this paper, we develop a data augmentation framework based on ensembling retriever models that captures the relevant text segments from unlabeled policy documents and expand the positive examples in the training set.In addition, to improve the diversity and quality of the augmented data, we leverage multiple pre-trained language models (LMs) and cascade them with noise reduction filter models.Using our augmented data on the PrivacyQA benchmark, we elevate the existing baseline by a large margin (10% F1) and achieve a new state-of-the-art F1 score of 50%.Our ablation studies provide further insights into the effectiveness of our approach.
Md. Rizwan Parvez, Jianfeng Chi, Wasi Uddin Ahmad, Yuan Tian 0001, Kai-Wei Chang 0001
EACL5
2023 Text encoders bottleneck compositionality in contrastive vision-language models
abstract
Performant vision-language (VL) models like CLIP represent captions using a single vector.How much information about language is lost in this bottleneck?We first curate CompPrompts, a set of increasingly compositional image captions that VL models should be able to capture (e.g., single object, to ob-ject+property, to multiple interacting objects).Then, we train text-only recovery probes that aim to reconstruct captions from single-vector text representations produced by several VL models.This approach does not require images, allowing us to test on a broader range of scenes compared to prior work.We find that: 1) CLIP's text encoder falls short on more compositional inputs, including object relationships, attribute-object association, counting, and negations; 2) some text encoders work significantly better than others; and 3) text-only recovery performance predicts multimodal matching performance on ControlledImCaps: a new evaluation benchmark we collect and release consisting of fine-grained compositional images and captions.Specifically, our results suggest textonly recoverability is a necessary (but not sufficient) condition for modeling compositional factors in contrastive VL models.We release our datasets and code.
Amita Kamath, Jack Hessel, Kai-Wei Chang 0001
EMNLP3
2023 What's "up" with vision-language models? Investigating their struggle with spatial reasoning
abstract
Recent vision-language (VL) models are powerful, but can they reliably distinguish "right" from "left"?We curate three new corpora to quantify model comprehension of such basic spatial relations.These tests isolate spatial reasoning more precisely than existing datasets like VQAv2, e.g., our What'sUp benchmark contains sets of photographs varying only the spatial relations of objects, keeping their identity fixed (see Figure 1: models must comprehend not only the usual case of a dog under a table, but also, the same dog on top of the same table).We evaluate 18 VL models, finding that all perform poorly, e.g., BLIP finetuned on VQAv2, which nears human parity on VQAv2, achieves 56% accuracy on our benchmarks vs. humans at 99%.We conclude by studying causes of this surprising behavior, finding: 1) that popular vision-language pretraining corpora like LAION-2B contain little reliable data for learning spatial relationships; and 2) that basic modeling interventions like up-weighting preposition-containing instances or fine-tuning on our corpora are not sufficient to address the challenges our benchmarks pose.We are hopeful that these corpora will facilitate further research, and we release our data and code at https://github.com/amitakamath/ whatsup_vlms.
Amita Kamath, Jack Hessel, Kai-Wei Chang 0001
EMNLP3
2023 Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks
abstract
Instruction tuning (IT) achieves impressive zero-shot generalization results by training large language models (LLMs) on a massive amount of diverse tasks with instructions.However, how to select new tasks to improve the performance and generalizability of IT models remains an open question.Training on all existing tasks is impractical due to prohibiting computation requirements, and randomly selecting tasks can lead to suboptimal performance.In this work, we propose active instruction tuning based on prompt uncertainty, a novel framework to identify informative tasks, and then actively tune the models on the selected tasks.We represent the informativeness of new tasks with the disagreement of the current model outputs over perturbed prompts.Our experiments on NIV2 and Self-Instruct datasets demonstrate that our method consistently outperforms other baseline strategies for task selection, achieving better out-of-distribution generalization with fewer training tasks.Additionally, we introduce a task map that categorizes and diagnoses tasks based on prompt uncertainty and prediction probability.We discover that training on ambiguous (prompt-uncertain) tasks improves generalization while training on difficult (prompt-certain and low-probability) tasks offers no benefit, underscoring the importance of task selection for instruction tuning. 1
Po-Nien Kung, Fan Yin, Di Wu 0054, Kai-Wei Chang 0001, Nanyun Peng 0001
EMNLP4
2023 Rethinking Model Selection and Decoding for Keyphrase Generation with Pre-trained Sequence-to-Sequence Models
abstract
Keyphrase Generation (KPG) is a longstanding task in NLP with widespread applications.The advent of sequence-to-sequence (seq2seq) pre-trained language models (PLMs) has ushered in a transformative era for KPG, yielding promising performance improvements.However, many design decisions remain unexplored and are often made arbitrarily.This paper undertakes a systematic analysis of the influence of model selection and decoding strategies on PLM-based KPG.We begin by elucidating why seq2seq PLMs are apt for KPG, anchored by an attention-driven hypothesis.We then establish that conventional wisdom for selecting seq2seq PLMs lacks depth: (1) merely increasing model size or performing task-specific adaptation is not parameter-efficient; (2) although combining in-domain pre-training with task adaptation benefits KPG, it does partially hinder generalization.Regarding decoding, we demonstrate that while greedy search achieves strong F1 scores, it lags in recall compared with samplingbased methods.Based on these insights, we propose DESEL, a likelihood-based decodeselect algorithm for seq2seq PLMs.DESEL improves greedy search by an average of 4.7% semantic F1 across five datasets.Our collective findings pave the way for deeper future investigations into PLM-based KPG.
Di Wu 0054, Wasi Uddin Ahmad, Kai-Wei Chang 0001
EMNLP3
2023 LACMA: Language-Aligning Contrastive Learning with Meta-Actions for Embodied Instruction Following
abstract
End-to-end Transformers have demonstrated an impressive success rate for Embodied Instruction Following when the environment has been seen in training.However, they tend to struggle when deployed in an unseen environment.This lack of generalizability is due to the agent's insensitivity to subtle changes in natural language instructions.To mitigate this issue, we propose explicitly aligning the agent's hidden states with the instructions via contrastive learning.Nevertheless, the semantic gap between high-level language instructions and the agent's low-level action space remains an obstacle.Therefore, we further introduce a novel concept of meta-actions to bridge the gap.Meta-actions are ubiquitous action patterns that can be parsed from the original action sequence.These patterns represent higher-level semantics that are intuitively aligned closer to the instructions.When meta-actions are applied as additional training signals, the agent generalizes better to unseen environments.Compared to a strong multi-modal Transformer baseline, we achieve a significant 4.5% absolute gain in success rate in unseen environments of ALFRED Embodied Instruction Following.Additional analysis shows that the contrastive objective and meta-actions are complementary in achieving the best results, and the resulting agent better aligns its states with corresponding instructions, making it more suitable for real-world embodied agents. 1
Cheng-Fu Yang, Yen-Chun Chen 0001, Xiyang Dai, Lu Yuan 0001, Yu-Chiang Frank Wang, Kai-Wei Chang 0001
EMNLP7
2023 Dynosaur: A Dynamic Growth Paradigm for Instruction-Tuning Data Curation
abstract
Instruction tuning has emerged to enhance the capabilities of large language models (LLMs) to comprehend instructions and generate appropriate responses.Existing methods either manually annotate or employ LLM (e.g., GPTseries) to generate data for instruction tuning.However, they often overlook associating instructions with existing annotated datasets.In this paper, we propose DYNOSAUR, a dynamic growth paradigm for the automatic curation of instruction-tuning data.Based on the metadata of existing datasets, we use LLMs to automatically construct instruction-tuning data by identifying relevant data fields and generating appropriate instructions.By leveraging the existing annotated datasets, DYNOSAUR offers several advantages: 1) it reduces the API cost for generating instructions (e.g., it costs less than $12 USD by calling GPT-3.5-turbo for generating 800K instruction tuning samples; 2) it provides high-quality data for instruction tuning (e.g., it performs better than ALPACA and FLAN on SUPER-NI and LONGFORM with comparable data sizes); and 3) it supports the continuous improvement of models by generating instruction-tuning data when a new annotated dataset becomes available.We further investigate a continual learning scheme for learning with the ever-growing instruction-tuning dataset, and demonstrate that replaying tasks with diverse instruction embeddings not only helps mitigate forgetting issues but generalizes to unseen tasks better.
Da Yin, Xiao Liu 0032, Fan Yin, Ming Zhong 0005, Hritik Bansal, Jiawei Han 0001, Kai-Wei Chang 0001
EMNLP7
2023 Ensemble Knowledge Distillation of Self-Supervised Speech Models
abstract
Distilled self-supervised models have shown competitive performance and efficiency in recent years. However, there is a lack of experience in jointly distilling multiple self-supervised speech models. In our work, we performed Ensemble Knowledge Distillation (EKD) on various self-supervised speech models such as HuBERT, RobustHuBERT, and WavLM. We tried two different aggregation techniques, layerwise-average and layerwise-concatenation, to the representations of different teacher models and found that the former was more effective. On top of that, we proposed a multiple prediction head method for student models to predict different layer outputs of multiple teacher models simultaneously. The experimental results show that our method improves the performance of the distilled models on four downstream speech processing tasks, Phoneme Recognition, Speaker Identification, Emotion Recognition, and Automatic Speech Recognition in the hidden-set track of the SUPERB benchmark.
Kuan-Po Huang, Tzu-hsun Feng, Yu-Kuan Fu, Tsu-Yuan Hsu, Po-Chieh Yen, Wei-Cheng Tseng, Kai-Wei Chang 0001, Hung-yi Lee
ICASSP7
2023 CleanCLIP: Mitigating Data Poisoning Attacks in Multimodal Contrastive Learning
abstract
Multimodal contrastive pretraining has been used to train multimodal representation models, such as CLIP, on large amounts of paired image-text data. However, previous studies have revealed that such models are vulnerable to backdoor attacks. Specifically, when trained on backdoored examples, CLIP learns spurious correlations between the embedded backdoor trigger and the target label, aligning their representations in the joint embedding space. Injecting even a small number of poisoned examples, such as 75 examples in 3 million pretraining data, can significantly manipulate the model’s behavior, making it difficult to detect or unlearn such correlations. To address this issue, we propose CleanCLIP, a finetuning framework that weakens the learned spurious associations introduced by backdoor attacks by independently re-aligning the representations for individual modalities. We demonstrate that unsupervised finetuning using a combination of multimodal contrastive and unimodal self-supervised objectives for individual modalities can significantly reduce the impact of the backdoor attack. Additionally, we show that supervised finetuning on task-specific labeled image data removes the backdoor trigger from the CLIP vision encoder. We show empirically that CleanCLIP maintains model performance on benign examples while erasing a range of backdoor attacks on multimodal contrastive learning. Code and pretrained checkpoints are available at https://github.com/nishadsinghi/CleanCLIP.
Hritik Bansal, Fan Yin, Nishad Singhi, Aditya Grover, Kai-Wei Chang 0001
ICCV6
2023 Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning
Pan Lu, Liang Qiu 0001, Kai-Wei Chang 0001, Ying Nian Wu, Song-Chun Zhu, Tanmay Rajpurohit, Peter Clark, Ashwin Kalyan
ICLR3
2023 On the Paradox of Learning to Reason from Data
abstract
Logical reasoning is needed in a wide range of NLP tasks. Can a BERT model be trained end-to-end to solve logical reasoning problems presented in natural language? We attempt to answer this question in a confined problem space where there exists a set of parameters that perfectly simulates logical reasoning. We make observations that seem to contradict each other: BERT attains near-perfect accuracy on in-distribution test examples while failing to generalize to other data distributions over the exact same problem space. Our study provides an explanation for this paradox: instead of learning to emulate the correct reasoning function, BERT has, in fact, learned statistical features that inherently exist in logical reasoning problems. We also show that it is infeasible to jointly remove statistical features from data, illustrating the difficulty of learning to reason in general. Our result naturally extends to other neural models (e.g. T5) and unveils the fundamental difference between learning to reason and learning to achieve high performance on NLP benchmarks using statistical features.
Honghua Zhang, Liunian Harold Li, Kai-Wei Chang 0001, Guy Van den Broeck
IJCAI4
2023 ABC-KD: Attention-Based-Compression Knowledge Distillation for Deep Learning-Based Noise Suppression
abstract
Noise suppression (NS) models have been widely applied to enhance speech quality.Recently, Deep Learning-Based NS, which we denote as Deep Noise Suppression (DNS), became the mainstream NS method due to its excelling performance over traditional ones.However, DNS models face 2 major challenges for supporting the real-world applications.First, highperforming DNS models are usually large in size, causing deployment difficulties.Second, DNS models require extensive training data, including noisy audios as inputs and clean audios as labels.It is often difficult to obtain clean labels for training DNS models.We propose the use of knowledge distillation (KD) to resolve both challenges.Our study serves 2 main purposes.To begin with, we are among the first to comprehensively investigate mainstream KD techniques on DNS models to resolve the two challenges.Furthermore, we propose a novel Attention-Based-Compression KD method that outperforms all investigated mainstream KD frameworks on DNS task.
Yixin Wan, Xiulian Peng, Kai-Wei Chang 0001, Yan Lu 0001
INTERSPEECH4
2023 A Pseudo-Semantic Loss for Autoregressive Models with Logical Constraints
abstract
Neuro-symbolic AI bridges the gap between purely symbolic and neural approaches to learning. This often requires maximizing the likelihood of a symbolic constraint w.r.t the neural network's output distribution. Such output distributions are typically assumed to be fully-factorized. This limits the applicability of neuro-symbolic learning to the more expressive auto-regressive distributions, e.g., transformers. Under such distributions, computing the likelihood of even simple constraints is #P-hard. Instead of attempting to enforce the constraint on the entire likelihood distribution, we propose to do so on a random, local approximation thereof. More precisely, we approximate the likelihood of the constraint with the pseudolikelihood of the constraint centered around a model sample. Our approach is factorizable, allowing us to reuse solutions to sub-problems---a main tenet for the efficient computation of neuro-symbolic losses. It also provides a local, high fidelity approximation of the likelihood: it exhibits low entropy and KL-divergence around the model sample. We tested our approach on Sudoku and shortest-path prediction cast as auto-regressive generation, and observe that we greatly improve upon the base model's ability to predict logically-consistent outputs. We also tested our approach on the task of detoxifying large language models. We observe that using a simple constraint disallowing a list of toxic words, we are able to steer the model's outputs away from toxic generations, achieving SoTA compared to previous approaches.
Kareem Ahmed, Kai-Wei Chang 0001, Guy Van den Broeck
NeurIPS2
2023 AVIS: Autonomous Visual Information Seeking with Large Language Model Agent
abstract
In this paper, we propose an autonomous information seeking visual question answering framework, AVIS. Our method leverages a Large Language Model (LLM) to dynamically strategize the utilization of external tools and to investigate their outputs via tree search, thereby acquiring the indispensable knowledge needed to provide answers to the posed questions. Responding to visual questions that necessitate external knowledge, such as "What event is commemorated by the building depicted in this image?", is a complex task. This task presents a combinatorial search space that demands a sequence of actions, including invoking APIs, analyzing their responses, and making informed decisions. We conduct a user study to collect a variety of instances of human decision-making when faced with this task. This data is then used to design a system comprised of three components: an LLM-powered planner that dynamically determines which tool to use next, an LLM-powered reasoner that analyzes and extracts key information from the tool outputs, and a working memory component that retains the acquired information throughout the process. The collected user behavior serves as a guide for our system in two key ways. First, we create a transition graph by analyzing the sequence of decisions made by users. This graph delineates distinct states and confines the set of actions available at each state. Second, we use examples of user decision-making to provide our LLM-powered planner and reasoner with relevant contextual instances, enhancing their capacity to make informed decisions. We show that AVIS achieves state-of-the-art results on knowledge-based visual question answering benchmarks such as Infoseek and OK-VQA.
Ziniu Hu, Ahmet Iscen, Chen Sun 0002, Kai-Wei Chang 0001, Yizhou Sun, David A. Ross, Cordelia Schmid, Alireza Fathi
NeurIPS4
2023 DesCo: Learning Object Recognition with Rich Language Descriptions
abstract
Recent development in vision-language approaches has instigated a paradigm shift in learning visual recognition models from language supervision. These approaches align objects with language queries (e.g. "a photo of a cat") and thus improve the models' adaptability to novel objects and domains. Recent studies have attempted to query these models with complex language expressions that include specifications of fine-grained details, such as colors, shapes, and relations. However, simply incorporating language descriptions into queries does not guarantee accurate interpretation by the models. In fact, our experiments show that GLIP, a state-of-the-art vision-language model for object detection, often disregards contextual information in the language descriptions and instead relies heavily on detecting objects solely by their names. To tackle the challenge, we propose a new description-conditioned (DesCo) paradigm of learning object recognition models with rich language descriptions consisting of two innovations: 1) we employ a large language model as a commonsense knowledge engine to generate rich language descriptions of objects; 2) we design context-sensitive queries to improve the model's ability in deciphering intricate nuances embedded within descriptions and enforce the model to focus on context rather than object names alone. On two novel object detection benchmarks, LVIS and OminiLabel, under the zero-shot detection setting, our approach achieves 34.8 APr minival (+9.1) and 29.3 AP (+3.6), respectively, surpassing the prior state-of-the-art models, GLIP and FIBER, by a large margin.
Liunian Harold Li, Zi-Yi Dou, Nanyun Peng 0001, Kai-Wei Chang 0001
NeurIPS4
2023 Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
abstract
Large language models (LLMs) have achieved remarkable progress in solving various natural language processing tasks due to emergent reasoning abilities. However, LLMs have inherent limitations as they are incapable of accessing up-to-date information (stored on the Web or in task-specific knowledge bases), using external tools, and performing precise mathematical and logical reasoning. In this paper, we present Chameleon, an AI system that mitigates these limitations by augmenting LLMs with plug-and-play modules for compositional reasoning. Chameleon synthesizes programs by composing various tools (e.g., LLMs, off-the-shelf vision models, web search engines, Python functions, and heuristic-based modules) for accomplishing complex reasoning tasks. At the heart of Chameleon is an LLM-based planner that assembles a sequence of tools to execute to generate the final response. We showcase the effectiveness of Chameleon on two multi-modal knowledge-intensive reasoning tasks: ScienceQA and TabMWP. Chameleon, powered by GPT-4, achieves an 86.54% overall accuracy on ScienceQA, improving the best published few-shot result by 11.37%. On TabMWP, GPT-4-powered Chameleon improves the accuracy by 17.0%, lifting the state of the art to 98.78%. Our analysis also shows that the GPT-4-powered planner exhibits more consistent and rational tool selection via inferring potential constraints from instructions, compared to a ChatGPT-powered planner.
Pan Lu, Baolin Peng, Hao Cheng 0002, Michel Galley, Kai-Wei Chang 0001, Ying Nian Wu, Song-Chun Zhu, Jianfeng Gao 0001
NeurIPS5
2023 Incorporating Fairness in Large Scale NLU Systems
abstract
NLU models power several user facing experiences such as conversations agents and chat bots. Building NLU models typically consist of 3 stages: a) building or finetuning a pre-trained model b) distilling or fine-tuning the pre-trained model to build task specific models and, c) deploying the task-specific model to production. In this presentation, we will identify fairness considerations that can be incorporated in the aforementioned three stages in the life-cycle of NLU model building: (i) selection/building of a large scale language model, (ii) distillation/fine-tuning the large model into task specific model and, (iii) deployment of the task specific model. We will present select metrics that can be used to quantify fairness in NLU models and fairness enhancement techniques that can be deployed in each of these stages. Finally, we will share some recommendations to successfully implement fairness considerations when building an industrial scale NLU system.
Rahul Gupta 0001, Lisa Bauer, Kai-Wei Chang 0001, Jwala Dhamala, Aram Galstyan, Palash Goyal, Avni Khatri, Rohit Parimi, Charith Peris, Apurv Verma, Richard S. Zemel, Premkumar Natarajan
WSDM3
2022 PYLON: A PyTorch Framework for Learning with Constraints
abstract
Deep learning excels at learning task information from large amounts of data, but struggles with learning from declarative high-level knowledge that can be more succinctly expressed directly. In this work, we introduce PYLON, a neuro-symbolic training framework that builds on PyTorch to augment procedurally trained models with declaratively specified knowledge. PYLON lets users programmatically specify constraints as Python functions and compiles them into a differentiable loss, thus training predictive models that fit the data whilst satisfying the specified constraints. PYLON includes both exact as well as approximate compilers to efficiently compute the loss, employing fuzzy logic, sampling methods, and circuits, ensuring scalability even to complex models and constraints. Crucially, a guiding principle in designing PYLON is the ease with which any existing deep learning codebase can be extended to learn from constraints in a few lines code: a function that expresses the constraint, and a single line to compile it into a loss. Our demo comprises of models in NLP, computer vision, logical games, and knowledge graphs that can be interactively trained using constraints as supervision.
Kareem Ahmed, Tao Li 0039, Thy Ton, Quan Guo, Kai-Wei Chang 0001, Parisa Kordjamshidi, Vivek Srikumar, Guy Van den Broeck, Sameer Singh 0001
AAAI5
2022 SGEITL: Scene Graph Enhanced Image-Text Learning for Visual Commonsense Reasoning
abstract
Answering complex questions about images is an ambitious goal for machine intelligence, which requires a joint understanding of images, text, and commonsense knowledge, as well as a strong reasoning ability. Recently, multimodal Transformers have made a great progress in the task of Visual Commonsense Reasoning (VCR), by jointly understanding visual objects and text tokens through layers of cross-modality attention. However, these approaches do not utilize the rich structure of the scene and the interactions between objects which are essential in answering complex commonsense questions. We propose a Scene Graph Enhanced Image-Text Learning (SGEITL) framework to incorporate visual scene graph in commonsense reasoning. In order to exploit the scene graph structure, at the model structure level, we propose a multihop graph transformer for regularizing attention interaction among hops. As for pre-training, a scene-graph-aware pre-training method is proposed to leverage structure knowledge extracted in visual scene graph. Moreover, we introduce a method to train and generate domain relevant visual scene graph using textual annotations in a weakly-supervised manner. Extensive experiments on VCR and other tasks show significant performance boost compared with the state-of-the-art methods, and prove the efficacy of each proposed component.
Zhecan Wang, Haoxuan You, Liunian Harold Li, Alireza Zareian, Suji Park, Yiqing Liang, Kai-Wei Chang 0001, Shih-Fu Chang
AAAI7
2022 Multilingual Generative Language Models for Zero-Shot Cross-Lingual Event Argument Extraction
abstract
We present a study on leveraging multilingual pre-trained generative language models for zero-shot cross-lingual event argument extraction (EAE).By formulating EAE as a language generation task, our method effectively encodes event structures and captures the dependencies between arguments.We design language-agnostic templates to represent the event argument structures, which are compatible with any language, hence facilitating the cross-lingual transfer.Our proposed model finetunes multilingual pre-trained generative language models to generate sentences that fill in the language-agnostic template with arguments extracted from the input passage.The model is trained on source languages and is then directly applied to target languages for event argument extraction.Experiments demonstrate that the proposed model outperforms the current state-of-the-art models on zero-shot cross-lingual EAE.Comprehensive studies and error analyses are presented to better understand the advantages and the current limitations of using generative language models for zero-shot cross-lingual transfer EAE. *The authors contribute equally.Attacker Place Target Attacker Target 接近高级军官的消息灵通人士 说,南斯拉夫 军队 不会离 开军营去干涉 反对派 起义。 Australian commandos , who have been operating deep in Iraq , destroyed a command and control post and killed a number of soldiers.
Kuan-Hao Huang, I-Hung Hsu, Premkumar Natarajan, Kai-Wei Chang 0001, Nanyun Peng 0001
ACL (1)4
2022 Measuring Fairness of Text Classifiers via Prediction Sensitivity
abstract
Satyapriya Krishna, Rahul Gupta, Apurv Verma, Jwala Dhamala, Yada Pruksachatkun, Kai-Wei Chang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Satyapriya Krishna, Rahul Gupta 0001, Apurv Verma, Jwala Dhamala, Yada Pruksachatkun, Kai-Wei Chang 0001
ACL (1)6
2022 On the Sensitivity and Stability of Model Interpretations in NLP
abstract
Recent years have witnessed the emergence of a variety of post-hoc interpretations that aim to uncover how natural language processing (NLP) models make predictions.Despite the surge of new interpretation methods, it remains an open problem how to define and quantitatively measure the faithfulness of interpretations, i.e., to what extent interpretations reflect the reasoning process by a model.We propose two new criteria, sensitivity and stability, that provide complementary notions of faithfulness to the existed removal-based criteria.Our results show that the conclusion for how faithful interpretations are could vary substantially based on different notions.Motivated by the desiderata of sensitivity and stability, we introduce a new class of interpretation methods that adopt techniques from adversarial robustness.Empirical results show that our proposed methods are effective under the new criteria and overcome limitations of gradient-based methods on removal-based criteria.Besides text classification, we also apply interpretation methods and metrics to dependency parsing.Our results shed light on understanding the diverse set of interpretations.
Fan Yin, Zhouxing Shi, Cho-Jui Hsieh, Kai-Wei Chang 0001
ACL (1)4
2022 Grounded Language-Image Pre-training
abstract
This paper presents a grounded language-image pretraining (GLIP) model for learning object-level, language-aware, and semantic-rich visual representations. GLIP unifies object detection and phrase grounding for pre-training. The unification brings two benefits: 1) it allows GLIP to learn from both detection and grounding data to improve both tasks and bootstrap a good grounding model; 2) GLIP can leverage massive image-text pairs by generating grounding boxes in a self-training fashion, making the learned representations semantic-rich. In our experiments, we pre-train GLIP on 27M grounding data, including 3M human-annotated and 24M web-crawled image-text pairs. The learned representations demonstrate strong zero-shot and few-shot transferability to various object-level recognition tasks. 1) When directly evaluated on COCO and LVIS (without seeing any images in COCO during pre-training), GLIP achieves 49.8 AP and 26.9 AP, respectively, surpassing many supervised baselines.11Supervised baselines on COCO object detection: Faster-RCNN w/ ResNet50 (40.2) or ResNet101 (42.0), and DyHead w/ Swin-Tiny (49.7). 2) After fine-tuned on COCO, GLIP achieves 60.8 AP on val and 61.5 AP on test-dev, surpassing prior SoTA. 3) When transferred to 13 downstream object detection tasks, a 1-shot GLIP rivals with a fully-supervised Dynamic Head. Code will be released at https://github.com/microsoft/GLIP.
Liunian Harold Li, Pengchuan Zhang, Haotian Zhang 0005, Chunyuan Li, Yiwu Zhong, Lu Yuan 0001, Lei Zhang 0001, Jenq-Neng Hwang, Kai-Wei Chang 0001, Jianfeng Gao 0001
CVPR11
2022 How well can Text-to-Image Generative Models understand Ethical Natural Language Interventions?
abstract
Text-to-image generative models have achieved unprecedented success in generating highquality images based on natural language descriptions.However, it is shown that these models tend to favor specific social groups when prompted with neutral text descriptions (e.g., 'a photo of a lawyer').Following Zhao et al. (2021), we study the effect on the diversity of the generated images when adding ethical intervention that supports equitable judgment (e.g., 'if all individuals can be a lawyer irrespective of their gender') in the input prompts.To this end, we introduce an Ethical NaTural Language Interventions in Text-to-Image GENeration (ENTIGEN) benchmark dataset to evaluate the change in image generations conditional on ethical interventions across three social axesgender, skin color, and culture.Through ENTI-GEN framework, we find that the generations from minDALL•E, DALL•E-mini and Stable Diffusion cover diverse social groups while preserving the image quality.Preliminary studies indicate that a large change in the model predictions is triggered by certain phrases such as 'irrespective of gender' in the context of gender bias in the ethical interventions.We release code and annotated data at https://github. com/Hritikbansal/entigen_emnlp.
Hritik Bansal, Da Yin, Masoud Monajatipoor, Kai-Wei Chang 0001
EMNLP4
2022 Empowering Language Models with Knowledge Graph Reasoning for Open-Domain Question Answering
abstract
Answering open-domain questions requires world knowledge about in-context entities.As pre-trained Language Models (LMs) lack the power to store all required knowledge, external knowledge sources, such as knowledge graphs, are often used to augment LMs.In this work, we propose knOwledge REasOning empowered Language Model (OREOLM), which consists of a novel Knowledge Interaction Layer that can be flexibly plugged into existing Transformer-based LMs to interact with a differentiable Knowledge Graph Reasoning module collaboratively.In this way, LM guides KG to walk towards the desired answer, while the retrieved knowledge improves LM.By adopting OREOLM to RoBERTa and T5, we show significant performance gain, achieving state-of-art results in the Closed-Book setting.The performance enhancement is mainly from the KG reasoning's capacity to infer missing relational facts.In addition, OREOLM provides reasoning paths as rationales to interpret the model's decision.
Ziniu Hu, Yichong Xu, Wenhao Yu 0002, Shuohang Wang, Ziyi Yang 0011, Chenguang Zhu 0001, Kai-Wei Chang 0001, Yizhou Sun
EMNLP7
2022 Understanding ME? Multimodal Evaluation for Fine-grained Visual Commonsense
abstract
Visual commonsense understanding requires Vision Language (VL) models to not only understand image and text but also crossreference in-between to fully integrate and achieve comprehension of the visual scene described.Recently, various approaches have been developed and have achieved high performance on visual commonsense benchmarks.However, it is unclear whether the models really understand the visual scene and underlying commonsense knowledge due to limited evaluation data resources.To provide an indepth analysis, we present a Multimodal Evaluation (ME) pipeline to automatically generate question-answer pairs to test models' understanding of the visual scene, text, and related knowledge.We then take a step further to show that training with the ME data boosts model's performance in standard VCR evaluation.Lastly, our in-depth analysis and comparison reveal interesting findings: (1) semantically low-level information can assist learning of high-level information but not the opposite;(2) visual information is generally under utilization compared with text.
Zhecan Wang, Haoxuan You, Yicheng He, Kai-Wei Chang 0001, Shih-Fu Chang
EMNLP5
2022 GeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language Models
abstract
Recent work has shown that Pre-trained Language Models (PLMs) store the relational knowledge learned from data and utilize it for performing downstream tasks.However, commonsense knowledge across different regions may vary.For instance, the color of bridal dress is white in American weddings whereas it is red in Chinese weddings.In this paper, we introduce a benchmark dataset, Geo-diverse Commonsense Multilingual Language Models Analysis (GEOMLAMA), for probing the diversity of the relational knowledge in multilingual PLMs.GEOMLAMA contains 3,125 prompts in English, Chinese, Hindi, Persian, and Swahili, with a wide coverage of concepts shared by people from American, Chinese, Indian, Iranian and Kenyan cultures.We benchmark 11 standard multilingual PLMs on GE-OMLAMA.Interestingly, we find that 1) larger multilingual PLMs variants do not necessarily store geo-diverse concepts better than its smaller variant; 2) multilingual PLMs are not intrinsically biased towards knowledge from the Western countries (the United States); 3) the native language of a country may not be the best language to probe its knowledge and 4) a language may better probe knowledge about a nonnative country than its native country.
Da Yin, Hritik Bansal, Masoud Monajatipoor, Liunian Harold Li, Kai-Wei Chang 0001
EMNLP5
2022 ADDMU: Detection of Far-Boundary Adversarial Examples with Data and Model Uncertainty Estimation
abstract
Adversarial Examples Detection (AED) is a crucial defense technique against adversarial attacks and has drawn increasing attention from the Natural Language Processing (NLP) community.Despite the surge of new AED methods, our studies show that existing methods heavily rely on a shortcut to achieve good performance.In other words, current search-based adversarial attacks in NLP stop once model predictions change, and thus most adversarial examples generated by those attacks are located near model decision boundaries.To surpass this shortcut and fairly evaluate AED methods, we propose to test AED methods with Far Boundary (FB) adversarial examples.Existing methods show worse than random guess performance under this scenario.To overcome this limitation, we propose a new technique, ADDMU, adversary detection with data and model uncertainty, which combines two types of uncertainty estimation for both regular and FB adversarial example detection.Our new method outperforms previous methods by 3.6 and 6.0 AUC points under each scenario.Finally, our analysis shows that the two types of uncertainty provided by ADDMU can be leveraged to characterize adversarial examples and identify the ones that contribute most to model's robustness in adversarial training.
Fan Yin, Yao Li 0015, Cho-Jui Hsieh, Kai-Wei Chang 0001
EMNLP4
2022 Toward Degradation-Robust Voice Conversion
abstract
Any-to-any voice conversion technologies convert the vocal timbre of an utterance to any speaker even unseen during training. Although there have been several state-of-the-art any-to-any voice conversion models, they were all based on clean utterances to convert successfully. However, in real-world scenarios, it is difficult to collect clean utterances of a speaker, and they are usually degraded by noises or reverberations. It thus becomes highly desired to understand how these degradations affect voice conversion and build a degradation-robust model. We report in this paper the first comprehensive study on the degradation robustness of any-to-any voice conversion. We show that the performance of state-of-the-art models nowadays was severely hampered given degraded utterances. To this end, we then propose speech enhancement concatenation and denoising training to improve the robustness. In addition to common degradations, we also consider adversarial noises, which alter the model output significantly yet are human-imperceptible. It was shown that both concatenations with off-the-shelf speech enhancement models and denoising training on voice conversion models could improve the robustness, while each of them had pros and cons.
Chien-Yu Huang, Kai-Wei Chang 0001, Hung-yi Lee
ICASSP2
2022 How Much Can CLIP Benefit Vision-and-Language Tasks?
Sheng Shen 0001, Liunian Harold Li, Hao Tan 0002, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang 0001, Zhewei Yao, Kurt Keutzer
ICLR6
2022 An Exploration of Prompt Tuning on Generative Spoken Language Model for Speech Processing Tasks
abstract
Speech representations learned from Self-supervised learning (SSL) models can benefit various speech processing tasks. However, utilizing SSL representations usually requires fine-tuning the pre-trained models or designing task-specific downstream models and loss functions, causing much memory usage and human labor. Recently, prompting in Natural Language Processing (NLP) has been found to be an efficient technique to leverage pre-trained language models (LMs). Specifically, prompt tuning optimizes a limited number of task-specific parameters with a fixed pre-trained model; as a result, only a small set of parameters is needed to be stored for each task. Prompt tuning improves computation and memory efficiency by leveraging the pre-trained LM's prediction ability. Nevertheless, such a paradigm is little studied in the speech community. We report in this paper the first exploration of the prompt tuning paradigm for speech processing tasks based on Generative Spoken Language Model (GSLM). Experiment results show that the prompt tuning technique achieves competitive performance in speech classification tasks with fewer trainable parameters than fine-tuning specialized downstream models. We further study the technique in challenging sequence generation tasks. Prompt tuning also demonstrates its potential, while the limitation and possible research directions are discussed in this paper. The source code is available on https://github.com/ga642381/SpeechPrompt.
Kai-Wei Chang 0001, Wei-Cheng Tseng, Shang-Wen Li 0001, Hung-yi Lee
INTERSPEECH1
2022 BERTHop: An Effective Vision-and-Language Model for Chest X-ray Disease Diagnosis
Masoud Monajatipoor, Mozhdeh Rouhsedaghat, Liunian Harold Li, C.-C. Jay Kuo, Aichi Chien, Kai-Wei Chang 0001
MICCAI (5)6
2022 DEGREE: A Data-Efficient Generation-Based Event Extraction Model
abstract
I-Hung Hsu, Kuan-Hao Huang, Elizabeth Boschee, Scott Miller, Prem Natarajan, Kai-Wei Chang, Nanyun Peng. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
I-Hung Hsu, Kuan-Hao Huang, Elizabeth Boschee, Premkumar Natarajan, Kai-Wei Chang 0001, Nanyun Peng 0001
NAACL-HLT6
2022 Socially Aware Bias Measurements for Hindi Language Representations
abstract
Vijit Malik, Sunipa Dev, Akihiro Nishi, Nanyun Peng, Kai-Wei Chang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Vijit Malik, Sunipa Dev, Akihiro Nishi, Nanyun Peng 0001, Kai-Wei Chang 0001
NAACL-HLT5
2022 Semantic Probabilistic Layers for Neuro-Symbolic Learning
abstract
We design a predictive layer for structured-output prediction (SOP) that can be plugged into any neural network guaranteeing its predictions are consistent with a set of predefined symbolic constraints. Our Semantic Probabilistic Layer (SPL) can model intricate correlations, and hard constraints, over a structured output space all while being amenable to end-to-end learning via maximum likelihood.SPLs combine exact probabilistic inference with logical reasoning in a clean and modular way, learning complex distributions and restricting their support to solutions of the constraint. As such, they can faithfully, and efficiently, model complex SOP tasks beyond the reach of alternative neuro-symbolic approaches. We empirically demonstrate that SPLs outperform these competitors in terms of accuracy on challenging SOP tasks such as hierarchical multi-label classification, pathfinding and preference learning, while retaining perfect constraint satisfaction.
Kareem Ahmed, Stefano Teso, Kai-Wei Chang 0001, Guy Van den Broeck, Antonio Vergari
NeurIPS3
2022 Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
abstract
When answering a question, humans utilize the information available across different modalities to synthesize a consistent and complete chain of thought (CoT). This process is normally a black box in the case of deep learning models like large-scale language models. Recently, science question benchmarks have been used to diagnose the multi-hop reasoning ability and interpretability of an AI system. However, existing datasets fail to provide annotations for the answers, or are restricted to the textual-only modality, small scales, and limited domain diversity. To this end, we present Science Question Answering (ScienceQA), a new benchmark that consists of ~21k multimodal multiple choice questions with a diverse set of science topics and annotations of their answers with corresponding lectures and explanations. We further design language models to learn to generate lectures and explanations as the chain of thought (CoT) to mimic the multi-hop reasoning process when answering ScienceQA questions. ScienceQA demonstrates the utility of CoT in language models, as CoT improves the question answering performance by 1.20% in few-shot GPT-3 and 3.99% in fine-tuned UnifiedQA. We also explore the upper bound for models to leverage explanations by feeding those in the input; we observe that it improves the few-shot performance of GPT-3 by 18.96%. Our analysis further shows that language models, similar to humans, benefit from explanations to learn from fewer data and achieve the same performance with just 40% of the data. The data and code are available at https://scienceqa.github.io.
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 0001, Kai-Wei Chang 0001, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, Ashwin Kalyan
NeurIPS5
2022 Controllable Text Generation with Neurally-Decomposed Oracle
abstract
We propose a general and efficient framework to control auto-regressive generation models with NeurAlly-Decomposed Oracle (NADO). Given a pre-trained base language model and a sequence-level boolean oracle function, we aim to decompose the oracle function into token-level guidance to steer the base model in text generation. Specifically, the token-level guidance is provided by NADO, a neural model trained with examples sampled from the base model, demanding no additional auxiliary labeled data. Based on posterior regularization, we present the close-form optimal solution to incorporate the decomposed token-level guidance into the base model for controllable generation. We further discuss how the neural approximation affects the quality of the solution. These experiments conducted on two different applications: (1) text generation with lexical constraints and (2) machine translation with formality control demonstrate that our framework efficiently guides the base model towards the given oracle while keeping high generation quality.
Sidi Lu, Nanyun Peng 0001, Kai-Wei Chang 0001
NeurIPS4
2022 On the Discrimination Risk of Mean Aggregation Feature Imputation in Graphs
abstract
In human networks, nodes belonging to a marginalized group often have a disproportionate rate of unknown or missing features. This, in conjunction with graph structure and known feature biases, can cause graph feature imputation algorithms to predict values for unknown features that make the marginalized group's feature values more distinct from the the dominant group's feature values than they are in reality. We call this distinction the discrimination risk. We prove that a higher discrimination risk can amplify the unfairness of a machine learning model applied to the imputed data. We then formalize a general graph feature imputation framework called mean aggregation imputation and theoretically and empirically characterize graphs in which applying this framework can yield feature values with a high discrimination risk. We propose a simple algorithm to ensure mean aggregation-imputed features provably have a low discrimination risk, while minimally sacrificing reconstruction error (with respect to the imputation objective). We evaluate the fairness and accuracy of our solution on synthetic and real-world credit networks.
Arjun Subramonian, Kai-Wei Chang 0001, Yizhou Sun
NeurIPS2
2022 An Analysis of The Effects of Decoding Algorithms on Fairness in Open-Ended Language Generation
abstract
Several prior works have shown that language models (LMs) can generate text containing harmful social biases and stereotypes. While decoding algorithms play a central role in determining properties of LM generated text, their impact on the fairness of the generations has not been studied. We present a systematic analysis of the impact of decoding algorithms on LM fairness, and analyze the trade-off between fairness, diversity and quality. Our experiments with top-p, top-k and temperature decoding algorithms, in open-ended language generation, show that fairness across demographic groups changes significantly with change in decoding algorithm's hyper-parameters. Notably, decoding algorithms that output more diverse text also output more texts with negative sentiment and regard. We present several findings and provide recommendations on standardized reporting of decoding details in fairness evaluations and optimization of decoding algorithms for fairness alongside quality and diversity.
Jwala Dhamala, Rahul Gupta 0001, Kai-Wei Chang 0001, Aram Galstyan
SLT4
2022 Superb @ SLT 2022: Challenge on Generalization and Efficiency of Self-Supervised Speech Representation Learning
abstract
We present the SUPERB challenge at SLT 2022, which aims at learning self-supervised speech representation for better performance, generalization, and efficiency. The challenge builds upon the SUPERB benchmark and implements metrics to measure the computation requirements of self-supervised learning (SSL) representation and to evaluate its generalizability and performance across the diverse SUPERB tasks. The SUPERB benchmark provides comprehensive coverage of popular speech processing tasks, from speech and speaker recognition to audio generation and semantic understanding. As SSL has gained interest in the speech community and showed promising outcomes, we envision the challenge to uplevel the impact of SSL techniques by motivating more practical designs of techniques beyond task performance. We summarize the results of 14 submitted models in this paper. We also discuss the main findings from those submissions and the future directions of SSL research.
Tzu-hsun Feng, Shuyan Dong, Ching-Feng Yeh, Shu-Wen Yang, Tzu-Quan Lin, Jiatong Shi, Kai-Wei Chang 0001, Zili Huang, Xuankai Chang, Shinji Watanabe 0001, Abdel-rahman Mohamed, Shang-Wen Li 0001, Hung-yi Lee
SLT7
2022 Neuro-symbolic entropy regularization
abstract
In structured output prediction, the goal is to jointly predict several output variables that together encode a structured object – a path in a graph, an entity-relation triple, or an ordering of objects. Such a large output space makes learning hard and requires vast amounts of labeled data. Different approaches leverage alternate sources of supervision. One approach – entropy regularization – posits that decision boundaries should lie in low-probability regions. It extracts supervision from unlabeled examples, but remains agnostic to the structure of the output space. Conversely, neuro-symbolic approaches exploit the knowledge that not every prediction corresponds to a valid structure in the output space. Yet, they do not further restrict the learned output distribution.This paper introduces a framework that unifies both approaches. We propose a loss, neuro-symbolic entropy regularization, that encourages the model to confidently predict a valid object. It is obtained by restricting entropy regularization to the distribution over only the valid structures. This loss can be computed efficiently when the output constraint is expressed as a tractable logic circuit. Moreover, it seamlessly integrates with other neuro-symbolic losses that eliminate invalid predictions. We demonstrate the efficacy of our approach on a series of semi-supervised and fully-supervised structured-prediction experiments, where it leads to models whose predictions are more accurate as well as more likely to be valid.
Kareem Ahmed, Kai-Wei Chang 0001, Guy Van den Broeck
UAI3
2021 GATE: Graph Attention Transformer Encoder for Cross-lingual Relation and Event Extraction
abstract
Recent progress in cross-lingual relation and event extraction use graph convolutional networks (GCNs) with universal dependency parses to learn language-agnostic sentence representations such that models trained on one language can be applied to other languages. However, GCNs struggle to model words with long-range dependencies or are not directly connected in the dependency tree. To address these challenges, we propose to utilize the self-attention mechanism where we explicitly fuse structural information to learn the dependencies between words with different syntactic distances. We introduce GATE, a Graph Attention Transformer Encoder, and test its cross-lingual transferability on relation and event extraction tasks. We perform experiments on the ACE05 dataset that includes three typologically different languages: English, Chinese, and Arabic. The evaluation results show that GATE outperforms three recently proposed methods by a large margin. Our detailed analysis reveals that due to the reliance on syntactic dependencies, GATE produces robust representations that facilitate transfer across languages.
Wasi Uddin Ahmad, Nanyun Peng 0001, Kai-Wei Chang 0001
AAAI3
2021 Clinical Temporal Relation Extraction with Probabilistic Soft Logic Regularization and Global Inference
abstract
There has been a steady need in the medical community to precisely extract the temporal relations between clinical events. In particular, temporal information can facilitate a variety of downstream applications such as case report retrieval and medical question answering. Existing methods either require expensive feature engineering or are incapable of modeling the global relational dependencies among the events. In this paper, we propose a novel method, Clinical Temporal ReLation Exaction with Probabilistic Soft Logic Regularization and Global Inference (CTRL-PG) to tackle the problem at the document level. Extensive experiments on two benchmark datasets, I2B2-2012 and TB-Dense, demonstrate that CTRL-PG significantly outperforms baseline methods for temporal relation extraction.
Yichao Zhou 0001, Rujun Han, J. Harry Caufield, Kai-Wei Chang 0001, Yizhou Sun, Peipei Ping, Wei Wang 0010
AAAI5
2021 Select, Extract and Generate: Neural Keyphrase Generation with Layer-wise Coverage Attention
abstract
Wasi Ahmad, Xiao Bai, Soomin Lee, Kai-Wei Chang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Wasi Uddin Ahmad, Xiao Bai 0002, Kai-Wei Chang 0001
ACL/IJCNLP (1)4
2021 Intent Classification and Slot Filling for Privacy Policies
abstract
Wasi Ahmad, Jianfeng Chi, Tu Le, Thomas Norton, Yuan Tian, Kai-Wei Chang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Wasi Uddin Ahmad, Jianfeng Chi, Tu Le, Thomas Norton, Yuan Tian 0001, Kai-Wei Chang 0001
ACL/IJCNLP (1)6
2021 Syntax-augmented Multilingual BERT for Cross-lingual Transfer
abstract
In recent years, we have seen a colossal effort in pre-training multilingual text encoders using large-scale corpora in many languages to facilitate cross-lingual transfer learning. However, due to typological differences across languages, the cross-lingual transfer is challenging. Nevertheless, language syntax, e.g., syntactic dependencies, can bridge the typological gap. Previous works have shown that pre-trained multilingual encoders, such as mBERT (CITATION), capture language syntax, helping cross-lingual transfer. This work shows that explicitly providing language syntax and training mBERT using an auxiliary objective to encode the universal dependency tree structure helps cross-lingual transfer. We perform rigorous experiments on four NLP tasks, including text classification, question answering, named entity recognition, and task-oriented semantic parsing. The experiment results show that syntax-augmented mBERT improves cross-lingual transfer on popular benchmarks, such as PAWS-X and MLQA, by 1.4 and 1.6 points on average across all languages. In the generalized transfer setting, the performance boosted significantly, with 3.9 and 3.1 points on average in PAWS-X and MLQA.
Wasi Uddin Ahmad, Haoran Li 0007, Kai-Wei Chang 0001, Yashar Mehdad
ACL/IJCNLP (1)3
2021 Societal Biases in Language Generation: Progress and Challenges
abstract
Emily Sheng, Kai-Wei Chang, Prem Natarajan, Nanyun Peng. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Emily Sheng, Kai-Wei Chang 0001, Premkumar Natarajan, Nanyun Peng 0001
ACL/IJCNLP (1)2
2021 Defense against Synonym Substitution-based Adversarial Attacks via Dirichlet Neighborhood Ensemble
abstract
Yi Zhou, Xiaoqing Zheng, Cho-Jui Hsieh, Kai-Wei Chang, Xuanjing Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yi Zhou 0018, Xiaoqing Zheng, Cho-Jui Hsieh, Kai-Wei Chang 0001, Xuanjing Huang 0001
ACL/IJCNLP (1)4
2021 Generating Syntactically Controlled Paraphrases without Using Annotated Parallel Pairs
abstract
Paraphrase generation plays an essential role in natural language process (NLP), and it has many downstream applications.However, training supervised paraphrase models requires many annotated paraphrase pairs, which are usually costly to obtain.On the other hand, the paraphrases generated by existing unsupervised approaches are usually syntactically similar to the source sentences and are limited in diversity.In this paper, we demonstrate that it is possible to generate syntactically various paraphrases without the need for annotated paraphrase pairs.We propose Syntactically controlled Paraphrase Generator (SynPG), an encoder-decoder based model that learns to disentangle the semantics and the syntax of a sentence from a collection of unannotated texts.The disentanglement enables SynPG to control the syntax of output paraphrases by manipulating the embedding in the syntactic space.Extensive experiments using automatic metrics and human evaluation show that SynPG performs better syntactic control than unsupervised baselines, while the quality of the generated paraphrases is competitive.We also demonstrate that the performance of SynPG is competitive or even better than supervised models when the unannotated data is large.Finally, we show that the syntactically controlled paraphrases generated by SynPG can be utilized for data augmentation to improve the robustness of NLP models.
Kuan-Hao Huang, Kai-Wei Chang 0001
EACL2
2021 Harms of Gender Exclusivity and Challenges in Non-Binary Representation in Language Technologies
abstract
Content Warning: This paper contains examples of stereotypes and associations, misgendering, erasure, and other harms that could be offensive and triggering to trans and nonbinary individuals.Gender is widely discussed in the context of language tasks and when examining the stereotypes propagated by language models.However, current discussions primarily treat gender as binary, which can perpetuate harms such as the cyclical erasure of non-binary gender identities.These harms are driven by model and dataset biases, which are consequences of the non-recognition and lack of understanding of non-binary genders in society.In this paper, we explain the complexity of gender and language around it, and survey non-binary persons to understand harms associated with the treatment of gender as binary in English language technologies.We also detail how current language representations (e.g., GloVe, BERT) capture and perpetuate these harms and related challenges that need to be acknowledged and addressed for representations to equitably encode gender information.
Sunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian, Jeff M. Phillips, Kai-Wei Chang 0001
EMNLP (1)6
2021 Improving Zero-Shot Cross-Lingual Transfer Learning via Robust Training
abstract
Pre-trained multilingual language encoders, such as multilingual BERT and XLM-R, show great potential for zero-shot cross-lingual transfer.However, these multilingual encoders do not precisely align words and phrases across languages.Especially, learning alignments in the multilingual embedding space usually requires sentence-level or word-level parallel corpora, which are expensive to be obtained for low-resource languages.An alternative is to make the multilingual encoders more robust; when fine-tuning the encoder using downstream task, we train the encoder to tolerate noise in the contextual embedding spaces such that even if the representations of different languages are not aligned well, the model can still achieve good performance on zero-shot cross-lingual transfer.In this work, we propose a learning strategy for training robust models by drawing connections between adversarial examples and the failure cases of zero-shot cross-lingual transfer.We adopt two widely used robust training methods, adversarial training and randomized smoothing, to train the desired robust model.The experimental results demonstrate that robust training improves zero-shot cross-lingual transfer on text classification tasks.The improvement is more significant in the generalized crosslingual transfer setting, where the pair of input sentences belong to two different languages.
Kuan-Hao Huang, Wasi Uddin Ahmad, Nanyun Peng 0001, Kai-Wei Chang 0001
EMNLP (1)4
2021 Searching for an Effective Defender: Benchmarking Defense against Adversarial Word Substitution
abstract
Recent studies have shown that deep neural network-based models are vulnerable to intentionally crafted adversarial examples, and various methods have been proposed to defend against adversarial word-substitution attacks for neural NLP models.However, there is a lack of systematic study on comparing different defense approaches under the same attacking setting.In this paper, we seek to fill the gap through comprehensive studies on the behavior of neural text classifiers trained with various defense methods against representative adversarial attacks.In addition, we propose an effective method to further improve the robustness of neural text classifiers against such attacks, and achieved the highest accuracy on both clean and adversarial examples on AGNEWS and IMDB datasets, outperforming existing methods by a significant margin.We hope this study could provide useful clues for future research on text adversarial defense.Codes are available at https:// github.com/RockyLzy/TextDefender.
Zongyi Li, Jianhan Xu, Jiehang Zeng, Linyang Li, Xiaoqing Zheng, Qi Zhang 0001, Kai-Wei Chang 0001, Cho-Jui Hsieh
EMNLP (1)7
2021 Broaden the Vision: Geo-Diverse Visual Commonsense Reasoning
abstract
Commonsense is defined as the knowledge that is shared by everyone.However, certain types of commonsense knowledge are correlated with culture and geographic locations and they are only shared locally.For example, the scenarios of wedding ceremonies vary across regions due to different customs influenced by historical and religious factors.Such regional characteristics, however, are generally omitted in prior work.In this paper, we construct a Geo-Diverse Visual Commonsense Reasoning dataset (GD-VCR) to test vision-and-language models' ability to understand cultural and geo-location-specific commonsense.In particular, we study two state-of-the-art Vision-and-Language models, VisualBERT and ViLBERT trained on VCR, a standard multimodal commonsense benchmark with images primarily from Western regions.We then evaluate how well the trained models can generalize to answering the questions in GD-VCR.We find that the performance of both models for non-Western regions including East Asia, South Asia, and Africa is significantly lower than that for Western region.We analyze the reasons behind the performance disparity and find that the performance gap is larger on QA pairs that: 1) are concerned with culture-related scenarios, e.g., weddings, religious activities, and festivals; 2) require high-level geo-diverse commonsense reasoning rather than low-order perception and recognition.Dataset and code are released at https://github.com/ WadeYin9712/GD-VCR.
Da Yin, Liunian Harold Li, Ziniu Hu, Nanyun Peng 0001, Kai-Wei Chang 0001
EMNLP (1)5
2021 On the Transferability of Adversarial Attacks against Neural Text Classifier
abstract
Deep neural networks are vulnerable to adversarial attacks, where a small perturbation to an input alters the model prediction.In many cases, malicious inputs intentionally crafted for one model can fool another model.In this paper, we present the first study to systematically investigate the transferability of adversarial examples for text classification models and explore how various factors, including network architecture, tokenization scheme, word embedding, and model capacity, affect the transferability of adversarial examples.Based on these studies, we propose a genetic algorithm to find an ensemble of models that can be used to induce adversarial examples to fool almost all existing models.Such adversarial examples reflect the defects of the learning process and the data bias in the training set.Finally, we derive word replacement rules that can be used for model diagnostics from these adversarial examples. A.2 Transferability among Different Neural ModelsWe show in Figure 4 the transferability rate among all neural models in the model pool.The column and row headers indicate the IDs of source and target models respectively.The mapping of IDs and the corresponding models is shown in Figure 8.We generate adversarial examples by attacking a source model, and report the transferability rates on a target (or victim) model. 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 1 .71.50.68.66.21.30.50.39.48.48.22.27.51.49.57.50.21.29 .44 .38 2 .55.63.58.59.21.33 .46.47.46.48.20 .31.49.59.52 .50.18.30.45.47 3 .68.53.70.70.22.28.49.41.50.48.21.26.52 .50.58.55.20 .27.44 .41 4 .
Liping Yuan, Xiaoqing Zheng, Yi Zhou 0018, Cho-Jui Hsieh, Kai-Wei Chang 0001
EMNLP (1)5
2021 CREATe: Clinical Report Extraction and Annotation Technology
abstract
Clinical case reports are written descriptions of the unique aspects of a particular clinical case, playing an essential role in sharing clinical experiences about atypical disease phenotypes and new therapies. However, to our knowledge, there has been no attempt to develop an end-to-end system to annotate, index, or otherwise curate these reports. In this paper, we propose a novel computational resource platform, CREATe, for extracting, indexing, and querying the contents of clinical case reports. CREATe fosters an environment of sustainable resource support and discovery, enabling researchers to overcome the challenges of information science. An online video of the demonstration can be viewed at https://youtu.be/Q8owBQYTjDc.
Yichao Zhou 0001, Bowen Zhang 0002, J. Harry Caufield, Kai-Wei Chang 0001, Yizhou Sun, Peipei Ping, Wei Wang 0010
ICDE6
2021 An Integer Linear Programming Framework for Mining Constraints from Data
abstract
Structured output prediction problems (e.g., sequential tagging, hierarchical multi-class classification) often involve constraints over the output space. These constraints interact with the learned models to filter infeasible solutions and facilitate in building an accountable system. However, despite constraints are useful, they are often based on hand-crafted rules. This raises a question – can we mine constraints and rules from data based on a learning algorithm? In this paper, we present a general framework for mining constraints from data. In particular, we consider the inference in structured output prediction as an integer linear programming (ILP) problem. Then, given the coefficients of the objective function and the corresponding solution, we mine the underlying constraints by estimating the outer and inner polytopes of the feasible set. We verify the proposed constraint mining algorithm in various synthetic and real-world applications and demonstrate that the proposed approach successfully identifies the feasible set at scale. In particular, we show that our approach can learn to solve 9x9 Sudoku puzzles and minimal spanning tree problems from examples without providing the underlying rules. Our algorithm can also integrate with a neural network model to learn the hierarchical label structure of a multi-label classification task. Besides, we provide theoretical analysis about the tightness of the polytopes and the reliability of the mined constraints.
Kai-Wei Chang 0001
ICML2
2021 Unified Pre-training for Program Understanding and Generation
abstract
Wasi Ahmad, Saikat Chakraborty, Baishakhi Ray, Kai-Wei Chang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Wasi Uddin Ahmad, Saikat Chakraborty 0001, Baishakhi Ray, Kai-Wei Chang 0001
NAACL-HLT4
2021 Disentangling Semantics and Syntax in Sentence Embeddings with Pre-trained Language Models
abstract
Pre-trained language models have achieved huge success on a wide range of NLP tasks.However, contextual representations from pretrained models contain entangled semantic and syntactic information, and therefore cannot be directly used to derive useful semantic sentence embeddings for some tasks.Paraphrase pairs offer an effective way of learning the distinction between semantics and syntax, as they naturally share semantics and often vary in syntax.In this work, we present ParaBART, a semantic sentence embedding model that learns to disentangle semantics and syntax in sentence embeddings obtained by pre-trained language models.ParaBART is trained to perform syntax-guided paraphrasing, based on a source sentence that shares semantics with the target paraphrase, and a parse tree that specifies the target syntax.In this way, ParaBART learns disentangled semantic and syntactic representations from their respective inputs with separate encoders.Experiments in English show that ParaBART outperforms state-of-theart sentence embedding models on unsupervised semantic similarity tasks.Additionally, we show that our approach can effectively remove syntactic information from semantic sentence embeddings, leading to better robustness against syntactic variation on downstream semantic tasks.
James Y. Huang, Kuan-Hao Huang, Kai-Wei Chang 0001
NAACL-HLT3
2021 Unsupervised Vision-and-Language Pre-training Without Parallel Images and Captions
abstract
Liunian Harold Li, Haoxuan You, Zhecan Wang, Alireza Zareian, Shih-Fu Chang, Kai-Wei Chang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Liunian Harold Li, Haoxuan You, Zhecan Wang, Alireza Zareian, Shih-Fu Chang, Kai-Wei Chang 0001
NAACL-HLT6
2021 Evaluating the Values of Sources in Transfer Learning
abstract
Transfer learning that adapts a model trained on data-rich sources to low-resource targets has been widely applied in natural language processing (NLP).However, when training a transfer model over multiple sources, not every source is equally useful for the target.To better transfer a model, it is essential to understand the values of the sources.In this paper, we develop SEAL-Shap, an efficient source valuation framework for quantifying the usefulness of the sources (e.g., domains/languages) in transfer learning based on the Shapley value method.Experiments and comprehensive analyses on both cross-domain and cross-lingual transfers demonstrate that our framework is not only effective in choosing useful transfer sources but also the source values match the intuitive source-target similarity.
Md. Rizwan Parvez, Kai-Wei Chang 0001
NAACL-HLT2
2021 "Nice Try, Kiddo": Investigating Ad Hominems in Dialogue Responses
abstract
Emily Sheng, Kai-Wei Chang, Prem Natarajan, Nanyun Peng. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Emily Sheng, Kai-Wei Chang 0001, Premkumar Natarajan, Nanyun Peng 0001
NAACL-HLT2
2021 Adapting Coreference Resolution for Processing Violent Death Narratives
abstract
Ankith Uppunda, Susan Cochran, Jacob Foster, Alina Arseniev-Koehler, Vickie Mays, Kai-Wei Chang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Ankith Uppunda, Susan D. Cochran, Jacob G. Foster, Alina Arseniev-Koehler, Vickie M. Mays, Kai-Wei Chang 0001
NAACL-HLT6
2021 Double Perturbation: On the Robustness of Robustness and Counterfactual Bias Evaluation
abstract
Chong Zhang, Jieyu Zhao, Huan Zhang, Kai-Wei Chang, Cho-Jui Hsieh. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Jieyu Zhao 0001, Huan Zhang 0001, Kai-Wei Chang 0001, Cho-Jui Hsieh
NAACL-HLT4
2020 Learning Directional Sentence-Pair Embedding for Natural Language Reasoning (Student Abstract)
abstract
Enabling the models with the ability of reasoning and inference over text is one of the core missions of natural language understanding. Despite deep learning models have shown strong performance on various cross-sentence inference benchmarks, recent work has shown that they are leveraging spurious statistical cues rather than capturing deeper implied relations between pairs of sentences. In this paper, we show that the state-of-the-art language encoding models are especially bad at modeling directional relations between sentences by proposing a new evaluation task: Cause-and-Effect relation prediction task. Back by our curated Cause-and-Effect Relation dataset (Cℰℛ), we also demonstrate that a mutual attention mechanism can guide the model to focus on capturing directional relations between sentences when added to existing transformer-based models. Experiment results show that the proposed approach improves the performance on downstream applications, such as the abductive reasoning task.
Zhenxin Xiao, Kai-Wei Chang 0001
AAAI3
2020 A Transformer-based Approach for Source Code Summarization
abstract
Generating a readable summary that describes the functionality of a program is known as source code summarization.In this task, learning code representation by modeling the pairwise relationship between code tokens to capture their long-range dependencies is crucial.To learn code representation for summarization, we explore the Transformer model that uses a self-attention mechanism and has shown to be effective in capturing long-range dependencies.In this work, we show that despite the approach is simple, it outperforms the state-of-the-art techniques by a significant margin.We perform extensive analysis and ablation studies that reveal several important findings, e.g., the absolute encoding of source code tokens' position hinders, while relative encoding significantly improves the summarization performance.We have made our code publicly available 1 to facilitate future research.
Wasi Uddin Ahmad, Saikat Chakraborty 0001, Baishakhi Ray, Kai-Wei Chang 0001
ACL4
2020 Towards Understanding Gender Bias in Relation Extraction
abstract
Andrew Gaut, Tony Sun, Shirlyn Tang, Yuxin Huang, Jing Qian, Mai ElSherief, Jieyu Zhao, Diba Mirza, Elizabeth Belding, Kai-Wei Chang, William Yang Wang. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Andrew Gaut, Tony Sun, Shirlyn Tang, Mai ElSherief, Jieyu Zhao 0001, Diba Mirza, Elizabeth M. Belding, Kai-Wei Chang 0001, William Yang Wang
ACL10
2020 Mitigating Gender Bias Amplification in Distribution by Posterior Regularization
abstract
Advanced machine learning techniques have boosted the performance of natural language processing.Nevertheless, recent studies, e.g., Zhao et al. (2017) show that these techniques inadvertently capture the societal bias hidden in the corpus and further amplify it.However, their analysis is conducted only on models' top predictions.In this paper, we investigate the gender bias amplification issue from the distribution perspective and demonstrate that the bias is amplified in the view of predicted probability distribution over labels.We further propose a bias mitigation approach based on posterior regularization.With little performance loss, our method can almost remove the bias amplification in the distribution.Our study sheds the light on understanding the bias amplification.* Both authors contributed equally to this work and are listed in alphabetical order.
Shengyu Jia, Jieyu Zhao 0001, Kai-Wei Chang 0001
ACL4
2020 What Does BERT with Vision Look At?
abstract
Pre-trained visually grounded language models such as ViLBERT, LXMERT, and UNITER have achieved significant performance improvement on vision-and-language tasks but what they learn during pre-training remains unclear. In this work, we demonstrate that certain attention heads of a visually grounded language model actively ground elements of language to image regions. Specifically, some heads can map entities to image regions, performing the task known as entity grounding. Some heads can even detect the syntactic relations between non-entity words and image regions, tracking, for example, associations between verbs and regions corresponding to their arguments. We denote this ability as syntactic grounding. We verify grounding both quantitatively and qualitatively, using Flickr30K Entities as a testbed.
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, Kai-Wei Chang 0001
ACL5
2020 On the Robustness of Language Encoders against Grammatical Errors
abstract
We conduct a thorough study to diagnose the behaviors of pre-trained language encoders (ELMo, BERT, and RoBERTa) when confronted with natural grammatical errors.Specifically, we collect real grammatical errors from non-native speakers and conduct adversarial attacks to simulate these errors on clean text data.We use this approach to facilitate debugging models on downstream applications.Results confirm that the performance of all tested models is affected but the degree of impact varies.To interpret model behaviors, we further design a linguistic acceptability task to reveal their abilities in identifying ungrammatical sentences and the position of errors.We find that fixed contextual encoders with a simple classifier trained on the prediction of sentence correctness are able to locate error positions.We also design a cloze test for BERT and discover that BERT captures the interaction between errors and specific tokens in context.Our results shed light on understanding the robustness and behaviors of language encoders against grammatical errors.
Fan Yin, Quanyu Long, Kai-Wei Chang 0001
ACL4
2020 SentiBERT: A Transferable Transformer-Based Architecture for Compositional Sentiment Semantics
abstract
We propose SentiBERT, a variant of BERT that effectively captures compositional sentiment semantics.The model incorporates contextualized representation with binary constituency parse tree to capture semantic composition.Comprehensive experiments demonstrate that SentiBERT achieves competitive performance on phrase-level sentiment classification.We further demonstrate that the sentiment composition learned from the phrase-level annotations on SST can be transferred to other sentiment analysis tasks as well as related tasks, such as emotion classification tasks.Moreover, we conduct ablation studies and design visualization methods to understand SentiBERT.We show that SentiBERT is better than baseline approaches in capturing negation and the contrastive relation and model the compositional sentiment semantics.
Da Yin, Kai-Wei Chang 0001
ACL3
2020 Gender Bias in Multilingual Embeddings and Cross-Lingual Transfer
abstract
Multilingual representations embed words from many languages into a single semantic space such that words with similar meanings are close to each other regardless of the language.These embeddings have been widely used in various settings, such as cross-lingual transfer, where a natural language processing (NLP) model trained on one language is deployed to another language.While the crosslingual transfer techniques are powerful, they carry gender bias from the source to target languages.In this paper, we study gender bias in multilingual embeddings and how it affects transfer learning for NLP applications.We create a multilingual dataset for bias analysis and propose several ways for quantifying bias in multilingual representations from both the intrinsic and extrinsic perspectives.Experimental results show that the magnitude of bias in the multilingual representations changes differently when we align the embeddings to different target spaces and that the alignment direction can also have an influence on the bias in transfer learning.We further provide recommendations for using the multilingual word representations for downstream tasks.
Jieyu Zhao 0001, Subhabrata Mukherjee, Saghar Hosseini, Kai-Wei Chang 0001, Ahmed Awadallah 0001
ACL4
2020 "The Boating Store Had Its Best Sail Ever": Pronunciation-attentive Contextualized Pun Recognition
abstract
Humor plays an important role in human languages and it is essential to model humor when building intelligence systems. Among different forms of humor, puns perform wordplay for humorous effects by employing words with double entendre and high phonetic similarity. However, identifying and modeling puns are challenging as puns usually involved implicit semantic or phonological tricks. In this paper, we propose Pronunciation-attentive Contextualized Pun Recognition (PCPR) to perceive human humor, detect if a sentence contains puns and locate them in the sentence. PCPR derives contextualized representation for each word in a sentence by capturing the association between the surrounding context and its corresponding phonetic symbols. Extensive experiments are conducted on two benchmark datasets. Results demonstrate that the proposed approach significantly outperforms the state-of-the-art methods in pun detection and location tasks. In-depth analyses verify the effectiveness and robustness of PCPR.
Yichao Zhou 0001, Jyun-Yu Jiang, Jieyu Zhao 0001, Kai-Wei Chang 0001, Wei Wang 0010
ACL4
2020 LOGAN: Local Group Bias Detection by Clustering
abstract
Machine learning techniques have been widely used in natural language processing (NLP). However, as revealed by many recent studies, machine learning models often inherit and amplify the societal biases in data. Various metrics have been proposed to quantify biases in model predictions. In particular, several of them evaluate disparity in model performance between protected groups and advantaged groups in the test corpus. However, we argue that evaluating bias at the corpus level is not enough for understanding how biases are embedded in a model. In fact, a model with similar aggregated performance between different groups on the entire data may behave differently on instances in a local region. To analyze and detect such local bias, we propose LOGAN, a new bias detection technique based on clustering. Experiments on toxicity classification and object classification tasks show that LOGAN identifies bias in a local region and allows us to better analyze the biases in model predictions.
Jieyu Zhao 0001, Kai-Wei Chang 0001
EMNLP (1)2
2020 Robustness Verification for Transformers
Zhouxing Shi, Huan Zhang 0001, Kai-Wei Chang 0001, Minlie Huang, Cho-Jui Hsieh
ICLR3
2020 GPT-GNN: Generative Pre-Training of Graph Neural Networks
abstract
Graph neural networks (GNNs) have been demonstrated to be powerful in modeling graph-structured data. However, training GNNs requires abundant task-specific labeled data, which is often arduously expensive to obtain. One effective way to reduce the labeling effort is to pre-train an expressive GNN model on unlabelled data with self-supervision and then transfer the learned model to downstream tasks with only a few labels. In this paper, we present the GPT-GNN framework to initialize GNNs by generative pre-training. GPT-GNN introduces a self-supervised attributed graph generation task to pre-train a GNN so that it can capture the structural and semantic properties of the graph. We factorize the likelihood of graph generation into two components: 1) attribute generation and 2) edge generation. By modeling both components, GPT-GNN captures the inherent dependency between node attributes and graph structure during the generative process. Comprehensive experiments on the billion-scale open academic graph and Amazon recommendation data demonstrate that GPT-GNN significantly outperforms state-of-the-art GNN models without pre-training by up to 9.1% across various downstream tasks?
Ziniu Hu, Yuxiao Dong, Kuansan Wang, Kai-Wei Chang 0001, Yizhou Sun
KDD4
2020 Automatic Perturbation Analysis for Scalable Certified Robustness and Beyond
abstract
Linear relaxation based perturbation analysis (LiRPA) for neural networks, which computes provable linear bounds of output neurons given a certain amount of input perturbation, has become a core component in robustness verification and certified defense. The majority of LiRPA-based methods focus on simple feed-forward networks and need particular manual derivations and implementations when extended to other architectures. In this paper, we develop an automatic framework to enable perturbation analysis on any neural network structures, by generalizing existing LiRPA algorithms such as CROWN to operate on general computational graphs. The flexibility, differentiability and ease of use of our framework allow us to obtain state-of-the-art results on LiRPA based certified defense on fairly complicated networks like DenseNet, ResNeXt and Transformer that are not supported by prior works. Our framework also enables loss fusion, a technique that significantly reduces the computational complexity of LiRPA for certified defense. For the first time, we demonstrate LiRPA based certified defense on Tiny ImageNet and Downscaled ImageNet where previous approaches cannot scale to due to the relatively large number of classes. Our work also yields an open-source library for the community to apply LiRPA to areas beyond certified defense without much LiRPA expertise, e.g., we create a neural network with a provably flat optimization landscape by applying LiRPA to network parameters. Our open source library is available at https://github.com/KaidiXu/auto_LiRPA.
Kaidi Xu, Zhouxing Shi, Huan Zhang 0001, Kai-Wei Chang 0001, Minlie Huang, Bhavya Kailkhura, Xue Lin 0001, Cho-Jui Hsieh
NeurIPS5
2020 Distributed block-diagonal approximation methods for regularized empirical risk minimization
abstract
Abstract In recent years, there is a growing need to train machine learning models on a huge volume of data. Therefore, designing efficient distributed optimization algorithms for empirical risk minimization (ERM) has become an active and challenging research topic. In this paper, we propose a flexible framework for distributed ERM training through solving the dual problem, which provides a unified description and comparison of existing methods. Our approach requires only approximate solutions of the sub-problems involved in the optimization process, and is versatile to be applied on many large-scale machine learning problems including classification, regression, and structured prediction. We show that our framework enjoys global linear convergence for a broad class of non-strongly-convex problems, and some specific choices of the sub-problems can even achieve much faster convergence than existing approaches by a refined analysis. This improved convergence rate is also reflected in the superior empirical performance of our method.
Ching-Pei Lee, Kai-Wei Chang 0001
Mach. Learn.2
2019 Few-Shot Representation Learning for Out-Of-Vocabulary Words
abstract
Existing approaches for learning word embeddings often assume there are sufficient occurrences for each word in the corpus, such that the representation of words can be accurately estimated from their contexts.However, in real-world scenarios, out-of-vocabulary (a.k.a.OOV) words that do not appear in training corpus emerge frequently.It is challenging to learn accurate representations of these words with only a few observations.In this paper, we formulate the learning of OOV embeddings as a few-shot regression problem, and address it by training a representation function to predict the oracle embedding vector (defined as embedding trained with abundant observations) based on limited observations.Specifically, we propose a novel hierarchical attention-based architecture to serve as the neural regression function, with which the context information of a word is encoded and aggregated from K observations.Furthermore, our approach can leverage Model-Agnostic Meta-Learning (MAML) for adapting the learned model to the new corpus fast and robustly.Experiments show that the proposed approach significantly outperforms existing methods in constructing accurate embeddings for OOV words, and improves downstream tasks where these embeddings are utilized.
Ziniu Hu, Ting Chen 0007, Kai-Wei Chang 0001, Yizhou Sun
ACL (1)3
2019 Mitigating Gender Bias in Natural Language Processing: Literature Review
abstract
Tony Sun, Andrew Gaut, Shirlyn Tang, Yuxin Huang, Mai ElSherief, Jieyu Zhao, Diba Mirza, Elizabeth Belding, Kai-Wei Chang, William Yang Wang. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Tony Sun, Andrew Gaut, Shirlyn Tang, Mai ElSherief, Jieyu Zhao 0001, Diba Mirza, Elizabeth M. Belding, Kai-Wei Chang 0001, William Yang Wang
ACL (1)9
2019 Cross-Lingual Dependency Parsing with Unlabeled Auxiliary Languages
abstract
Cross-lingual transfer learning has become an important weapon to battle the unavailability of annotated resources for low-resource languages.One of the fundamental techniques to transfer across languages is learning language-agnostic representations, in the form of word embeddings or contextual encodings.In this work, we propose to leverage unannotated sentences from auxiliary languages to help learning language-agnostic representations.Specifically, we explore adversarial training for learning contextual encoders that produce invariant representations across languages to facilitate cross-lingual transfer.We conduct experiments on cross-lingual dependency parsing where we train a dependency parser on a source language and transfer it to a wide range of target languages.Experiments on 28 target languages demonstrate that adversarial training significantly improves the overall transfer performances under several different settings.We conduct a careful analysis to evaluate the language-agnostic representations resulted from adversarial training.
Wasi Uddin Ahmad, Zhisong Zhang, Xuezhe Ma, Kai-Wei Chang 0001, Nanyun Peng 0001
CoNLL4
2019 Learning to Represent Bilingual Dictionaries
abstract
Bilingual word embeddings have been widely used to capture the correspondence of lexical semantics in different human languages.However, the cross-lingual correspondence between sentences and words is less studied, despite that this correspondence can significantly benefit many applications such as crosslingual semantic search and textual inference.To bridge this gap, we propose a neural embedding model that leverages bilingual dictionaries 1 .The proposed model is trained to map the lexical definitions to the cross-lingual target words, for which we explore with different sentence encoding techniques.To enhance the learning process on limited resources, our model adopts several critical learning strategies, including multi-task learning on different bridges of languages, and joint learning of the dictionary model with a bilingual word embedding model.We conduct experiments on two new tasks.In the cross-lingual reverse dictionary retrieval task, we demonstrate that our model is capable of comprehending bilingual concepts based on descriptions, and the proposed learning strategies are effective.In the bilingual paraphrase identification task, we show that our model effectively associates sentences in different languages via a shared embedding space, and outperforms existing approaches in identifying bilingual paraphrases.
Muhao Chen 0001, Yingtao Tian, Haochen Chen, Kai-Wei Chang 0001, Steven Skiena, Carlo Zaniolo
CoNLL4
2019 Target Language-Aware Constrained Inference for Cross-lingual Dependency Parsing
abstract
Tao Meng, Nanyun Peng, Kai-Wei Chang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Nanyun Peng 0001, Kai-Wei Chang 0001
EMNLP/IJCNLP (1)3
2019 Robust Text Classifier on Test-Time Budgets
abstract
Md Rizwan Parvez, Tolga Bolukbasi, Kai-Wei Chang, Venkatesh Saligrama. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Md. Rizwan Parvez, Tolga Bolukbasi, Kai-Wei Chang 0001, Venkatesh Saligrama
EMNLP/IJCNLP (1)3
2019 The Woman Worked as a Babysitter: On Biases in Language Generation
abstract
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, Nanyun Peng. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Emily Sheng, Kai-Wei Chang 0001, Premkumar Natarajan, Nanyun Peng 0001
EMNLP/IJCNLP (1)2
2019 Retrofitting Contextualized Word Embeddings with Paraphrases
abstract
Weijia Shi, Muhao Chen, Pei Zhou, Kai-Wei Chang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Muhao Chen 0001, Kai-Wei Chang 0001
EMNLP/IJCNLP (1)4
2019 Learning to Discriminate Perturbations for Blocking Adversarial Attacks in Text Classification
abstract
Yichao Zhou, Jyun-Yu Jiang, Kai-Wei Chang, Wei Wang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yichao Zhou 0001, Jyun-Yu Jiang, Kai-Wei Chang 0001, Wei Wang 0010
EMNLP/IJCNLP (1)3
2019 Examining Gender Bias in Languages with Grammatical Gender
abstract
Pei Zhou, Weijia Shi, Jieyu Zhao, Kuan-Hao Huang, Muhao Chen, Ryan Cotterell, Kai-Wei Chang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Jieyu Zhao 0001, Kuan-Hao Huang, Muhao Chen 0001, Ryan Cotterell, Kai-Wei Chang 0001
EMNLP/IJCNLP (1)7
2019 Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image Representations
abstract
In this work, we present a framework to measure and mitigate intrinsic biases with respect to protected variables -such as gender- in visual recognition tasks. We show that trained models significantly amplify the association of target labels with gender beyond what one would expect from biased datasets. Surprisingly, we show that even when datasets are balanced such that each label co-occurs equally with each gender, learned models amplify the association between labels and gender, as much as if data had not been balanced! To mitigate this, we adopt an adversarial approach to remove unwanted features corresponding to protected variables from intermediate representations in a deep neural network - and provide a detailed analysis of its effectiveness. Experiments on two datasets: the COCO dataset (objects), and the imSitu dataset (actions), show reductions in gender bias amplification while maintaining most of the accuracy of the original models.
Jieyu Zhao 0001, Mark Yatskar, Kai-Wei Chang 0001, Vicente Ordonez
ICCV4
2019 Context Attentive Document Ranking and Query Suggestion
abstract
We present a context-aware neural ranking model to exploit users' on-task search activities and enhance retrieval performance. In particular, a two-level hierarchical recurrent neural network is introduced to learn search context representation of individual queries, search tasks, and corresponding dependency structure by jointly optimizing two companion retrieval tasks: document ranking and query suggestion. To identify variable dependency structure between search context and users' ongoing search activities, attention at both levels of recurrent states are introduced. Extensive experiment comparisons against a rich set of baseline methods and an in-depth ablation analysis confirm the value of our proposed approach for modeling search context buried in search tasks.
Wasi Uddin Ahmad, Kai-Wei Chang 0001, Hongning Wang
SIGIR2
2019 Multifaceted protein-protein interaction prediction based on Siamese residual RCNN
abstract
MOTIVATION: Sequence-based protein-protein interaction (PPI) prediction represents a fundamental computational biology problem. To address this problem, extensive research efforts have been made to extract predefined features from the sequences. Based on these features, statistical algorithms are learned to classify the PPIs. However, such explicit features are usually costly to extract, and typically have limited coverage on the PPI information. RESULTS: We present an end-to-end framework, PIPR (Protein-Protein Interaction Prediction Based on Siamese Residual RCNN), for PPI predictions using only the protein sequences. PIPR incorporates a deep residual recurrent convolutional neural network in the Siamese architecture, which leverages both robust local features and contextualized information, which are significant for capturing the mutual influence of proteins sequences. PIPR relieves the data pre-processing efforts that are required by other systems, and generalizes well to different application scenarios. Experimental evaluations show that PIPR outperforms various state-of-the-art systems on the binary PPI prediction problem. Moreover, it shows a promising performance on more challenging problems of interaction type prediction and binding affinity estimation, where existing approaches fall short. AVAILABILITY AND IMPLEMENTATION: The implementation is available at https://github.com/muhaochen/seq_ppi.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Muhao Chen 0001, Chelsea J.-T. Ju, Xuelu Chen, Kai-Wei Chang 0001, Carlo Zaniolo, Wei Wang 0010
Bioinform.6
2019 Efficient Contextual Representation Learning With Continuous Outputs
abstract
Contextual representation models have achieved great success in improving various downstream natural language processing tasks. However, these language-model-based encoders are difficult to train due to their large parameter size and high computational complexity. By carefully examining the training procedure, we observe that the softmax layer, which predicts a distribution of the target word, often induces significant overhead, especially when the vocabulary size is large. Therefore, we revisit the design of the output layer and consider directly predicting the pre-trained embedding of the target word for a given context. When applied to ELMo, the proposed approach achieves a 4-fold speedup and eliminates 80% trainable parameters while achieving competitive performance on downstream tasks. Further analysis shows that the approach maintains the speed advantage under various settings, even when the sentence encoder is scaled up.
Liunian Harold Li, Patrick H. Chen, Cho-Jui Hsieh, Kai-Wei Chang 0001
Trans. Assoc. Comput. Linguistics4
2018 Building Language Models for Text with Named Entities
abstract
Text in many domains involves a significant amount of named entities.Predicting the entity names is often challenging for a language model as they appear less frequent on the training corpus.In this paper, we propose a novel and effective approach to building a discriminative language model which can learn the entity names by leveraging their entity type information.We also introduce two benchmark datasets based on recipes and Java programming codes, on which we evaluate the proposed model.Experimental results show that our model achieves 52.2% better perplexity in recipe generation and 22.06% on code generation than the stateof-the-art language models.
Md. Rizwan Parvez, Saikat Chakraborty 0001, Baishakhi Ray, Kai-Wei Chang 0001
ACL (1)4
2018 Generating Natural Language Adversarial Examples
abstract
Deep neural networks (DNNs) are vulnerable to adversarial examples, perturbations to correctly classified examples which can cause the model to misclassify.In the image domain, these perturbations are often virtually indistinguishable to human perception, causing humans and state-of-the-art models to disagree.However, in the natural language domain, small perturbations are clearly perceptible, and the replacement of a single word can drastically alter the semantics of the document.Given these challenges, we use a black-box population-based optimization algorithm to generate semantically and syntactically similar adversarial examples that fool well-trained sentiment analysis and textual entailment models with success rates of 97% and 70%, respectively.We additionally demonstrate that 92.3% of the successful sentiment analysis adversarial examples are classified to their original label by 20 human annotators, and that the examples are perceptibly quite similar.Finally, we discuss an attempt to use adversarial training as a defense, but fail to yield improvement, demonstrating the strength and diversity of our adversarial examples.We hope our findings encourage researchers to pursue improving the robustness of DNNs in the natural language domain.
Moustafa Farid Alzantot, Yash Sharma 0001, Ahmed Elgohary, Bo-Jhang Ho, Mani Srivastava 0001, Kai-Wei Chang 0001
EMNLP6
2018 Learning Gender-Neutral Word Embeddings
abstract
Word embedding models have become a fundamental component in a wide range of Natural Language Processing (NLP) applications.However, embeddings trained on human-generated corpora have been demonstrated to inherit strong gender stereotypes that reflect social constructs.To address this concern, in this paper, we propose a novel training procedure for learning gender-neutral word embeddings.Our approach aims to preserve gender information in certain dimensions of word vectors while compelling other dimensions to be free of gender influence.Based on the proposed method, we generate a Gender-Neutral variant of GloVe (GN-GloVe).Quantitative and qualitative experiments demonstrate that GN-GloVe successfully isolates gender information without sacrificing the functionality of the embedding model.
Jieyu Zhao 0001, Yichao Zhou 0001, Zeyu Li 0001, Wei Wang 0010, Kai-Wei Chang 0001
EMNLP5
2018 Multi-Task Learning for Document Ranking and Query Suggestion
Wasi Uddin Ahmad, Kai-Wei Chang 0001, Hongning Wang
ICLR (Poster)2
2018 Counterexamples for Robotic Planning Explained in Structured Language
abstract
Automated techniques such as model checking have been used to verify models of robotic mission plans based on Markov decision processes (MDPs) and generate counterexamples that may help diagnose requirement violations. However, such artifacts may be too complex for humans to understand, because existing representations of counterexamples typically include a large number of paths or a complex automaton. To help improve the interpretability of counterexamples, we define a notion of explainable counterexample, which includes a set of structured natural language sentences to describe the robotic behavior that lead to a requirement violation in an MDP model of robotic mission plan. We propose an approach based on mixed-integer linear programming for generating explainable counterexamples that are minimal, sound and complete. We demonstrate the usefulness of the proposed approach via a case study of warehouse robots planning.
Lu Feng 0001, Mahsa Ghasemi, Kai-Wei Chang 0001, Ufuk Topcu
ICRA3
2018 Co-training Embeddings of Knowledge Graphs and Entity Descriptions for Cross-lingual Entity Alignment
abstract
Multilingual knowledge graph (KG) embeddings provide latent semantic representations of entities and structured knowledge with cross-lingual inferences, which benefit various knowledge-driven cross-lingual NLP tasks. However, precisely learning such cross-lingual inferences is usually hindered by the low coverage of entity alignment in many KGs. Since many multilingual KGs also provide literal descriptions of entities, in this paper, we introduce an embedding-based approach which leverages a weakly aligned multilingual KG for semi-supervised cross-lingual learning using entity descriptions. Our approach performs co-training of two embedding models, i.e. a multilingual KG embedding model and a multilingual literal description embedding model. The models are trained on a large Wikipedia-based trilingual dataset where most entity alignment is unknown to training. Experimental results show that the performance of the proposed approach on the entity alignment task improves at each iteration of co-training, and eventually reaches a stage at which it significantly surpasses previous approaches. We also show that our approach has promising abilities for zero-shot entity alignment, and cross-lingual KG completion.
Muhao Chen 0001, Yingtao Tian, Kai-Wei Chang 0001, Steven Skiena, Carlo Zaniolo
IJCAI3
2018 A Corpus to Learn Refer-to-as Relations for Nominals
Wasi Uddin Ahmad, Kai-Wei Chang 0001
LREC2
2018 A Corpus of Drug Usage Guidelines Annotated with Type of Advice
Sarah Masud Preum, Md. Rizwan Parvez, Kai-Wei Chang 0001, John A. Stankovic
LREC3
2018 Learning Word Embeddings for Low-Resource Languages by PU Learning
abstract
Chao Jiang, Hsiang-Fu Yu, Cho-Jui Hsieh, Kai-Wei Chang. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Hsiang-Fu Yu, Cho-Jui Hsieh, Kai-Wei Chang 0001
NAACL-HLT4
2018 Intent-aware Query Obfuscation for Privacy Protection in Personalized Web Search
abstract
Modern web search engines exploit users' search history to personalize search results, with a goal of improving their service utility on a per-user basis. But it is this very dimension that leads to the risk of privacy infringement and raises serious public concerns. In this work, we propose a client-centered intent-aware query obfuscation solution for protecting user privacy in a personalized web search scenario. In our solution, each user query is submitted with l additional cover queries and corresponding clicks, which act as decoys to mask users' genuine search intent from a search engine. The cover queries are sequentially sampled from a set of hierarchically organized language models to ensure the coherency of fake search intents in a cover search task. Our approach emphasizes the plausibility of generated cover queries, not only to the current genuine query but also to previous queries in the same task, to increase the complexity for a search engine to identify a user's true intent. We also develop two new metrics from an information theoretic perspective to evaluate the effectiveness of provided privacy protection. Comprehensive experiment comparisons with state-of-the-art query obfuscation techniques are performed on the public AOL search log, and the propitious results substantiate the effectiveness of our solution.
Wasi Uddin Ahmad, Kai-Wei Chang 0001, Hongning Wang
SIGIR2
2017 Resource Constrained Structured Prediction
Tolga Bolukbasi, Kai-Wei Chang 0001, Joseph Wang 0001, Venkatesh Saligrama
AAAI2
2017 Men Also Like Shopping: Reducing Gender Bias Amplification using Corpus-level Constraints
abstract
Language is increasingly being used to define rich visual recognition problems with supporting image collections sourced from the web.Structured prediction models are used in these tasks to take advantage of correlations between co-occurring labels and visual input but risk inadvertently encoding social biases found in web corpora.In this work, we study data and models associated with multilabel object classification and visual semantic role labeling.We find that (a) datasets for these tasks contain significant gender bias and (b) models trained on these datasets further amplify existing bias.For example, the activity cooking is over 33% more likely to involve females than males in a training set, and a trained model further amplifies the disparity to 68% at test time.We propose to inject corpus-level constraints for calibrating existing structured prediction models and design an algorithm based on Lagrangian relaxation for collective inference.Our method results in almost no performance loss for the underlying recognition task but decreases the magnitude of bias amplification by 47.5% and 40.5% for multilabel classification and visual semantic role labeling, respectively.
Jieyu Zhao 0001, Mark Yatskar, Vicente Ordonez, Kai-Wei Chang 0001
EMNLP5
2016 Learning from Explicit and Implicit Supervision Jointly For Algebra Word Problems
abstract
Automatically solving algebra word problems has raised considerable interest recently.Existing state-of-the-art approaches mainly rely on learning from human annotated equations.In this paper, we demonstrate that it is possible to efficiently mine algebra problems and their numerical solutions with little to no manual effort.To leverage the mined dataset, we propose a novel structured-output learning algorithm that aims to learn from both explicit (e.g., equations) and implicit (e.g., solutions) supervision signals jointly.Enabled by this new algorithm, our model gains 4.6% absolute improvement in accuracy on the ALG-514 benchmark compared to the one without using implicit supervision.The final model also outperforms the current state-of-the-art approach by 3%.
Shyam Upadhyay, Ming-Wei Chang, Kai-Wei Chang 0001, Scott Yih
EMNLP3
2016 Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings
abstract
The blind application of machine learning runs the risk of amplifying biases present in data. Such a danger is facing us with word embedding, a popular framework to represent text data as vectors which has been used in many machine learning and natural language processing tasks. We show that even word embeddings trained on Google News articles exhibit female/male gender stereotypes to a disturbing extent. This raises concerns because their widespread use, as we describe, often tends to amplify these biases. Geometrically, gender bias is first shown to be captured by a direction in the word embedding. Second, gender neutral words are shown to be linearly separable from gender definition words in the word embedding. Using these properties, we provide a methodology for modifying an embedding to remove gender stereotypes, such as the association between the words receptionist and female, while maintaining desired associations such as between the words queen and female. Using crowd-worker evaluation as well as standard benchmarks, we empirically demonstrate that our algorithms significantly reduce gender bias in embeddings while preserving the its useful properties such as the ability to cluster related concepts and to solve analogy tasks. The resulting embeddings can be used in applications without amplifying gender bias.
Tolga Bolukbasi, Kai-Wei Chang 0001, James Zou 0001, Venkatesh Saligrama, Adam Tauman Kalai
NIPS2
2016 A Credit Assignment Compiler for Joint Prediction
abstract
Many machine learning applications involve jointly predicting multiple mutually dependent output variables. Learning to search is a family of methods where the complex decision problem is cast into a sequence of decisions via a search space. Although these methods have shown promise both in theory and in practice, implementing them has been burdensomely awkward. In this paper, we show the search space can be defined by an arbitrary imperative program, turning learning to search into a credit assignment compiler. Altogether with the algorithmic improvements for the compiler, we radically reduce the complexity of programming and the running time. We demonstrate the feasibility of our approach on multiple joint prediction tasks. In all cases, we obtain accuracies as high as alternative approaches, at drastically reduced execution and programming time.
Kai-Wei Chang 0001, He He 0001, Stéphane Ross, Hal Daumé III, John Langford 0001
NIPS1
2015 Structural Learning with Amortized Inference
abstract
Training a structured prediction model involves performing several loss-augmented inference steps. Over the lifetime of the training, many of these inference problems, although different, share the same solution. We propose AI-DCD, an Amortized Inference framework for Dual Coordinate Descent method, an approximate learning algorithm, that accelerates the training process by exploiting this redundancy of solutions, without compromising the performance of the model. We show the efficacy of our method by training a structured SVM using dual coordinate descent for an entityrelation extraction task. Our method learns the same model as an exact training algorithm would, but call the inference engine only in 10% – 24% of the inference problems encountered during training. We observe similar gains on a multi-label classification task and with a Structured Perceptron model for the entity-relation task.
Kai-Wei Chang 0001, Shyam Upadhyay, Gourab Kundu, Dan Roth 0001
AAAI1
2015 A Joint Framework for Coreference Resolution and Mention Head Detection
abstract
In coreference resolution, a fair amount of research treats mention detection as a preprocessed step and focuses on developing algorithms for clustering coreferred mentions. However, there are significant gaps between the performance on gold mentions and the performance on the real problem, when mentions are predicted from raw text via an imperfect Mention Detection (MD) module. Motivated by the goal of reducing such gaps, we develop an ILP-based joint coreference resolution and mention head formulation that is shown to yield significant improvements on coreference from raw text, outperforming existing state-ofart systems on both the ACE-2004 and the CoNLL-2012 datasets. At the same time, our joint approach is shown to improve mention detection by close to 15% F1. One key insight underlying our approach is that identifying and co-referring mention heads is not only sufficient but is more robust than working with complete mentions.
Haoruo Peng, Kai-Wei Chang 0001, Dan Roth 0001
CoNLL2
2015 Learning to Search Better than Your Teacher
abstract
Methods for learning to search for structured prediction typically imitate a reference policy, with existing theoretical guarantees demonstrating low regret compared to that reference. This is unsatisfactory in many applications where the reference policy is suboptimal and the goal of learning is to improve upon it. Can learning to search work even when the reference is poor? We provide a new learning to search algorithm, LOLS, which does well relative to the reference policy, but additionally guarantees low regret compared to deviations from the learned policy: a local-optimality guarantee. Consequently, LOLS can improve upon the reference policy, unlike previous algorithms. This enables us to develop structured contextual bandits, a partial information structured prediction setting with many potential applications.
Kai-Wei Chang 0001, Akshay Krishnamurthy, Alekh Agarwal, Hal Daumé III, John Langford 0001
ICML1
2015 Hands-on Learning to Search for Structured Prediction
abstract
Hal Daumé III, John Langford, Kai-Wei Chang, He He, Sudha Rao. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Tutorial Abstracts. 2015.
Hal Daumé III, John Langford 0001, Kai-Wei Chang 0001, He He 0001, Sudha Rao
HLT-NAACL3
2014 Typed Tensor Decomposition of Knowledge Bases for Relation Extraction
abstract
While relation extraction has traditionally been viewed as a task relying solely on textual data, recent work has shown that by taking as input existing facts in the form of entity-relation triples from both knowl-edge bases and textual data, the perfor-mance of relation extraction can be im-proved significantly. Following this new paradigm, we propose a tensor decompo-sition approach for knowledge base em-bedding that is highly scalable, and is es-pecially suitable for relation extraction. By leveraging relational domain knowl-edge about entity type information, our learning algorithm is significantly faster than previous approaches and is better able to discover new relations missing from the database. In addition, when ap-plied to a relation extraction task, our ap-proach alone is comparable to several ex-isting systems, and improves the weighted mean average precision of a state-of-the-art method by 10 points when used as a subcomponent. 1
Kai-Wei Chang 0001, Scott Yih, Bishan Yang, Christopher Meek
EMNLP1
2014 A Discriminative Latent Variable Model for Online Clustering
abstract
This paper presents a latent variable structured prediction model for discriminative supervised clustering of items called the Latent Left-linking Model (L3M). We present an online clustering algorithm for L3M based on a feature-based item similarity function. We provide a learning framework for estimating the similarity function and present a fast stochastic gradient-based learning technique. In our experiments on coreference resolution and document clustering, L3 M outperforms several existing online as well as batch supervised clustering techniques.
Rajhans Samdani, Kai-Wei Chang 0001, Dan Roth 0001
ICML2
2013 A Constrained Latent Variable Model for Coreference Resolution
abstract
Coreference resolution is a well known clustering task in Natural Language Processing.In this paper, we describe the Latent Left Linking model (L 3 M), a novel, principled, and linguistically motivated latent structured prediction approach to coreference resolution.We show that L 3 M admits efficient inference and can be augmented with knowledge-based constraints; we also present a fast stochastic gradient based learning.Experiments on ACE and Ontonotes data show that L 3 M and its constrained version, CL 3 M, are more accurate than several state-of-the-art approaches as well as some structured prediction models proposed in the literature.
Kai-Wei Chang 0001, Rajhans Samdani, Dan Roth 0001
EMNLP1
2013 Multi-Relational Latent Semantic Analysis
abstract
We present Multi-Relational Latent Semantic Analysis (MRLSA) which generalizes Latent Semantic Analysis (LSA).MRLSA provides an elegant approach to combining multiple relations between words by constructing a 3-way tensor.Similar to LSA, a lowrank approximation of the tensor is derived using a tensor decomposition.Each word in the vocabulary is thus represented by a vector in the latent semantic space and each relation is captured by a latent square matrix.The degree of two words having a specific relation can then be measured through simple linear algebraic operations.We demonstrate that by integrating multiple relations from both homogeneous and heterogeneous information sources, MRLSA achieves stateof-the-art performance on existing benchmark datasets for two relations, antonymy and is-a.
Kai-Wei Chang 0001, Scott Yih, Christopher Meek
EMNLP1
2013 Tractable Semi-supervised Learning of Complex Structured Prediction Models
Kai-Wei Chang 0001, S. Sundararajan, S. Sathiya Keerthi
ECML/PKDD (3)1
2013 Multi-core Structural SVM Training
Kai-Wei Chang 0001, Vivek Srikumar, Dan Roth 0001
ECML/PKDD (2)1
2012 Efficient Pattern-Based Time Series Classification on GPU
abstract
Time series shapelet discovery algorithm finds subsequences from a set of time series for use as primitives for time series classification. This algorithm has drawn a lot of interest because of the interpretability of its results. However, computation requirements restrict the algorithm from dealing with large data sets and may limit its application in many domains. In this paper, we address this issue by redesigning the algorithm for implementation on highly parallel Graphics Process Units (GPUs). We investigate several concepts of GPU programming and propose a dynamic programming algorithm that is suitable for implementation on GPUs. Results show that the proposed GPU implementation significantly reduces the running time of the shapelet discovery algorithm. For example, on the largest sample dataset from the original authors, the running time is reduced from half a day to two minutes.
Kai-Wei Chang 0001, Biplab Deka, Wen-Mei W. Hwu, Dan Roth 0001
ICDM1
2012 Large Linear Classification When Data Cannot Fit in Memory
abstract
Recent advances in linear classification have shown that for applications such as document classification, the training process can be extremely efficient. However, most of the existing training methods are designed by assuming that data can be stored in the computer memory. These methods cannot be easily applied to data larger than the memory capacity due to the random access to the disk. We propose and analyze a block minimization framework for data larger than the memory size. At each step a block of data is loaded from the disk and handled by certain learning methods. We investigate two implementations of the proposed framework for primal and dual SVMs, respectively. Because data cannot fit in memory, many design considerations are very different from those for traditional algorithms. We discuss and compare with existing approaches that are able to handle data larger than memory. Experiments using data sets 20 times larger than the memory demonstrate the effectiveness of the proposed method.
Hsiang-Fu Yu, Cho-Jui Hsieh, Kai-Wei Chang 0001, Chih-Jen Lin
ACM Trans. Knowl. Discov. Data3
2011 Large Linear Classification When Data Cannot Fit in Memory
abstract
Linear classification is a useful tool for dealing with large-scale data in applications such as document classification and natural language processing. Recent developments of linear classification have shown that the training process can be efficiently conducted. However, when the data size exceeds the memory capacity, most training methods suffer from very slow convergence due to the severe disk swapping. Although some methods have attempted to handle such a situation, they are usually too complicated to support some important functions such as parameter selection. In this paper, we introduce a block minimization framework for data larger than memory. Under the framework, a solver splits data into blocks and stores them into separate files. Then, at each time, the solver trains a data block loaded from disk. Although the framework is simple, the experimental results show that it effectively handles a data set 20 times larger than the memory capacity.
Hsiang-Fu Yu, Cho-Jui Hsieh, Kai-Wei Chang 0001, Chih-Jen Lin
IJCAI3
2011 Selective block minimization for faster convergence of limited memory large-scale linear models
abstract
As the size of data sets used to build classifiers steadily increases, training a linear model efficiently with limited memory becomes essential. Several techniques deal with this problem by loading blocks of data from disk one at a time, but usually take a considerable number of iterations to converge to a reasonable model. Even the best block minimization techniques [1] require many block loads since they treat all training examples uniformly. As disk I/O is expensive, reducing the amount of disk access can dramatically decrease the training time. This paper introduces a selective block minimization (SBM) algorithm, a block minimization method that makes use of selective sampling. At each step, SBM updates the model using data consisting of two parts: (1) new data loaded from disk and (2) a set of informative samples already in memory from previous steps. We prove that, by updating the linear model in the dual form, the proposed method fully utilizes the data in memory and converges to a globally optimal solution on the entire data. Experiments show that the SBM algorithm dramatically reduces the number of blocks loaded from disk and consequently obtains an accurate and stable model quickly on both binary and multi-class classification.
Kai-Wei Chang 0001, Dan Roth 0001
KDD1
2010 Large linear classification when data cannot fit in memory
abstract
Recent advances in linear classification have shown that for applications such as document classification, the training can be extremely efficient. However, most of the existing training methods are designed by assuming that data can be stored in the computer memory. These methods cannot be easily applied to data larger than the memory capacity due to the random access to the disk. We propose and analyze a block minimization framework for data larger than the memory size. At each step a block of data is loaded from the disk and handled by certain learning methods. We investigate two implementations of the proposed framework for primal and dual SVMs, respectively. As data cannot fit in memory, many design considerations are very different from those for traditional algorithms. Experiments using data sets 20 times larger than the memory demonstrate the effectiveness of the proposed method.
Hsiang-Fu Yu, Cho-Jui Hsieh, Kai-Wei Chang 0001, Chih-Jen Lin
KDD3
2010 Training and Testing Low-degree Polynomial Data Mappings via Linear SVM
Yin-Wen Chang, Cho-Jui Hsieh, Kai-Wei Chang 0001, Michael Ringgaard, Chih-Jen Lin
J. Mach. Learn. Res.3
2010 Iterative Scaling and Coordinate Descent Methods for Maximum Entropy Models
Fang-Lan Huang, Cho-Jui Hsieh, Kai-Wei Chang 0001, Chih-Jen Lin
J. Mach. Learn. Res.3
2010 A Comparison of Optimization Methods and Software for Large-scale L1-regularized Linear Classification
Guo-Xun Yuan, Kai-Wei Chang 0001, Cho-Jui Hsieh, Chih-Jen Lin
J. Mach. Learn. Res.2
2008 A dual coordinate descent method for large-scale linear SVM
abstract
In many applications, data appear with a huge number of instances as well as features. Linear Support Vector Machines (SVM) is one of the most popular tools to deal with such large-scale sparse data. This paper presents a novel dual coordinate descent method for linear SVM with L1-and L2-loss functions. The proposed method is simple and reaches an ε-accurate solution in O(log(1/ε)) iterations. Experiments indicate that our method is much faster than state of the art solvers such as Pegasos, TRON, SVMperf, and a recent primal coordinate descent implementation.
Cho-Jui Hsieh, Kai-Wei Chang 0001, Chih-Jen Lin, S. Sathiya Keerthi, S. Sundararajan
ICML2
2008 A sequential dual method for large scale multi-class linear svms
abstract
Efficient training of direct multi-class formulations of linear Support Vector Machines is very useful in applications such as text classification with a huge number examples as well as features. This paper presents a fast dual method for this training. The main idea is to sequentially traverse through the training set and optimize the dual variables associated with one example at a time. The speed of training is enhanced further by shrinking and cooling heuristics. Experiments indicate that our method is much faster than state of the art solvers such as bundle, cutting plane and exponentiated gradient methods.
S. Sathiya Keerthi, S. Sundararajan, Kai-Wei Chang 0001, Cho-Jui Hsieh, Chih-Jen Lin
KDD3
2008 Coordinate Descent Method for Large-scale L2-loss Linear Support Vector Machines
Kai-Wei Chang 0001, Cho-Jui Hsieh, Chih-Jen Lin
J. Mach. Learn. Res.1
2008 LIBLINEAR: A Library for Large Linear Classification
Rong-En Fan, Kai-Wei Chang 0001, Cho-Jui Hsieh, Xiang-Rui Wang, Chih-Jen Lin
J. Mach. Learn. Res.2