EDBT 2026 Demo / reviewers in the wild / expert
Bill Y. Lin
dblp:190/4518 · also Bill Yuchen Lin, Yuchen Lin 0001
· DBLP profile ↗
48ranked-venue papers
16as first author
37since 2021 · last 2026
0000-0002-1149-0186ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 47 · 16 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Temporal Sampling for Forgotten Reasoning in LLMsabstractYuetai Li, Zhangchen Xu, Fengqing Jiang, Bhaskar Ramasubramanian, Luyao Niu, Bill Yuchen Lin, Xiang Yue, Radha Poovendran. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuetai Li, Zhangchen Xu, Fengqing Jiang, Bhaskar Ramasubramanian, Luyao Niu, Bill Y. Lin, Xiang Yue, Radha Poovendran |
ACL (1) | 6 |
| 2025 | ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat TemplatesabstractLarge language models (LLMs) are expected to follow instructions from users and engage in conversations. Techniques to enhance LLMs' instruction-following capabilities typically fine-tune them using data structured according to a predefined chat template. Although chat templates are shown to be effective in optimizing LLM performance, their impact on safety alignment of LLMs has been less understood, which is crucial for deploying LLMs safely at scale. In this paper, we investigate how chat templates affect safety alignment of LLMs. We identify a common vulnerability, named ChatBug, that is introduced by chat templates. Our key insight to identify ChatBug is that the chat templates provide a rigid format that need to be followed by LLMs, but not by users. Hence, a malicious user may not necessarily follow the chat template when prompting LLMs. Instead, malicious users could leverage their knowledge of the chat template and accordingly craft their prompts to bypass safety alignments of LLMs. We study two attacks to exploit the ChatBug vulnerability. Additionally, we demonstrate that the success of multiple existing attacks can be attributed to the ChatBug vulnerability. We show that a malicious user can exploit the ChatBug vulnerability of eight state-of-the-art (SOTA) LLMs and effectively elicit unintended responses from these models. Moreover, we show that ChatBug can be exploited by existing jailbreak attacks to enhance their attack success rates. We investigate potential countermeasures to ChatBug. Our results show that while adversarial training effectively mitigates the ChatBug vulnerability, the victim model incurs significant performance degradation. These results highlight the trade-off between safety alignment and helpfulness. Developing new methods for instruction tuning to balance this trade-off is an open and critical direction for future research. Fengqing Jiang, Zhangchen Xu, Luyao Niu, Bill Y. Lin, Radha Poovendran |
AAAI | 4 |
| 2025 | CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-TeamingabstractYu Ying Chiu, Liwei Jiang, Bill Yuchen Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, Yejin Choi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yu Ying Chiu, Bill Y. Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, Yejin Choi 0001 |
ACL (1) | 3 |
| 2025 | VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward ModelsabstractVision-language generative reward models (VL-GenRMs) play a crucial role in aligning and evaluating multimodal AI systems, yet their own evaluation remains under-explored. Current assessment methods primarily rely on AI-annotated preference labels from traditional VL tasks, which can introduce biases and often fail to effectively challenge state-of-the-art models. To address these limitations, we introduce VL-RewardBench, a comprehensive benchmark spanning general multimodal queries, visual hallucination detection, and complex reasoning tasks. Through our AI-assisted annotation pipeline that combines sample selection with human verification, we curate 1,250 high-quality examples specifically designed to probe VL-GenRMs limitations. Comprehensive evaluation across 16 leading large vision-language models demonstrates VL-RewardBench’s effectiveness as a challenging testbed, where even GPT-4o achieves only 65.4% accuracy, and state-of-the-art open-source models such as Qwen2-VL-72B, struggle to surpass random-guessing. Importantly, performance on VL-RewardBench strongly correlates (Pearson’s r > 0.9) with MMMU-Pro accuracy using Best-of-N sampling with VL-GenRMs. Analysis experiments uncover three critical insights for improving VL-GenRMs: (i) models predominantly fail at basic visual perception tasks rather than reasoning tasks; (ii) inference-time scaling benefits vary dramatically by model capacity; and (iii) training VL-GenRMs to learn to judge substantially boosts judgment capability (+14.7% accuracy for a 7B VL-GenRM). We believe VL-RewardBench along with the experimental insights will become a valuable resource for advancing VL-GenRMs. Project page: https://vl-rewardbench.github.io. Lei Li 0039, Yuancheng Wei, Zhihui Xie 0002, Xuqing Yang, Yifan Song 0002, Peiyi Wang, Chenxin An, Tianyu Liu 0001, Sujian Li, Bill Y. Lin, Lingpeng Kong, Qi Liu 0049 |
CVPR | 10 |
| 2025 | WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the WildabstractWe introduce WildBench, an automated evaluation framework designed to benchmark large language models (LLMs) using challenging, real-world user queries. WildBench consists of 1,024 tasks carefully selected from over one million human-chatbot conversation logs. For automated evaluation with WildBench, we have developed two metrics, WB-Reward and WB-Score, which are computable using advanced LLMs such as GPT-4-turbo. WildBench evaluation uses task-specific checklists to evaluate model outputs systematically and provides structured explanations that justify the scores and comparisons, resulting in more reliable and interpretable automatic judgments. WB-Reward employs fine-grained pairwise comparisons between model responses, generating five potential outcomes: much better, slightly better, slightly worse, much worse, or a tie. Unlike previous evaluations that employed a single baseline model, we selected three baseline models at varying performance levels to ensure a comprehensive pairwise evaluation. Additionally, we propose a simple method to mitigate length bias, by converting outcomes of “slightly better/worse” to “tie” if the winner response exceeds the loser one by more than K characters. WB-Score evaluates the quality of model outputs individually, making it a fast and cost-efficient evaluation metric. WildBench results demonstrate a strong correlation with the human-voted Elo ratings from Chatbot Arena on hard tasks. Specifically, WB-Reward achieves a Pearson correlation of 0.98 with top-ranking models. Additionally, WB-Score reaches 0.95, surpassing both ArenaHard’s 0.91 and AlpacaEval2.0’s 0.89 for length-controlled win rates, as well as the 0.87 for regular win rates. Bill Y. Lin, Yuntian Deng, Khyathi Raghavi Chandu, Abhilasha Ravichander, Valentina Pyatkin, Nouha Dziri, Ronan Le Bras 0001, Yejin Choi 0001 |
ICLR | 1 |
| 2025 | Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with NothingabstractHigh-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent existing open-source data creation methods from scaling effectively, potentially limiting the diversity and quality of public alignment datasets. Is it possible to synthesize high-quality instruction data at scale by extracting it directly from an aligned LLM? We present a self-synthesis method for generating large-scale alignment data named Magpie. Our key observation is that aligned LLMs like Llama-3-Instruct can generate a user query when we input only the pre-query templates up to the position reserved for user messages, thanks to their auto-regressive nature. We use this method to prompt Llama-3-Instruct and generate 4 million instructions along with their corresponding responses. We further introduce extensions of Magpie for filtering, generating multi-turn, preference optimization, domain-specific and multilingual datasets. We perform a comprehensive analysis of the Magpie-generated data. To compare Magpie-generated data with other public instruction datasets (e.g., ShareGPT, WildChat, Evol-Instruct, UltraChat, OpenHermes, Tulu-V2-Mix, GenQA), we fine-tune Llama-3-8B-Base with each dataset and evaluate the performance of the fine-tuned models. Our results indicate that using Magpie for supervised fine-tuning (SFT) solely can surpass the performance of previous public datasets utilized for both SFT and preference optimization, such as direct preference optimization with UltraFeedback. We also show that in some tasks, models supervised fine-tuned with Magpie perform comparably to the official Llama-3-8B-Instruct, despite the latter being enhanced with 10 million data points through SFT and subsequent preference optimization. This advantage is evident on alignment benchmarks such as AlpacaEval, ArenaHard, and WildBench. Zhangchen Xu, Fengqing Jiang, Luyao Niu, Yuntian Deng, Radha Poovendran, Yejin Choi 0001, Bill Y. Lin |
ICLR | 7 |
| 2025 | Latent Action Pretraining from VideosabstractWe introduce Latent Action Pretraining for general Action models (LAPA), the first unsupervised method for pretraining Vision-Language-Action (VLA) models without ground-truth robot action labels. Existing Vision-Language-Action models require action labels typically collected by human teleoperators during pretraining, which significantly limits possible data sources and scale. In this work, we propose a method to learn from internet-scale videos that do not have robot action labels. We first train an action quantization model leveraging VQ-VAE-based objective to learn discrete latent actions between image frames, then pretrain a latent VLA model to predict these latent actions from observations and task descriptions, and finally finetune the VLA on small-scale robot manipulation data to map from latent to robot actions. Experimental results demonstrate that our method significantly outperforms existing techniques that train robot manipulation policies from large-scale videos. Furthermore, it outperforms the state-of-the-art VLA model trained with robotic action labels on real-world manipulation tasks that require language conditioning, generalization to unseen objects, and semantic generalization to unseen instructions. Training only on human manipulation videos also shows positive transfer, opening up the potential for leveraging web-scale data for robotics foundation models. Seonghyeon Ye, Joel Jang, Byeongguk Jeon, Se June Joo, Baolin Peng, Ajay Mandlekar, Reuben Tan, Yu-Wei Chao, Bill Y. Lin, Lars Liden, Kimin Lee, Jianfeng Gao 0001, Luke Zettlemoyer, Dieter Fox, Minjoon Seo |
ICLR | 10 |
| 2025 | ZebraLogic: On the Scaling Limits of LLMs for Logical ReasoningabstractWe investigate the logical reasoning capabilities of Large Language Models (LLMs) and their scalability across complex deductive tasks. Using ZebraLogic, a newly developed benchmark dataset of logic grid puzzles derived from constraint satisfaction problems (CSPs), we systematically evaluate LLM performance. ZebraLogic spans a broad range of search space complexities and incorporates diverse logical constraints, providing a controlled environment to assess reasoning abilities. Our results reveal a significant decline in accuracy as problem complexity increases—a phenomenon we term the “curse of complexity.” Notably, this limitation persists even with scaling model size and inference-time computation, suggesting fundamental constraints in current LLM reasoning capabilities. Additionally, we explore strategies such as Best-of-N sampling, backtracking mechanisms, and self-verification prompts to enhance logical reasoning performance. Our findings provide critical insights into the scaling behavior of LLMs, highlight their limitations, and outline potential directions for advancing their reasoning capabilities. Bill Y. Lin, Ronan Le Bras 0001, Kyle Richardson 0001, Ashish Sabharwal, Radha Poovendran, Peter Clark, Yejin Choi 0001 |
ICML | 1 |
| 2025 | The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language ModelsabstractSeungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Cho, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Yuchen Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, Minjoon Seo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Choi 0001, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Y. Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee 0002, Minjoon Seo |
NAACL (Long Papers) | 27 |
| 2025 | Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language ModelsabstractAbhilasha Ravichander, Jillian Fisher, Taylor Sorensen, Ximing Lu, Maria Antoniak, Bill Yuchen Lin, Niloofar Mireshghallah, Chandra Bhagavatula, Yejin Choi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Abhilasha Ravichander, Jillian Fisher, Taylor Sorensen, Ximing Lu, Maria Antoniak, Bill Y. Lin, Niloofar Mireshghallah, Chandra Bhagavatula, Yejin Choi 0001 |
NAACL (Long Papers) | 6 |
| 2025 | The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-DeterminismabstractYifan Song, Guoyin Wang, Sujian Li, Bill Yuchen Lin. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yifan Song 0002, Sujian Li, Bill Y. Lin |
NAACL (Long Papers) | 4 |
| 2025 | Stronger Models are Not Always Stronger Teachers for Instruction TuningabstractZhangchen Xu, Fengqing Jiang, Luyao Niu, Bill Yuchen Lin, Radha Poovendran. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zhangchen Xu, Fengqing Jiang, Luyao Niu, Bill Y. Lin, Radha Poovendran |
NAACL (Long Papers) | 4 |
| 2024 | Trial and Error: Exploration-Based Trajectory Optimization of LLM AgentsabstractLarge Language Models (LLMs) have become integral components in various autonomous agent systems.In this study, we present an exploration-based trajectory optimization approach, referred to as ETO.This learning method is designed to enhance the performance of open LLM agents.Contrary to previous studies that exclusively train on successful expert trajectories, our method allows agents to learn from their exploration failures.This leads to improved performance through an iterative optimization framework.During the exploration phase, the agent interacts with the environment while completing given tasks, gathering failure trajectories to create contrastive trajectory pairs.In the subsequent training phase, the agent utilizes these trajectory preference pairs to update its policy using contrastive learning methods like DPO (Rafailov et al., 2023).This iterative cycle of exploration and training fosters continued improvement in the agents.Our experiments on three complex tasks demonstrate that ETO consistently surpasses baseline performance by a large margin.Furthermore, an examination of task-solving efficiency and potential in scenarios lacking expert trajectory underscores the effectiveness of our approach.1 Yifan Song 0002, Da Yin, Xiang Yue, Sujian Li, Bill Y. Lin |
ACL (1) | 6 |
| 2024 | SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware DecodingabstractZhangchen Xu, Fengqing Jiang, Luyao Niu, Jinyuan Jia, Bill Yuchen Lin, Radha Poovendran. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhangchen Xu, Fengqing Jiang, Luyao Niu, Jinyuan Jia 0001, Bill Y. Lin, Radha Poovendran |
ACL (1) | 5 |
| 2024 | Agent Lumos: Unified and Modular Training for Open-Source Language AgentsabstractDa Yin, Faeze Brahman, Abhilasha Ravichander, Khyathi Chandu, Kai-Wei Chang, Yejin Choi, Bill Yuchen Lin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Da Yin, Faeze Brahman, Abhilasha Ravichander, Khyathi Raghavi Chandu, Kai-Wei Chang 0001, Yejin Choi 0001, Bill Y. Lin |
ACL (1) | 7 |
| 2024 | VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video GenerationabstractXuan He, Dongfu Jiang, Ge Zhang, Max Ku, Achint Soni, Sherman Siu, Haonan Chen, Abhranil Chandra, Ziyan Jiang, Aaran Arulraj, Kai Wang, Quy Duc Do, Yuansheng Ni, Bohan Lyu, Yaswanth Narsupalli, Rongqi Fan, Zhiheng Lyu, Bill Yuchen Lin, Wenhu Chen. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Dongfu Jiang, Ge Zhang 0009, Max Ku, Achint Soni, Sherman Siu, Abhranil Chandra, Ziyan Jiang, Aaran Arulraj, Kai Wang 0068, Quy Duc Do, Yuansheng Ni, Bohan Lyu 0001, Yaswanth Narsupalli, Rongqi Fan, Zhiheng Lyu, Bill Y. Lin, Wenhu Chen |
EMNLP | 18 |
| 2024 | Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language ModelsabstractSeungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, Minjoon Seo. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Y. Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee 0002, Minjoon Seo |
EMNLP | 4 |
| 2024 | The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context LearningabstractAlignment tuning has become the de facto standard practice for enabling base large language models (LLMs) to serve as open-domain AI assistants. The alignment tuning process typically involves instruction learning through supervised fine-tuning (SFT) and preference tuning via reinforcement learning from human feedback (RLHF). A recent study, LIMA (Zhou et al., 2023), shows that using merely 1K examples for SFT can achieve significant alignment performance as well, suggesting that the effect of alignment tuning might be "superficial." This raises questions about how exactly the alignment tuning transforms a base LLM.
We analyze the effect of alignment tuning by examining the token distribution shift between base LLMs and their aligned counterparts (e.g., Llama-2 and Llama-2-chat). Our findings reveal that base LLMs and their alignment-tuned versions perform nearly identically in decoding on the majority of token positions (i.e., they share the top-ranked tokens). Most distribution shifts occur with stylistic tokens (e.g., discourse markers, safety disclaimers). This direct evidence strongly supports the hypothesis that alignment tuning primarily learns to adopt the language style of AI assistants, and that the knowledge required for answering user queries predominantly comes from the base LLMs themselves.
Based on these findings, we rethink the alignment of LLMs by posing the research question: how effectively can we align base LLMs without SFT or RLHF? To address this, we introduce a simple, tuning-free alignment method, URIAL (Untuned LLMs with Restyled In-context Alignment). URIAL achieves effective alignment purely through in-context learning (ICL) with base LLMs, requiring as few as three constant stylistic examples and a system prompt. We conduct a fine-grained and interpretable evaluation on a diverse set of examples, named just-eval-instruct. Results demonstrate that base LLMs with URIAL can match or even surpass the performance of LLMs aligned with SFT (Mistral-7b-Instruct) or SFT+RLHF (Llama-2-70b-chat). We show that the gap between tuning-free and tuning-based alignment methods can be significantly reduced through strategic prompting and ICL. Our findings on the superficial nature of alignment tuning and results with URIAL suggest that deeper analysis and theoretical understanding of alignment is crucial to future LLM research. Bill Y. Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Raghavi Chandu, Chandra Bhagavatula, Yejin Choi 0001 |
ICLR | 1 |
| 2024 | WildGuard: Open One-stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMsabstractWe introduce WildGuard---an open, light-weight moderation tool for LLM safety that achieves three goals: (1) identifying malicious intent in user prompts, (2) detecting safety risks of model responses, and (3) determining model refusal rate. Together, WildGuard serves the increasing needs for automatic safety moderation and evaluation of LLM interactions, providing a one-stop tool with enhanced accuracy and broad coverage across 13 risk categories. While existing open moderation tools such as Llama-Guard2 score reasonably well in classifying straightforward model interactions, they lag far behind a prompted GPT-4, especially in identifying adversarial jailbreaks and in evaluating models' refusals, a key measure for evaluating safety behaviors in model responses. To address these challenges, we construct WildGuardMix, a large-scale and carefully balanced multi-task safety moderation dataset with 92K labeled examples that cover vanilla (direct) prompts and adversarial jailbreaks, paired with various refusal and compliance responses. WildGuardMix is a combination of WildGuardTrain, the training data of WildGuard, and WildGuardTest, a high-quality human-annotated moderation test set with 5K labeled items covering broad risk scenarios.Through extensive evaluations on WildGuardTest and ten existing public benchmarks, we show that WildGuard establishes state-of-the-art performance in open-source safety moderation across all the three tasks compared to ten strong existing open-source moderation models (e.g., up to 25.3% improvement on refusal detection). Importantly, WildGuard matches and sometimes exceeds GPT-4 performance (e.g., up to 4.8% improvement on prompt harmfulness identification). WildGuard serves as a highly effective safety moderator in an LLM interface, reducing the success rate of jailbreak attacks from 79.8% to 2.4%. We will make all our data, models and training/evaluation code publicly available under CC BY 4.0 license. Seungju Han 0002, Kavel Rao, Allyson Ettinger, Bill Y. Lin, Nathan Lambert 0001, Yejin Choi 0001, Nouha Dziri |
NeurIPS | 5 |
| 2024 | WildVision: Evaluating Vision-Language Models in the Wild with Human PreferencesabstractRecent breakthroughs in vision-language models (VLMs) emphasize the necessity of benchmarking human preferences in real-world multimodal interactions. To address this gap, we launched WildVision-Arena (WV-Arena), an online platform that collects human preferences to evaluate VLMs. We curated WV-Bench by selecting 500 high-quality samples from 8,000 user submissions in WV-Arena. WV-Bench uses GPT-4 as the judge to compare each VLM with Claude-3-Sonnet, achieving a Spearman correlation of 0.94 with the WV-Arena Elo. This significantly outperforms other benchmarks like MMVet, MMMU, and MMStar.Our comprehensive analysis of 20K real-world interactions reveals important insights into the failure cases of top-performing VLMs. For example, we find that although GPT-4V surpasses many other models like Reka-Flash, Opus, and Yi-VL-Plus in simple visual recognition and reasoning tasks, it still faces challenges with subtle contextual cues, spatial reasoning, visual imagination, and expert domain knowledge. Additionally, current VLMs exhibit issues with hallucinations and safety when intentionally provoked. We are releasing our chat and feedback data to further advance research in the field of VLMs. Dongfu Jiang, Wenhu Chen, William Yang Wang, Yejin Choi 0001, Bill Y. Lin |
NeurIPS | 6 |
| 2023 | On Grounded Planning for Embodied Tasks with Language ModelsabstractLanguage models (LMs) have demonstrated their capability in possessing commonsense knowledge of the physical world, a crucial aspect of performing tasks in everyday life. However, it remains unclear whether they have the capacity to generate grounded, executable plans for embodied tasks. This is a challenging task as LMs lack the ability to perceive the environment through vision and feedback from the physical environment. In this paper, we address this important research question and present the first investigation into the topic. Our novel problem formulation, named G-PlanET, inputs a high-level goal and a data table about objects in a specific environment, and then outputs a step-by-step actionable plan for a robotic agent to follow. To facilitate the study, we establish an evaluation protocol and design a dedicated metric, KAS, to assess the quality of the plans. Our experiments demonstrate that the use of tables for encoding the environment and an iterative decoding strategy can significantly enhance the LMs' ability in grounded planning. Our analysis also reveals interesting and non-trivial findings. Bill Y. Lin, Chengsong Huang, Qian Liu 0033, Wenda Gu, Sam Sommerer, Xiang Ren 0001 |
AAAI | 1 |
| 2023 | LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative FusionabstractWe present LLM-BL E N D E R, an ensembling framework designed to attain consistently superior performance by leveraging the diverse strengths of multiple open-source large language models (LLMs).Our framework consists of two modules: PAIRRANKER and GEN-FUSER, addressing the observation that optimal LLMs for different examples can significantly vary.PAIRRANKER employs a specialized pairwise comparison method to distinguish subtle differences between candidate outputs.It jointly encodes the input text and a pair of candidates, using cross-attention encoders to determine the superior one.Our results demonstrate that PAIRRANKER exhibits the highest correlation with ChatGPT-based ranking.Then, GENFUSER aims to merge the top-ranked candidates, generating an improved output by capitalizing on their strengths and mitigating their weaknesses.To facilitate largescale evaluation, we introduce a benchmark dataset, MixInstruct, which is a mixture of multiple instruction datasets featuring oracle pairwise comparisons.Our LLM-BL E N D E R significantly outperform individual LLMs and baseline methods across various metrics, establishing a substantial performance gap. 1 2 Open Assistant 12.61% Koala 6.71% Alpaca 11.61% Baize 11.61% StableLM 1.90% FLAN-T5 0.80% Vicuna 21.22% Dolly V2 4.50% MOSS 12.91% ChatGLM 8.51% MPT 7.61% Percentage of Examples Where Each Model Ranks First Which LLM should I use for my input?All!I can ensemble! Dongfu Jiang, Xiang Ren 0001, Bill Y. Lin |
ACL (1) | 3 |
| 2023 | AutoTriggER: Label-Efficient and Robust Named Entity Recognition with Auxiliary Trigger ExtractionabstractDong-Ho Lee, Ravi Kiran Selvam, Sheikh Muhammad Sarwar, Bill Yuchen Lin, Fred Morstatter, Jay Pujara, Elizabeth Boschee, James Allan, Xiang Ren. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Ravi Kiran Selvam, Sheikh Muhammad Sarwar, Bill Y. Lin, Fred Morstatter, Jay Pujara, Elizabeth Boschee, James Allan 0001, Xiang Ren 0001 |
EACL | 4 |
| 2023 | Inference-Time Policy Adapters (IPA): Tailoring Extreme-Scale LMs without Fine-tuningabstractXiming Lu, Faeze Brahman, Peter West, Jaehun Jung, Khyathi Chandu, Abhilasha Ravichander, Prithviraj Ammanabrolu, Liwei Jiang, Sahana Ramnath, Nouha Dziri, Jillian Fisher, Bill Lin, Skyler Hallinan, Lianhui Qin, Xiang Ren, Sean Welleck, Yejin Choi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Ximing Lu, Faeze Brahman, Peter West, Jaehun Jung, Khyathi Raghavi Chandu, Abhilasha Ravichander, Prithviraj Ammanabrolu, Sahana Ramnath, Nouha Dziri, Jillian Fisher, Bill Y. Lin, Skyler Hallinan, Lianhui Qin, Xiang Ren 0001, Sean Welleck, Yejin Choi 0001 |
EMNLP | 12 |
| 2023 | Faith and Fate: Limits of Transformers on CompositionalityabstractTransformer large language models (LLMs) have sparked admiration for their exceptional performance on tasks that demand intricate multi-step reasoning. Yet, these models simultaneously show failures on surprisingly trivial problems.
This begs the question: Are these errors incidental, or do they signal more substantial limitations?
In an attempt to demystify transformer LLMs, we investigate the limits of these models across three representative compositional tasks---multi-digit multiplication, logic grid puzzles, and a classic dynamic programming problem. These tasks require breaking problems down into sub-steps and synthesizing these steps into a precise answer. We formulate compositional tasks as computation graphs to systematically quantify the level of complexity, and break down reasoning steps into intermediate sub-procedures.
Our empirical findings suggest that transformer LLMs solve compositional tasks by reducing multi-step compositional reasoning into linearized subgraph matching, without necessarily developing systematic problem-solving skills. To round off our empirical study, we provide theoretical arguments on abstract multi-step reasoning problems that highlight how autoregressive generations' performance can rapidly decay with increased task complexity. Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Li 0069, Bill Y. Lin, Sean Welleck, Peter West, Chandra Bhagavatula, Ronan Le Bras 0001, Jena D. Hwang, Soumya Sanyal 0001, Xiang Ren 0001, Allyson Ettinger, Zaïd Harchaoui, Yejin Choi 0001 |
NeurIPS | 6 |
| 2023 | SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive TasksabstractWe introduce SwiftSage, a novel agent framework inspired by the dual-process theory of human cognition, designed to excel in action planning for complex interactive reasoning tasks. SwiftSage integrates the strengths of behavior cloning and prompting large language models (LLMs) to enhance task completion performance. The framework comprises two primary modules: the Swift module, representing fast and intuitive thinking, and the Sage module, emulating deliberate thought processes. The Swift module is a small encoder-decoder LM fine-tuned on the oracle agent's action trajectories, while the Sage module employs LLMs such as GPT-4 for subgoal planning and grounding. We develop a heuristic method to harmoniously integrate the two modules, resulting in a more efficient and robust problem-solving process. In 30 tasks from the ScienceWorld benchmark, SwiftSage significantly outperforms other methods such as SayCan, ReAct, and Reflexion, demonstrating its effectiveness in solving complex interactive tasks. Bill Y. Lin, Yicheng Fu, Karina Yang, Faeze Brahman, Shiyu Huang 0001, Chandra Bhagavatula, Prithviraj Ammanabrolu, Yejin Choi 0001, Xiang Ren 0001 |
NeurIPS | 1 |
| 2023 | Knowledge-Augmented Methods for Natural Language ProcessingabstractKnowledge in NLP has been a rising trend especially after the advent of large-scale pre-trained models. Knowledge is critical to equip statistics-based models with common sense, logic and other external information. In this tutorial, we will introduce recent state-of-the-art works in applying knowledge in language understanding, language generation and commonsense reasoning. Chenguang Zhu 0001, Yichong Xu, Xiang Ren 0001, Bill Y. Lin, Meng Jiang 0001, Wenhao Yu 0002 |
WSDM | 4 |
| 2022 | On Continual Model Refinement in Out-of-Distribution Data StreamsabstractReal-world natural language processing (NLP) models need to be continually updated to fix the prediction errors in out-of-distribution (OOD) data streams while overcoming catastrophic forgetting.However, existing continual learning (CL) problem setups cannot cover such a realistic and complex scenario.In response to this, we propose a new CL problem formulation dubbed continual model refinement (CMR).Compared to prior CL settings, CMR is more practical and introduces unique challenges (boundary-agnostic and non-stationary distribution shift, diverse mixtures of multiple OOD data clusters, error-centric streams, etc.).We extend several existing CL approaches to the CMR setting and evaluate them extensively.For benchmarking and analysis, we propose a general sampling algorithm to obtain dynamic OOD data streams with controllable nonstationarity, as well as a suite of metrics measuring various aspects of online performance.Our experiments and detailed analysis reveal the promise and challenges of the CMR problem, supporting that studying CMR in dynamic OOD streams can benefit the longevity of deployed NLP models in production. 1 Bill Y. Lin, Sida I. Wang, Xi Victoria Lin, Robin Jia, Xiang Ren 0001, Scott Yih |
ACL (1) | 1 |
| 2022 | Reflect, Not Reflex: Inference-Based Common Ground Improves Dialogue Response QualityabstractHuman communication relies on common ground (CG), the mutual knowledge and beliefs shared by participants, to produce coherent and interesting conversations.In this paper, we demonstrate that current response generation (RG) models produce generic and dull responses in dialogues because they act reflexively, failing to explicitly model CG, both due to the lack of CG in training data and the standard RG training procedure.We introduce Reflect, a dataset that annotates dialogues with explicit CG (materialized as inferences approximating shared knowledge and beliefs) and solicits 9k diverse human-generated responses each following one common ground.Using Reflect, we showcase the limitations of current dialogue data and RG models: less than half of the responses in current data is rated as high quality (sensible, specific, and interesting) and models trained using this data have even lower quality, while most Reflect responses are judged high quality.Next, we analyze whether CG can help models produce better quality responses by using Reflect CG to guide RG models.Surprisingly, we find that simply prompting GPT3 to "think" about CG generates 30% more quality responses, showing promising benefits to integrating CG into the RG process. 1 Hyundong Cho, Pegah Jandaghi, Bill Y. Lin, Jay Pujara, Xiang Ren 0001 |
EMNLP | 5 |
| 2022 | On the Robustness of Reading Comprehension Models to Entity RenamingabstractJun Yan, Yang Xiao, Sagnik Mukherjee, Bill Yuchen Lin, Robin Jia, Xiang Ren. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jun Yan 0012, Sagnik Mukherjee, Bill Y. Lin, Robin Jia, Xiang Ren 0001 |
NAACL-HLT | 4 |
| 2022 | Unsupervised Cross-Task Generalization via Retrieval AugmentationabstractHumans can perform unseen tasks by recalling relevant skills acquired previously and then generalizing them to the target tasks, even if there is no supervision at all. In this paper, we aim to improve this kind of cross-task generalization ability of massive multi-task language models, such as T0 and FLAN, in an unsupervised setting. We propose a retrieval-augmentation method named ReCross that takes a few unlabelled examples as queries to retrieve a small subset of upstream data and uses them to update the multi-task model for better generalization. ReCross is a straightforward yet effective retrieval method that combines both efficient dense retrieval and effective pair-wise reranking. Our results and analysis show that it significantly outperforms both non-retrieval methods and other baseline methods. Bill Y. Lin, Kangmin Tan, Beiwen Tian, Xiang Ren 0001 |
NeurIPS | 1 |
| 2021 | IsoBN: Fine-Tuning BERT with Isotropic Batch NormalizationabstractFine-tuning pre-trained language models (PTLMs), such as BERT and its better variant RoBERTa, has been a common practice for advancing performance in natural language understanding (NLU) tasks. Recent advance in representation learning shows that isotropic (i.e., unit-variance and uncorrelated) embeddings can significantly improve performance on downstream tasks with faster convergence and better generalization. The isotropy of the pre-trained embeddings in PTLMs, however, is relatively under-explored. In this paper, we analyze the isotropy of the pre-trained [CLS] embeddings of PTLMs with straightforward visualization, and point out two major issues: high variance in their standard deviation, and high correlation between different dimensions. We also propose a new network regularization method, isotropic batch normalization (IsoBN) to address the issues, towards learning more isotropic representations in fine-tuning by dynamically penalizing dominating principal components. This simple yet effective fine-tuning method yields about 1.0 absolute increment on the average of seven NLU tasks. Wenxuan Zhou 0002, Bill Y. Lin, Xiang Ren 0001 |
AAAI | 2 |
| 2021 | Common Sense Beyond English: Evaluating and Improving Multilingual Language Models for Commonsense ReasoningabstractBill Yuchen Lin, Seyeon Lee, Xiaoyang Qiao, Xiang Ren. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Bill Y. Lin, Seyeon Lee, Xiaoyang Qiao, Xiang Ren 0001 |
ACL/IJCNLP (1) | 1 |
| 2021 | RockNER: A Simple Method to Create Adversarial Examples for Evaluating the Robustness of Named Entity Recognition ModelsabstractTo audit the robustness of named entity recognition (NER) models, we propose RockNER, a simple yet effective method to create natural adversarial examples.Specifically, at the entity level, we replace target entities with other entities of the same semantic class in Wikidata; at the context level, we use pre-trained language models (e.g., BERT) to generate word substitutions.Together, the two levels of attack produce natural adversarial examples that result in a shifted distribution from the training data on which our target models have been trained.We apply the proposed method to the OntoNotes dataset and create a new benchmark named OntoRock for evaluating the robustness of existing NER models via a systematic evaluation protocol.Our experiments and analysis reveal that even the best model has a significant performance drop, and these models seem to memorize in-domain entity patterns instead of reasoning from the context.Our work also studies the effects of a few simple data augmentation methods to improve the robustness of NER models. 1 Bill Y. Lin, Wenyang Gao, Jun Yan 0012, Ryan Moreno, Xiang Ren 0001 |
EMNLP (1) | 1 |
| 2021 | CrossFit: A Few-shot Learning Challenge for Cross-task Generalization in NLPabstractHumans can learn a new language task efficiently with only few examples, by leveraging their knowledge obtained when learning prior tasks.In this paper, we explore whether and how such cross-task generalization ability can be acquired, and further applied to build better few-shot learners across diverse NLP tasks.We introduce CROSSFIT , a problem setup for studying cross-task generalization ability, which standardizes seen/unseen task partitions, data access during different learning stages, and the evaluation protocols.To instantiate different seen/unseen task partitions in CROSS-FIT and facilitate in-depth analysis, we present the NLP Few-shot Gym, a repository of 160 diverse few-shot NLP tasks created from openaccess NLP datasets and converted to a unified text-to-text format.Our analysis reveals that the few-shot learning ability on unseen tasks can be improved via an upstream learning stage using a set of seen tasks.We also observe that the selection of upstream learning tasks can significantly influence few-shot performance on unseen tasks, asking further analysis on task similarity and transferability. 1 Qinyuan Ye, Bill Y. Lin, Xiang Ren 0001 |
EMNLP (1) | 2 |
| 2021 | RICA: Evaluating Robust Inference Capabilities Based on Commonsense AxiomsabstractPre-trained language models (PTLMs) have achieved impressive performance on commonsense inference benchmarks, but their ability to employ commonsense to make robust inferences, which is crucial for effective communications with humans, is debated.In the pursuit of advancing fluid human-AI communication, we propose a new challenge, RICA: Robust Inference using Commonsense Axioms, that evaluates robust commonsense inference despite textual perturbations.To generate data for this challenge, we develop a systematic and scalable procedure using commonsense knowledge bases and probe PTLMs across two different evaluation settings.Extensive experiments on our generated probe sets with more than 10k statements show that PTLMs perform no better than random guessing on the zero-shot setting, are heavily impacted by statistical biases, and are not robust to perturbation attacks.We also find that fine-tuning on similar statements offer limited gains, as PTLMs still fail to generalize to unseen inferences.Our new large-scale benchmark exposes a significant gap between PTLMs and human-level language understanding and offers a new challenge for PTLMs to demonstrate commonsense. 1 Logical TemplateRel(A,B,r) à Comp(Prop(A,p), Prop(B,p)) Rahul Khanna, Seyeon Lee, Bill Y. Lin, Daniel Ho, Jay Pujara, Xiang Ren 0001 |
EMNLP (1) | 4 |
| 2021 | Differentiable Open-Ended Commonsense ReasoningabstractBill Yuchen Lin, Haitian Sun, Bhuwan Dhingra, Manzil Zaheer, Xiang Ren, William Cohen. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Bill Y. Lin, Haitian Sun, Bhuwan Dhingra, Manzil Zaheer, Xiang Ren 0001, William W. Cohen |
NAACL-HLT | 1 |
| 2020 | Learning to Contextually Aggregate Multi-Source Supervision for Sequence LabelingabstractSequence labeling is a fundamental task for a range of natural language processing problems. When used in practice, its performance is largely influenced by the annotation quality and quantity, and meanwhile, obtaining ground truth labels is often costly. In many cases, ground truth labels do not exist, but noisy annotations or annotations from different domains are accessible. In this paper, we propose a novel framework Consensus Network (ConNet) that can be trained on annotations from multiple sources (e.g., crowd annotation, cross-domain data). It learns individual representation for every source and dynamically aggregates source-specific knowledge by a context-aware attention module. Finally, it leads to a model reflecting the agreement (consensus) among multiple sources. We evaluate the proposed framework in two practical settings of multi-source learning: learning with crowd annotations and unsupervised cross-domain model adaptation. Extensive experimental results show that our model achieves significant improvements over existing methods in both settings. We also demonstrate that the method can apply to various tasks and cope with different encoders. Ouyu Lan, Bill Y. Lin, Xiang Ren 0001 |
ACL | 3 |
| 2020 | TriggerNER: Learning with Entity Triggers as Explanations for Named Entity RecognitionabstractTraining neural models for named entity recognition (NER) in a new domain often requires additional human annotations that are usually expensive and time-consuming to collect.Thus, a crucial research question is how to obtain supervision in a cost-effective way.In this paper, we introduce "entity triggers," an effective proxy of human explanations for facilitating label-efficient learning of NER models.An entity trigger is defined as a group of words in a sentence that helps to explain why humans would recognize an entity in the sentence.We crowd-sourced 14k entity triggers for two well-studied NER datasets 1 .Our proposed model, Trigger Matching Network, jointly learns trigger representations and soft matching module with self-attention such that can generalize to unseen sentences easily for tagging.The framework is significantly more cost-effective than the traditional frameworks. Bill Y. Lin, Ryan Moreno, Prashant Shiralkar, Xiang Ren 0001 |
ACL | 1 |
| 2020 | Scalable Multi-Hop Relational Reasoning for Knowledge-Aware Question AnsweringabstractExisting work that augment question answering (QA) models with external knowledge (e.g., knowledge graphs) either struggle to model multi-hop relations efficiently, or lack transparency into the model's prediction rationale.In this paper, we propose a novel knowledge-aware approach that equips pretrained language models (PTLMs) with a multi-hop relational reasoning module, named multi-hop graph relation network (MHGRN).It performs multi-hop, multi-relational reasoning over subgraphs extracted from external knowledge graphs.The proposed reasoning module unifies path-based reasoning methods and graph neural networks and results in better interpretability and scalability.We also empirically show its effectiveness and scalability on CommonsenseQA and OpenbookQA datasets, and interpret its behaviors with case studies, with the code for experiments released 1 . Yanlin Feng, Bill Y. Lin, Peifeng Wang, Jun Yan 0012, Xiang Ren 0001 |
EMNLP (1) | 3 |
| 2020 | Birds have four legs?! NumerSense: Probing Numerical Commonsense Knowledge of Pre-Trained Language ModelsabstractRecent works show that pre-trained language models (PTLMs), such as BERT, possess certain commonsense and factual knowledge.They suggest that it is promising to use PTLMs as "neural knowledge bases" via predicting masked words.Surprisingly, we find that this may not work for numerical commonsense knowledge (e.g., a bird usually has two legs).In this paper, we investigate whether and to what extent we can induce numerical commonsense knowledge from PTLMs as well as the robustness of this process.To study this, we introduce a novel probing task with a diagnostic dataset, NUMERSENSE 1 , containing 13.6k masked-word-prediction probes (10.5k for fine-tuning and 3.1k for testing).Our analysis reveals that: (1) BERT and its stronger variant RoBERTa perform poorly on the diagnostic dataset prior to any fine-tuning; (2) finetuning with distant supervision brings some improvement; (3) the best supervised model still performs poorly as compared to human performance (54.06% vs 96.3% in accuracy). Bill Y. Lin, Seyeon Lee, Rahul Khanna, Xiang Ren 0001 |
EMNLP (1) | 1 |
| 2020 | FreeDOM: A Transferable Neural Architecture for Structured Information Extraction on Web DocumentsabstractExtracting structured data from HTML documents is a long-studied problem with a broad range of applications like augmenting knowledge bases, supporting faceted search, and providing domain-specific experiences for key verticals like shopping and movies. Previous approaches have either required a small number of examples for each target site or relied on carefully handcrafted heuristics built over visual renderings of websites. In this paper, we present a novel two-stage neural approach, named FreeDOM, which overcomes both these limitations. The first stage learns a representation for each DOM node in the page by combining both the text and markup information. The second stage captures longer range distance and semantic relatedness using a relational neural network. By combining these stages, FreeDOM is able to generalize to unseen sites after training on a small number of seed sites from that vertical without requiring expensive hand-crafted features over visual renderings of the page. Through experiments on a public dataset with 8 different verticals, we show that FreeDOM beats the previous state of the art by nearly 3.7 F1 points on average without requiring features over rendered pages or expensive hand-crafted features. Bill Y. Lin, Ying Sheng 0002, Nguyen Vo, Sandeep Tata |
KDD | 1 |
| 2020 | NERO: A Neural Rule Grounding Framework for Label-Efficient Relation ExtractionabstractDeep neural models for relation extraction tend to be less reliable when perfectly labeled data is limited, despite their success in label-sufficient scenarios. Instead of seeking more instance-level labels from human annotators, here we propose to annotate frequent surface patterns to form labeling rules. These rules can be automatically mined from large text corpora and generalized via a soft rule matching mechanism. Prior works use labeling rules in an exact matching fashion, which inherently limits the coverage of sentence matching and results in the low-recall issue. In this paper, we present a neural approach to ground rules for RE, named Nero, which jointly learns a relation extraction module and a soft matching module. One can employ any neural relation extraction models as the instantiation for the RE module. The soft matching module learns to match rules with semantically similar sentences such that raw corpora can be automatically labeled and leveraged by the RE module (in a much better coverage) as augmented supervision, in addition to the exactly matched sentences. Extensive experiments and analysis on two public and widely-used datasets demonstrate the effectiveness of the proposed Nero framework, comparing with both rule-based and semi-supervised methods. Through user studies, we find that the time efficiency for a human to annotate rules and sentences are similar (0.30 vs. 0.35 min per label). In particular, Nero’s performance using 270 rules is comparable to the models trained using 3,000 labeled sentences, yielding a 9.5x speedup. Moreover, Nero can predict for unseen relations at test time and provide interpretable predictions. We release our code1 to the community for future research. Wenxuan Zhou 0002, Bill Y. Lin, Ziqi Wang 0003, Junyi Du, Leonardo Neves, Xiang Ren 0001 |
WWW | 3 |
| 2019 | KagNet: Knowledge-Aware Graph Networks for Commonsense ReasoningabstractBill Yuchen Lin, Xinyue Chen, Jamin Chen, Xiang Ren. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Bill Y. Lin, Jamin Chen, Xiang Ren 0001 |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Dynamic Detection of Communities and Their Evolutions in Temporal Social NetworksabstractIn this paper, we propose a novel community detection model, which explores the dynamic community evolutions in temporal social networks by modeling temporal affiliation strength between users and communities. Instead of transforming dynamic networks into static networks, our model utilizes normal distribution to estimate the change of affiliation strength more concisely and comprehensively. Extensive quantitative and qualitative evaluation on large social network datasets shows that our model achieves improvements in terms of prediction accuracy and reveals distinctive insight about evolutions of temporal social networks. Yaowei Huang, Jinghuan Shang, Bill Y. Lin, Luoyi Fu, Xinbing Wang |
AAAI | 3 |
| 2018 | Mining Cross-Cultural Differences and Similarities in Social MediaabstractCross-cultural differences and similarities are common in cross-lingual natural language understanding, especially for research in social media.For instance, people of distinct cultures often hold different opinions on a single named entity.Also, understanding slang terms across languages requires knowledge of cross-cultural similarities.In this paper, we study the problem of computing such cross-cultural differences and similarities.We present a lightweight yet effective approach, and evaluate it on two novel tasks: 1) mining cross-cultural differences of named entities and 2) finding similar terms for slang across languages.Experimental results show that our framework substantially outperforms a number of baseline methods on both tasks.The framework could be useful for machine translation applications and research in computational social science. Bill Y. Lin, Frank F. Xu, Kenny Q. Zhu, Seung-won Hwang |
ACL (1) | 1 |
| 2018 | Neural Adaptation Layers for Cross-domain Named Entity RecognitionabstractRecent research efforts have shown that neural architectures can be effective in conventional information extraction tasks such as named entity recognition, yielding state-of-the-art results on standard newswire datasets.However, despite significant resources required for training such models, the performance of a model trained on one domain typically degrades dramatically when applied to a different domain, yet extracting entities from new emerging domains such as social media can be of significant interest.In this paper, we empirically investigate effective methods for conveniently adapting an existing, well-trained neural NER model for a new domain.Unlike existing approaches, we propose lightweight yet effective methods for performing domain adaptation for neural models.Specifically, we introduce adaptation layers on top of existing neural architectures, where no re-training using the source domain data is required.We conduct extensive empirical studies and show that our approach significantly outperforms stateof-the-art methods. Bill Y. Lin, Wei Lu 0011 |
EMNLP | 1 |
| 2018 | ExtRA: Extracting Prominent Review Aspects from Customer FeedbackabstractMany existing systems for analyzing and summarizing customer reviews about products or service are based on a number of prominent review aspects.Conventionally, the prominent review aspects of a product type are determined manually.This costly approach cannot scale to large and cross-domain services such as Amazon.com,Taobao.com or Yelp.comwhere there are a large number of product types and new products emerge almost everyday.In this paper, we propose a novel framework, for extracting the most prominent aspects of a given product type from textual reviews.The proposed framework, ExtRA, extracts K most prominent aspect terms or phrases which do not overlap semantically automatically without supervision.Extensive experiments show that ExtRA is effective and achieves the state-of-the-art performance on a dataset consisting of different product types. Zhiyi Luo, Shanshan Huang 0002, Frank F. Xu, Bill Y. Lin, Hanyuan Shi, Kenny Q. Zhu |
EMNLP | 4 |