Weiyan Shi 0001

dblp:218/5722-1 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
16since 2021 · last 2025
0000-0003-0850-5831ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 6 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Persuasion for Social Good: How to Build and Break AI
abstract
Persuasion is important in numerous situations like healthy habit promotion, and emotional support. As AI gets more involved in our daily life, it becomes critical to study how they can persuade humans and how persuasive they are. In this talk, I will cover (1) how to build such persuasive AI systems that can persuade, negotiate, and cooperate with other humans in the game of Diplomacy. (2) I will also discuss how humans perceive such specialized AI systems. This study validates the necessity of California's Autobot Law and proposes guidance to regulate such systems. (3) As these systems become more powerful, AI safety problems become more important. So I will describe how to persuade AI models to jailbreak them and study AI safety problems. Finally, I will conclude with my long-term vision to further study persuasion from a multi-angle approach that combines Artificial Intelligence, Human-Computer Interaction, and social sciences.
Weiyan Shi 0001
AAAI1
2025 Distilling an End-to-End Voice Assistant Without Instruction Training Data
abstract
Voice assistants, such as Siri and Google Assistant, typically model audio and text separately, resulting in lost speech information and increased complexity. Recent efforts to address this with end-to-end Speech Large Language Models (speech-in, text-out) trained with supervised finetuning (SFT) have led to models “forgetting” capabilities from text-only LLMs. Our work proposes an alternative paradigm for training Speech LLMs without instruction data, using the response of a text-only LLM to transcripts as self-supervision. Importantly, this process can be performed without annotated responses. We show that our Distilled Voice Assistant (DiVA) generalizes to Spoken Question Answering, Classification, and Translation. Furthermore, DiVA better matches user preferences, achieving a 72% win rate compared with state-of-the-art models like Qwen 2 Audio, despite using >100x less training compute.
William Barr Held, Weiyan Shi 0001, Minzhi Li, Michael J. Ryan, Diyi Yang
ACL (1)3
2025 NewsInterview: a Dataset and a Playground to Evaluate LLMs' Grounding Gap via Informational Interviews
abstract
Alexander Spangher, Michael Lu, Sriya Kalyan, Hyundong Justin Cho, Tenghao Huang, Weiyan Shi, Jonathan May. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Alexander Spangher, Michael Lu, Sriya Kalyan, Hyundong Cho, Tenghao Huang, Weiyan Shi 0001, Jonathan May
ACL (1)6
2025 Proactive Conversational Agents with Inner Thoughts
Xingyu Liu 0002, Shitao Fang, Weiyan Shi 0001, Chien-Sheng Wu, Takeo Igarashi, Xiang 'Anthony' Chen
CHI3
2025 LLMs Encode Harmfulness and Refusal Separately
abstract
LLMs are trained to refuse harmful instructions, but do they truly understand harmfulness beyond just refusing? Prior work has shown that LLMs’ refusal behaviors can be mediated by a one-dimensional subspace, i.e., a refusal direction. In this work, we identify a new dimension to analyze safety mechanisms in LLMs, i.e., harmfulness, which is encoded internally as a separate concept from refusal. And there exists a harmfulness direction that is distinct from the refusal direction. As causal evidence, steering along the harmfulness direction can lead LLMs to interpret harmless instructions as harmful, but steering along the refusal direction tends to elicit refusal responses directly without reversing the model’s judgment on harmfulness. Furthermore, using our identified harmfulness concept, we find that certain jailbreak methods work by reducing the refusal signals without suppressing the model’s internal belief of harmfulness. We also find that adversarially fine- tuning models to accept harmful instructions has minimal impact on the model’s internal belief of harmfulness. These insights lead to a practical safety application: The model’s latent harmfulness representation can serve as an intrinsic safeguard (Latent Guard) for detecting unsafe inputs and reducing over-refusals that is robust to finetuning attacks. For instance, our Latent Guard achieves performance comparable to or better than Llama Guard 3 8B, a dedicated finetuned safeguard model, across different jailbreak methods. Our findings suggest that LLMs’ internal understanding of harmfulness is more robust than their refusal decision to diverse input instructions, offering a new perspective to study AI safety.
Zhengxuan Wu, David Bau, Weiyan Shi 0001
NeurIPS5
2024 How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
abstract
Most traditional AI safety research views models as machines and centers on algorithmfocused attacks developed by security experts.As large language models (LLMs) become increasingly common and competent, non-expert users can also impose risks during daily interactions.Observing this, we shift the perspective, by treating LLMs as human-like communicators to examine the interplay between everyday language interaction and AI safety.Specifically, we study how to persuade LLMs to jailbreak them.First, we propose a persuasion taxonomy derived from decades of social science research.Then, we apply the taxonomy to automatically generate persuasive adversarial prompts (PAP) to jailbreak LLMs.Results show that persuasion significantly increases the jailbreak risk across all risk categories: PAP consistently achieves an attack success rate of over 92% on Llama-2-7b-Chat, GPT-3.5, and GPT-4 in 10 trials, surpassing recent algorithm-focused attacks.On the defense side, we explore various mechanisms against PAP, find a significant gap in existing defenses, and advocate for more fundamental solutions for AI safety 1 .
Yi Zeng 0005, Hongpeng Lin, Diyi Yang, Ruoxi Jia 0001, Weiyan Shi 0001
ACL (1)6
2024 The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation
abstract
Rongwu Xu, Brian Lin, Shujian Yang, Tianqi Zhang, Weiyan Shi, Tianwei Zhang, Zhixuan Fang, Wei Xu, Han Qiu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Rongwu Xu, Brian S. Lin, Shujian Yang, Weiyan Shi 0001, Tianwei Zhang 0004, Zhixuan Fang, Wei Xu 0039, Han Qiu 0001
ACL (1)5
2024 The Mirrored Influence Hypothesis: Efficient Data Influence Estimation by Harnessing Forward Passes
abstract
Large-scale black-box models have become ubiquitous across numerous applications. Understanding the influence of individual training data sources on predictions made by these models is crucial for improving their trustworthiness. Current influence estimation techniques involve computing gradients for every training point or repeated training on different subsets. These approaches face obvious computational challenges when scaled up to large datasets and models. In this paper, we introduce and explore the Mirrored Influence Hypothesis, highlighting a reciprocal nature of influence between training and test data. Specifically, it sug-gests that evaluating the influence of training data on test predictions can be reformulated as an equivalent, yet inverse problem: assessing how the predictions for training samples would be altered if the model were trained on specific test samples. Through both empirical and theoretical validations, we demonstrate the wide applicability of our hypothesis. Inspired by this, we introduce a new method for estimating the influence of training data, which requires calculating gradients for specific test samples, paired with a forward pass for each training point. This approach can capitalize on the common asymmetry in scenarios where the number of test samples under concurrent examination is much smaller than the scale of the training dataset, thus gaining a significant improvement in efficiency compared to existing approaches. We demonstrate the applicability of our method across a range of scenarios, including data attribution in diffusion models, data leakage detection, analy-sis of memorization, mislabeled data detection, and tracing behavior in language models.
Myeongseob Ko, Feiyang Kang, Weiyan Shi 0001, Ming Jin 0002, Zhou Yu 0005, Ruoxi Jia 0001
CVPR3
2024 Position: A Safe Harbor for AI Evaluation and Red Teaming
abstract
Independent evaluation and red teaming are critical for identifying the risks posed by generative AI systems. However, the terms of service and enforcement strategies used by prominent AI companies to deter model misuse have disincentives on good faith safety evaluations. This causes some researchers to fear that conducting such research or releasing their findings will result in account suspensions or legal reprisal. Although some companies offer researcher access programs, they are an inadequate substitute for independent research access, as they have limited community representation, receive inadequate funding, and lack independence from corporate incentives. We propose that major generative AI developers commit to providing a legal and technical safe harbor, protecting public interest safety research and removing the threat of account suspensions or legal reprisal. These proposals emerged from our collective experience conducting safety, privacy, and trustworthiness research on generative AI systems, where norms and incentives could be better aligned with public interests, without exacerbating model misuse. We believe these commitments are a necessary step towards more inclusive and unimpeded community efforts to tackle the risks of generative AI.
Shayne Longpre, Sayash Kapoor, Kevin Klyman, Ashwin Ramaswami, Rishi Bommasani, Borhane Blili-Hamelin, Yangsibo Huang, Aviya Skowron, Suhas Kotha, Yi Zeng 0005, Weiyan Shi 0001, Xianjun Yang, Reid Southen, Alexander Robey, Patrick Chao, Diyi Yang, Ruoxi Jia 0001, Daniel Kang 0001, Alex Pentland, Arvind Narayanan, Percy Liang, Peter Henderson 0002
ICML12
2024 When Life Gives You Lemons, Make Cherryade: Converting Feedback from Bad Responses into Good Labels
abstract
Weiyan Shi, Emily Dinan, Kurt Shuster, Jason Weston, Jing Xu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Weiyan Shi 0001, Emily Dinan, Kurt Shuster 0001, Jason Weston, Jing Xu 0014
NAACL-HLT1
2024 PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action
abstract
As language models (LMs) are widely utilized in personalized communication scenarios (e.g., sending emails, writing social media posts) and endowed with a certain level of agency, ensuring they act in accordance with the contextual privacy norms becomes increasingly critical. However, quantifying the privacy norm awareness of LMs and the emerging privacy risk in LM-mediated communication is challenging due to (1) the contextual and long-tailed nature of privacy-sensitive cases, and (2) the lack of evaluation approaches that capture realistic application scenarios. To address these challenges, we propose PrivacyLens, a novel framework designed to extend privacy-sensitive seeds into expressive vignettes and further into agent trajectories, enabling multi-level evaluation of privacy leakage in LM agents' actions. We instantiate PrivacyLens with a collection of privacy norms grounded in privacy literature and crowdsourced seeds. Using this dataset, we reveal a discrepancy between LM performance in answering probing questions and their actual behavior when executing user instructions in an agent setup. State-of-the-art LMs, like GPT-4 and Llama-3-70B, leak sensitive information in 25.68% and 38.69% of cases, even when prompted with privacy-enhancing instructions. We also demonstrate the dynamic nature of PrivacyLens by extending each seed into multiple trajectories to red-team LM privacy leakage risk. Dataset and code are available at https://github.com/SALT-NLP/PrivacyLens.
Yijia Shao, Tianshi Li 0001, Weiyan Shi 0001, Diyi Yang
NeurIPS3
2024 Dialoging Resonance in Human-Chatbot Conversation: How Users Perceive and Reciprocate Recommendation Chatbot's Self-Disclosure Strategy
abstract
Using chatbots to make recommendations is increasingly popular. The design of recommendation chatbots has mainly been taking an information-centric approach by focusing on the recommended content per se. Limited attention is on how social connection and relational strategies, such as self-disclosure from a chatbot, may influence users' perception and acceptance of the recommendation. In this work, we designed, implemented, and evaluated a social chatbot capable of performing three different levels of self-disclosure: factual information (low), cognitive opinions (medium), and emotions (high). In the evaluation, we recruited 372 participants to converse with the chatbot on two topics: movies and COVID-19 experiences. In each topic, the chatbot conducted small talks and made relevant recommendations to the topic. Participants were randomly assigned to four experimental conditions where the chatbot used factual, cognitive, emotional, and adaptive strategies to perform self-disclosures. By training a text classifier to identify users' level of self-disclosure in real-time, the adaptive chatbot can dynamically match its self-disclosure language to the level of disclosure exhibited by the users. Our results show that users reciprocate with higher-level self-disclosure when a recommendation chatbot displays emotions throughout the conversation. The utilization of emotional disclosure by the chatbot resulted in enhanced enjoyment during interactions and a more favorable perception of the bot. This, in turn, led to greater effectiveness in making recommendations, including a higher likelihood of accepting the recommendation. We discuss the understandings obtained and implications to future design.
Kaihui Liang, Weiyan Shi 0001, Hao-Chuan Wang, Zhou Yu 0005
Proc. ACM Hum. Comput. Interact.2
2023 Social Influence Dialogue Systems: A Survey of Datasets and Models For Social Influence Tasks
abstract
Dialogue systems capable of social influence such as persuasion, negotiation, and therapy, are essential for extending the use of technology to numerous realistic scenarios.However, existing research primarily focuses on either task-oriented or open-domain scenarios, a categorization that has been inadequate for capturing influence skills systematically.There exists no formal definition or category for dialogue systems with these skills and data-driven efforts in this direction are highly limited.In this work, we formally define and introduce the category of social influence dialogue systems that influence users' cognitive and emotional responses, leading to changes in thoughts, opinions, and behaviors through natural conversations.We present a survey of various tasks, datasets, and methods, compiling the progress across seven diverse domains.We discuss the commonalities and differences between the examined systems, identify limitations, and recommend future directions.This study serves as a comprehensive reference for social influence dialogue systems to inspire more dedicated research and discussion in this emerging area.
Kushal Chawla, Weiyan Shi 0001, Gale M. Lucas, Zhou Yu 0005, Jonathan Gratch
EACL2
2022 Just Fine-tune Twice: Selective Differential Privacy for Large Language Models
abstract
Protecting large language models from privacy leakage is becoming increasingly crucial with their wide adoption in real-world products.Yet applying differential privacy (DP), a canonical notion with provable privacy guarantees for machine learning models, to those models remains challenging due to the trade-off between model utility and privacy loss.Utilizing the fact that sensitive information in language data tends to be sparse, Shi et al. (2021) formalized a DP notion extension called Selective Differential Privacy (SDP) to protect only the sensitive tokens defined by a policy function.However, their algorithm only works for RNN-based models.In this paper, we develop a novel framework, Just Fine-tune Twice (JFT), that achieves SDP for state-of-the-art large transformer-based models.Our method is easy to implement: it first finetunes the model with redacted in-domain data, and then fine-tunes it again with the original in-domain data using a private training mechanism.Furthermore, we study the scenario of imperfect implementation of policy functions that misses sensitive tokens and develop systematic methods to handle it.Experiments show that our method achieves strong utility compared to previous baselines.We also analyze the SDP privacy guarantee empirically with the canary insertion attack 1 .
Weiyan Shi 0001, Ryan Shea, Si Chen 0008, Chiyuan Zhang, Ruoxi Jia 0001, Zhou Yu 0005
EMNLP1
2022 Selective Differential Privacy for Language Modeling
abstract
With the increasing applications of language models, it has become crucial to protect these models from leaking private information.Previous work has attempted to tackle this challenge by training RNN-based language models with differential privacy guarantees.However, applying classical differential privacy to language models leads to poor model performance as the underlying privacy notion is over-pessimistic and provides undifferentiated protection for all tokens in the data.Given that the private information in natural language is sparse (for example, the bulk of an email might not carry personally identifiable information), we propose a new privacy notion, selective differential privacy, to provide rigorous privacy guarantees on the sensitive portion of the data to improve model utility.To realize such a new notion, we develop a corresponding privacy mechanism, Selective-DPSGD, for RNN-based language models.Besides language modeling, we also apply the method to a more concrete application -dialog systems.Experiments on both language modeling and dialog system building show that the proposed privacy-preserving mechanism achieves better utilities while remaining safe under various privacy attacks compared to the baselines.The data and code are released to facilitate future research 1 .
Weiyan Shi 0001, Aiqi Cui, Evan Li, Ruoxi Jia 0001, Zhou Yu 0005
NAACL-HLT1
2022 Towards Socially Intelligent Agents with Mental State Transition and Human Value
abstract
Building a socially intelligent agent involves many challenges.One of which is to track the agent's mental state transition and teach the agent to make decisions guided by its value like a human.Towards this end, we propose to incorporate mental state simulation and value modeling into dialogue agents.First, we build a hybrid mental state parser that extracts information from both the dialogue and event observations and maintains a graphical representation of the agent's mind; Meanwhile, the transformer-based value model learns human preferences from the human value dataset, VALUENET.Empirical results show that the proposed model attains state-of-the-art performance on the dialogue/action/emotion prediction task in the fantasy text-adventure game dataset, LIGHT.We also show example cases to demonstrate: (i) how the proposed mental state parser can assist the agent's decision by grounding on the context like locations and objects, and (ii) how the value model can help the agent make decisions based on its personal priorities.
Liang Qiu 0001, Yuan Liang 0001, Pan Lu, Weiyan Shi 0001, Zhou Yu 0005, Song-Chun Zhu
SIGDIAL5
2020 End-to-End Trainable Non-Collaborative Dialog System
abstract
End-to-end task-oriented dialog models have achieved promising performance on collaborative tasks where users willingly coordinate with the system to complete a given task. While in non-collaborative settings, for example, negotiation and persuasion, users and systems do not share a common goal. As a result, compared to collaborate tasks, people use social content to build rapport and trust in these non-collaborative settings in order to advance their goals. To handle social content, we introduce a hierarchical intent annotation scheme, which can be generalized to different non-collaborative dialog tasks. Building upon TransferTransfo (Wolf et al. 2019), we propose an end-to-end neural network model to generate diverse coherent responses. Our model utilizes intent and semantic slots as the intermediate sentence representation to guide the generation process. In addition, we design a filter to select appropriate responses based on whether these intermediate representations fit the designed task and conversation constraints. Our non-collaborative dialog model guides users to complete the task while simultaneously keeps them engaged. We test our approach on our newly proposed AntiScam dataset and an existing PersuasionForGood dataset. Both automatic and human evaluations suggest that our model outperforms multiple baselines in these two non-collaborative tasks.
Yu Li 0013, Kun Qian 0016, Weiyan Shi 0001, Zhou Yu 0005
AAAI3
2020 Effects of Persuasive Dialogues: Testing Bot Identities and Inquiry Strategies
abstract
Intelligent conversational agents, or chatbots, can take on various identities and are increasingly engaging in more human-centered conversations with persuasive goals. However, little is known about how identities and inquiry strategies influence the conversation's effectiveness. We conducted an online study involving 790 participants to be persuaded by a chatbot for charity donation. We designed a two by four factorial experiment (two chatbot identities and four inquiry strategies) where participants were randomly assigned to different conditions. Findings showed that the perceived identity of the chatbot had significant effects on the persuasion outcome (i.e., donation) and interpersonal perceptions (i.e., competence, confidence, warmth, and sincerity). Further, we identified interaction effects among perceived identities and inquiry strategies. We discuss the findings for theoretical and practical implications for developing ethical and effective persuasive chatbots. Our published data, codes, and analyses serve as the first step towards building competent ethical persuasive chatbots.
Weiyan Shi 0001, Saurav Sahay, Zhou Yu 0005
CHI1
2020 INSPIRED: Toward Sociable Recommendation Dialog Systems
abstract
In recommendation dialogs, humans commonly disclose their preference and make recommendations in a friendly manner.However, this is a challenge in developing a sociable recommendation dialog system, due to the lack of dialog dataset annotated with such sociable strategies.Therefore, we present INSPIRED, a new dataset of 1,001 human-human dialogs for movie recommendation with measures for successful recommendations.To better understand how humans make recommendations in communication, we design an annotation scheme related to recommendation strategies based on social science theories and annotate these dialogs.Our analysis shows that sociable recommendation strategies, such as sharing personal opinions or communicating with encouragement, more frequently lead to successful recommendations.Based on our dataset, we train end-to-end recommendation dialog systems with and without our strategy labels.In both automatic and human evaluation, our model with strategy incorporation outperforms the baseline model.This work is a first step for building sociable recommendation dialog systems with a basis of social science theories 1 .
Shirley Anugrah Hayati, Dongyeop Kang, Qingxiaoyang Zhu, Weiyan Shi 0001, Zhou Yu 0005
EMNLP (1)4
2020 Structured Attention for Unsupervised Dialogue Structure Induction
abstract
Inducing a meaningful structural representation from one or a set of dialogues is a crucial but challenging task in computational linguistics. Advancement made in this area is critical for dialogue system design and discourse analysis. It can also be extended to solve grammatical inference. In this work, we propose to incorporate structured attention layers into a Variational Recurrent Neural Network (VRNN) model with discrete latent states to learn dialogue structure in an unsupervised fashion. Compared to a vanilla VRNN, structured attention enables a model to focus on different parts of the source sentence embeddings while enforcing a structural inductive bias. Experiments show that on two-party dialogue datasets, VRNN with structured attention learns semantic structures that are similar to templates used to generate this dialogue corpus. While on multi-party dialogue datasets, our model learns an interactive structure demonstrating its capability of distinguishing speakers or addresses, automatically disentangling dialogues without explicit human annotation.
Liang Qiu 0001, Weiyan Shi 0001, Yuan Liang 0001, Feng Shi 0006, Zhou Yu 0005, Song-Chun Zhu
EMNLP (1)3
2019 Persuasion for Good: Towards a Personalized Persuasive Dialogue System for Social Good
abstract
Developing intelligent persuasive conversational agents to change people's opinions and actions for social good is the frontier in advancing the ethical development of automated dialogue systems.To do so, the first step is to understand the intricate organization of strategic disclosures and appeals employed in human persuasion conversations.We designed an online persuasion task where one participant was asked to persuade the other to donate to a specific charity.We collected a large dataset with 1,017 dialogues and annotated emerging persuasion strategies from a subset.Based on the annotation, we built a baseline classifier with context information and sentence-level features to predict the 10 persuasion strategies used in the corpus.Furthermore, to develop an understanding of personalized persuasion processes, we analyzed the relationships between individuals' demographic and psychological backgrounds including personality, morality, value systems, and their willingness for donation.Then, we analyzed which types of persuasion strategies led to a greater amount of donation depending on the individuals' personal backgrounds.This work lays the ground for developing a personalized persuasive dialogue system. 1
Weiyan Shi 0001, Richard Kim, Zhou Yu 0005
ACL (1)2
2019 How to Build User Simulators to Train RL-based Dialog Systems
abstract
Weiyan Shi, Kun Qian, Xuewei Wang, Zhou Yu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Weiyan Shi 0001, Kun Qian 0016, Zhou Yu 0005
EMNLP/IJCNLP (1)1
2018 Sentiment Adaptive End-to-End Dialog Systems
abstract
End-to-end learning framework is useful for building dialog systems for its simplicity in training and efficiency in model updating.However, current end-to-end approaches only consider user semantic inputs in learning and under-utilize other user information.Therefore, we propose to include user sentiment obtained through multimodal information (acoustic, dialogic and textual), in the end-to-end learning framework to make systems more user-adaptive and effective.We incorporated user sentiment information in both supervised and reinforcement learning settings.In both settings, adding sentiment information reduced the dialog length and improved the task success rate on a bus information search task.This work is the first attempt to incorporate multimodal user information in the adaptive end-toend dialog system training framework and attained state-of-the-art performance.
Weiyan Shi 0001, Zhou Yu 0005
ACL (1)1