Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Shirley Anugrah Hayati

dblp:198/3799 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 58% Trustworthy machine learning · 17% Information extraction and text analysis · 15%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 6 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
instruction tuning
0.912025
Chain-of-Instructions: Compositional Instruction Tuning on Large Language Models · AAAI 2025
Machine learning › Trustworthy machine learning
language model interpretability
0.512021
Does BERT Learn as Humans Perceive? Understanding Linguistic Styles through Lexica · EMNLP (1) 2021
Natural language and speech › Information extraction and text analysis › text mining
writing style analysis
0.512021
Does BERT Learn as Humans Perceive? Understanding Linguistic Styles through Lexica · EMNLP (1) 2021
Program synthesis and code generation
neural program synthesis
0.312018
Retrieval-Based Neural Code Generation · EMNLP 2018
Program synthesis and code generation › code generation with language models
retrieval-augmented code generation
0.312018
Retrieval-Based Neural Code Generation · EMNLP 2018
Recommender systems › interactive recommendation
conversational recommendation
0.112020
INSPIRED: Toward Sociable Recommendation Dialog Systems · EMNLP (1) 2020

Methods — techniques the papers use, named apart from their topics

fine-tuning · 0.9end-to-end dialogue modeling · 0.9annotation scheme from social science theories · 0.9step-by-step recall prompting · 0.8large language model · 0.8criteria-based prompting · 0.8fine-tuned style classifiers · 0.5crowd annotation · 0.5BERT · 0.5subtree retrieval · 0.3neural encoder-decoder · 0.3dynamic-programming sentence similarity · 0.3
YearPublicationVenuePosition
2025 Chain-of-Instructions: Compositional Instruction Tuning on Large Language Models
abstract
Fine-tuning large language models (LLMs) with a collection of large and diverse instructions has improved the model’s generalization to different tasks, even for unseen tasks. However, most existing instruction datasets include only single instructions, and they struggle to follow complex instructions composed of multiple subtasks. In this work, we propose a novel concept of compositional instructions called chain-of-instructions (CoI), where the output of one instruction becomes an input for the next like a chain. Unlike the conventional practice of solving single instruction tasks, our proposed method encourages a model to solve each subtask step by step until the final answer is reached. CoI-tuning (i.e., fine-tuning with CoI instructions) improves the model’s ability to handle instructions composed of multiple subtasks as well as unseen composite tasks such as multilingual summarization. Overall, our study find that simple CoI tuning of existing instruction data can provide consistent generalization to solve more complex, unseen, and longer chains of instructions.
Shirley Anugrah Hayati, Taehee Jung, Tristan Bodding-Long, Sudipta Kar, Abhinav Sethy, Joo-Kyung Kim, Dongyeop Kang
AAAI1
2024 How Far Can We Extract Diverse Perspectives from Large Language Models?
abstract
Collecting diverse human opinions is costly and challenging.This leads to a recent trend in exploiting large language models (LLMs) for generating diverse data for potential scalable and efficient solutions.However, the extent to which LLMs can generate diverse perspectives on subjective topics is still unclear.In this study, we explore LLMs' capacity of generating diverse perspectives and rationales on subjective topics such as social norms and argumentative texts.We introduce the problem of extracting maximum diversity from LLMs.Motivated by how humans form opinions based on values, we propose a criteria-based prompting technique to ground diverse opinions.To see how far we can extract diverse perspectives from LLMs, or called diversity coverage, we employ a step-by-step recall prompting to generate more outputs from the model iteratively.Our methods, applied to various tasks, show that LLMs can indeed produce diverse opinions according to the degree of task subjectivity.We also find that LLMs performance of extracting maximum diversity is on par with human. 1
Shirley Anugrah Hayati, Minhwa Lee, Dheeraj Rajagopal, Dongyeop Kang
EMNLP1
2024 Consequences of Conflicts in Online Conversations
abstract
Interpersonal conflicts occur frequently in both offline and online groups, with conditions for conflict especially ripe online. This research attempts to understand the consequences of online group conflict and reporting it to group administrators, both for the protagonists in the conflict and observers. If group conflict is aversive, then group members should reduce their group participation after observing conflict. Theories of imitation and behavioral mimicry suggest that even onlookers will exhibit more conflict and negative language after observing conflict conversations in their group. In contrast, theories of deterrence suggest that both the instigator of the conflict and onlookers will reduce their conflict and onlookers might even increase their engagement if conflicts are reported to group administrators. The current study uses de-identified and aggregated data from Facebook group conversations and Mahalanobis distance matching to test these ideas. Results are consistent with the hypothesis that conflict in group conversations reduces engagement within the group and increases the amount of conflict and the negativity of language users express in the group. However, inconsistent with deterrence theories, conflict and language negativity increase and group engagement decreases when conflict is reported to group administrators.
Kristen M. Altenburger, Robert E. Kraut, Shirley Anugrah Hayati, Jane Dwivedi-Yu, Kaiyan Peng, Yi-Chia Wang
ICWSM3
2023 StyLEx: Explaining Style Using Human Lexical Annotations
abstract
Large pre-trained language models have achieved impressive results on various style classification tasks, but they often learn spurious domain-specific words to make predictions (Hayati et al., 2021).While human explanation highlights stylistic tokens as important features for this task, we observe that model explanations often do not align with them.To tackle this issue, we introduce StyLEx, a model that learns from human annotated explanations of stylistic features and jointly learns to perform the task and predict these features as model explanations.Our experiments show that StyLEx can provide human-like stylistic lexical explanations without sacrificing the performance of sentence-level style prediction on both indomain and out-of-domain datasets.Explanations from StyLEx show significant improvements in explanation metrics (sufficiency, plausibility) and when evaluated with human annotations.They are also more understandable by human judges compared to the widely-used saliency-based explanation baseline.1 * currently at Google
Shirley Anugrah Hayati, Kyumin Park, Dheeraj Rajagopal, Lyle H. Ungar, Dongyeop Kang
EACL1
2022 Modeling Motivational Interviewing Strategies on an Online Peer-to-Peer Counseling Platform
abstract
Millions of people participate in online peer-to-peer support sessions, yet there has been little prior research on systematic psychology-based evaluations of fine-grained peer-counselor behavior in relation to client satisfaction. This paper seeks to bridge this gap by mapping peer-counselor chat-messages to motivational interviewing (MI) techniques. We annotate 14,797 utterances from 734 chat conversations using 17 MI techniques and introduce four new interviewing codes such as ''chit-chat'' and ''inappropriate'' to account for the unique conversational patterns observed on online platforms. We automate the process of labeling peer-counselor responses to MI techniques by fine-tuning large domain-specific language models and then use these automated measures to investigate the behavior of the peer counselors via correlational studies. Specifically, we study the impact of MI techniques on the conversation ratings to investigate the techniques that predict clients' satisfaction with their counseling sessions. When counselors use techniques such as reflection and affirmation, clients are more satisfied. Examining volunteer counselors' change in usage of techniques suggest that counselors learn to use more introduction and open questions as they gain experience. This work provides a deeper understanding of the use of motivational interviewing techniques on peer-to-peer counselor platforms and sheds light on how to build better training programs for volunteer counselors on online platforms.
Raj Sanjay Shah, Faye Holt, Shirley Anugrah Hayati, Aastha Agarwal, Yi-Chia Wang, Robert E. Kraut, Diyi Yang
Proc. ACM Hum. Comput. Interact.3
2021 Does BERT Learn as Humans Perceive? Understanding Linguistic Styles through Lexica
abstract
People convey their intention and attitude through linguistic styles of the text that they write.In this study, we investigate lexicon usages across styles throughout two lenses: human perception and machine word importance, since words differ in the strength of the stylistic cues that they provide.To collect labels of human perception, we curate a new dataset, HUMMINGBIRD, on top of benchmarking style datasets.We have crowd workers highlight the representative words in the text that makes them think the text has the following styles: politeness, sentiment, offensiveness, and five emotion types.We then compare these human word labels with word importance derived from a popular fine-tuned style classifier like BERT.Our results show that the BERT often finds content words not relevant to the target style as important words used in style prediction, but humans do not perceive the same way even though for some styles (e.g., positive sentiment and joy) humanand machine-identified words share significant overlap for some styles.1
Shirley Anugrah Hayati, Dongyeop Kang, Lyle H. Ungar
EMNLP (1)1
2020 INSPIRED: Toward Sociable Recommendation Dialog Systems
abstract
In recommendation dialogs, humans commonly disclose their preference and make recommendations in a friendly manner.However, this is a challenge in developing a sociable recommendation dialog system, due to the lack of dialog dataset annotated with such sociable strategies.Therefore, we present INSPIRED, a new dataset of 1,001 human-human dialogs for movie recommendation with measures for successful recommendations.To better understand how humans make recommendations in communication, we design an annotation scheme related to recommendation strategies based on social science theories and annotate these dialogs.Our analysis shows that sociable recommendation strategies, such as sharing personal opinions or communicating with encouragement, more frequently lead to successful recommendations.Based on our dataset, we train end-to-end recommendation dialog systems with and without our strategy labels.In both automatic and human evaluation, our model with strategy incorporation outperforms the baseline model.This work is a first step for building sociable recommendation dialog systems with a basis of social science theories 1 .
Shirley Anugrah Hayati, Dongyeop Kang, Qingxiaoyang Zhu, Weiyan Shi 0001, Zhou Yu 0005
EMNLP (1)1
2018 Retrieval-Based Neural Code Generation
abstract
In models to generate program source code from natural language, representing this code in a tree structure has been a common approach.However, existing methods often fail to generate complex code correctly due to a lack of ability to memorize large and complex structures.We introduce RECODE, a method based on subtree retrieval that makes it possible to explicitly reference existing code examples within a neural code generation model.First, we retrieve sentences that are similar to input sentences using a dynamicprogramming-based sentence similarity scoring method.Next, we extract n-grams of action sequences that build the associated abstract syntax tree.Finally, we increase the probability of actions that cause the retrieved n-gram action subtree to be in the predicted code.We show that our approach improves the performance on two code generation tasks by up to +2.6 BLEU. 1
Shirley Anugrah Hayati, Raphaël Olivier, Pravalika Avvaru, Anthony Tomasic, Graham Neubig
EMNLP1