Chenhao Tan

dblp:95/8314 · DBLP profile ↗
← Back
68ranked-venue papers
12as first author
31since 2021 · last 2026
0000-0002-3981-2116ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 7 first-author · 19 since 2021Human-computer interaction and ubiquitous computing · 22 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 21 · 9 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 5 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 HypoEval: Hypothesis-Guided Evaluation for Natural Language Generation
abstract
Large language models (LLMs) have demonstrated great potential for automating the evaluation of natural language generation.Previous frameworks of LLM-as-a-judge fall short in two ways: they either use zero-shot setting without consulting any human input, which leads to low alignment, or fine-tune LLMs on labeled data, which requires a non-trivial number of samples.Moreover, previous methods often provide little reasoning behind automated evaluations.In this paper, we propose HYPO-EVAL, Hypothesis-guided Evaluation framework, which first uses a small corpus of human evaluations to generate more detailed rubrics for human judgments and then incorporates a checklist-like approach to combine LLM's assigned scores on each decomposed dimension to acquire overall scores 1 .With only 30 human evaluations, HypoEval achieves stateof-the-art performance in alignment with both human rankings (Spearman correlation) and human scores (Pearson correlation), on average outperforming G-Eval by 11.86% and finetuned LLAMA-3.1-8B-INSTRUCT with at least 3 times more human evaluations by 11.95%.Furthermore, we conduct systematic studies to assess the robustness of HYPOEVAL, highlighting its effectiveness as a reliable and interpretable automated evaluation framework.
Hanchen Li, Chenhao Tan
ACL (1)3
2026 Governance of AI-Generated Content: A Case Study on Social Media Platforms
abstract
Online platforms are seeing increasing amounts of AI-generated content—text and other forms of media that are made or co-created with generative AI. This trend suggests platforms may need to establish governance frameworks, including policies and enforcement strategies for how users create, post, share, and engage with such content to encourage responsible use. We investigate the governance of AI-generated content across 40 popular social media platforms. Just over two-thirds explicitly describe governance of AI-generated content spanning six themes. Most platforms focus on moderating AI-generated content that violates established content rules and discloses AI-generated content. Fewer platforms—those that are focused on creativity and knowledge-sharing—address other issues such as ownership and monetization. Based on these findings, we suggest stakeholders and policymakers develop more direct, comprehensive, and forward-looking AI-generated content governance, as well as tools and education for users about the use of such content.
Lan Gao 0001, Abani Ahmed, Oscar Chen, Margaux Reyl, Zayna Cheema, Nick Feamster, Chenhao Tan, Kurt Thomas, Marshini Chetty
CHI7
2025 The Impossibility of Fair LLMs
abstract
The rise of general-purpose artificial intelligence (AI) systems, particularly large language models (LLMs), has raised pressing moral questions about how to reduce bias and ensure fairness at scale.Researchers have documented a sort of "bias" in the significant correlations between demographics (e.g., race, gender) in LLM prompts and responses, but it remains unclear how LLM fairness could be evaluated with more rigorous definitions, such as group fairness or fair representations.We analyze a variety of technical fairness frameworks and find inherent challenges in each that make the development of a fair LLM intractable.We show that each framework either does not logically extend to the general-purpose AI context or is infeasible in practice, primarily due to the large amounts of unstructured training data and the many potential combinations of human populations, use cases, and sensitive attributes.These inherent challenges would persist for general-purpose AI, including LLMs, even if empirical challenges, such as limited participatory input and limited measurement methods, were overcome.Nonetheless, fairness will remain an important type of model evaluation, and there are still promising research directions, particularly the development of standards for the responsibility of LLM developers, contextspecific evaluations, and methods of iterative, participatory, and AI-assisted evaluation that could scale fairness across the diverse contexts of modern human-AI interaction.
Jacy Reese Anthis, Kristian Lum, Michael D. Ekstrand, Avi Feller, Chenhao Tan
ACL (1)5
2025 Literature Meets Data: A Synergistic Approach to Hypothesis Generation
abstract
AI holds promise for transforming scientific processes, including hypothesis generation.Prior work on hypothesis generation can be broadly categorized into theory-driven and datadriven approaches.While both have proven effective in generating novel and plausible hypotheses, it remains an open question whether they can complement each other.To address this, we develop the first method that combines literature-based insights with data to perform LLM-powered hypothesis generation.We apply our method on five different datasets and demonstrate that integrating literature and data outperforms other baselines (8.97% over fewshot, 15.75% over literature-based alone, and 3.37% over data-driven alone).Additionally, we conduct the first human evaluation to assess the utility of LLM-generated hypotheses in assisting human decision-making on two challenging tasks: deception detection and AI generated content detection.Our results show that human accuracy improves significantly by 7.44% and 14.19% on these tasks, respectively.These findings suggest that integrating literature-based and data-driven approaches provides a comprehensive and nuanced framework for hypothesis generation and could open new avenues for scientific inquiry.
Haokun Liu, Yangqiaoyu Zhou, Chenfei Yuan, Chenhao Tan
ACL (1)5
2025 MoVa: Towards Generalizable Classification of Human Morals and Values
abstract
Ziyu Chen, Junfei Sun, Chenxi Li, Tuan Dung Nguyen, Jing Yao, Xiaoyuan Yi, Xing Xie, Chenhao Tan, Lexing Xie. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Junfei Sun, Tuan Dung Nguyen, Jing Yao 0003, Xiaoyuan Yi, Xing Xie 0001, Chenhao Tan, Lexing Xie
EMNLP8
2025 Concept Incongruence: An Exploration of Time and Death in Role Playing
abstract
Consider this prompt "Draw a unicorn with two horns". Should large language models (LLMs) recognize that a unicorn has only one horn by definition and ask users for clarifications, or proceed to generate something anyway? We introduce *concept incongruence* to capture such phenomena where concept boundaries clash with each other, either in user prompts or in model representations, often leading to under-specified or mis-specified behaviors. In this work, we take the first step towards defining and analyzing model behavior under concept incongruence. Focusing on temporal boundaries in the Role-Play setting, we propose three behavioral metrics---abstention rate, conditional accuracy, and answer rate---to quantify model behavior under incongruence due to the role's death. We show that models fail to abstain after death and suffer from an accuracy drop compared to the Non-Role-Play setting. Through probing experiments, we identify two main causes: (i) unreliable encoding of the "death" state across different years, leading to unsatisfactory abstention behavior, and (ii) role playing causes shifts in the model’s temporal representations, resulting in accuracy drops. We leverage these insights to improve consistency in the model's abstention and answer behaviors. Our findings suggest that concept incongruence leads to unexpected model behaviors and point to future directions on improving model behavior under concept incongruence.
Xiaoyan Bai, Ike Peng, Chenhao Tan
NeurIPS4
2025 Absence Bench: Language Models Can't See What's Missing
abstract
Large language models (LLMs) are increasingly capable of processing long inputs and locating specific information within them, as evidenced by their performance on the Needle in a Haystack (NIAH) test. However, while models excel at recalling surprising information, they still struggle to identify clearly omitted information. We introduce AbsenceBench to assesses LLMs' capacity to detect missing information across three domains: numerical sequences, poetry, and GitHub pull requests. AbsenceBench asks models to identify which pieces of a document were deliberately removed, given access to both the original and edited contexts. Despite the apparent straightforwardness of these tasks, our experiments reveal that even state-of-the-art models like Claude-3.7-Sonnet achieve only 69.6% F1-score with a modest average context length of 5K tokens. Our analysis suggests this poor performance stems from a fundamental limitation: Transformer attention mechanisms cannot easily attend to "gaps" in documents since these absences don't correspond to any specific keys that can be attended to. Overall, our results and analysis provide a case study of the close proximity of tasks where models are already superhuman (NIAH) and tasks where models breakdown unexpectedly (AbsenceBench).
Harvey Yiyun Fu, Aryan Shrivastava, Jared Moore, Peter West, Chenhao Tan, Ari Holtzman
NeurIPS5
2025 Prompting as Scientific Inquiry
abstract
Prompting is the primary method by which we study and control large language models. It is also one of the most powerful: nearly every major capability attributed to LLMs—few-shot learning, chain-of-thought, constitutional AI—was first unlocked through prompting. Yet prompting is rarely treated as science and is frequently frowned upon as alchemy. We argue that this is a category error. If we treat LLMs as a new kind of organism—complex, opaque, and trained rather than programmed—then prompting is not a workaround. It is behavioral science. Mechanistic interpretability peers into the neural substrate, prompting probes the model in its native interface: language. We argue that prompting is not inferior, but rather a key component in the science of LLMs.
Ari Holtzman, Chenhao Tan
NeurIPS2
2025 "I Cannot Write This Because It Violates Our Content Policy": Understanding Content Moderation Policies and User Experiences in Generative AI Products
Lan Gao 0001, Oscar Chen, Rachel Lee, Nick Feamster, Chenhao Tan, Marshini Chetty
USENIX Security Symposium5
2024 "Community Guidelines Make this the Best Party on the Internet": An In-Depth Study of Online Platforms' Content Moderation Policies
abstract
Moderating user-generated content on online platforms is crucial for balancing user safety and freedom of speech. Particularly in the United States, platforms are not subject to legal constraints prescribing permissible content. Each platform has thus developed bespoke content moderation policies, but there is little work towards a comparative understanding of these policies across platforms and topics. This paper presents the first systematic study of these policies from the 43 largest online platforms hosting user-generated content, focusing on policies around copyright infringement, harmful speech, and misleading content. We build a custom web-scraper to obtain policy text and develop a unified annotation scheme to analyze the text for the presence of critical components. We find significant structural and compositional variation in policies across topics and platforms, with some variation attributable to disparate legal groundings. We lay the groundwork for future studies of ever-evolving content moderation policies and their impact on users.
Brennan Schaffner, Arjun Nitin Bhagoji, Siyuan Cheng 0018, Jacqueline Mei, Jay L. Shen, Marshini Chetty, Nick Feamster, Genevieve Lakier, Chenhao Tan
CHI10
2024 The First Workshop on AI Behavioral Science
abstract
This workshop initiates a new study field which may be named AI behavioral science. It discusses recent findings, methodologies, applications, and potential societal impacts that are related to analyzing, understanding, and directing the behaviors of AI models, especially those built upon large language models. This half-day workshop includes several keynote and invited talks, a poster session, and a panel discussion.
Himabindu Lakkaraju, Qiaozhu Mei, Chenhao Tan, Jie Tang 0001, Yutong Xie 0007
KDD3
2024 Community Archetypes: An Empirical Framework for Guiding Research Methodologies to Reflect User Experiences of Sense of Virtual Community on Reddit
abstract
Humans need a sense of community (SOC), and social media platforms afford opportunities to address this need by providing users with a sense of virtual community (SOVC). This paper explores SOVC on Reddit and is motivated by two goals: (1) providing researchers with an excellent resource for methodological decisions in studies of Reddit communities; and (2) creating the foundation for a new class of research methods and community support tools that reflect users' experiences of SOVC. To ensure that methods are respectfully and ethically designed in service and accountability to impacted communities, our work takes a qualitative and community-centered approach by engaging with two key stakeholder groups. First, we interviewed 21 researchers to understand how they study community" on Reddit. Second, we surveyed 12 subreddits to gain insight into user experiences of SOVC. Results show that some research methods can broadly reflect user experiences of SOVC regardless of the topic or type of subreddit. However, user responses also evidenced the existence of five distinct Community Archetypes: Topical Q&A, Learning & Perspective Broadening, Social Support, Content Generation, and Affiliation with an Entity. We offer the Community Archetypes framework to support future work in designing methods that align more closely with user experiences of SOVC and to create community support tools that can meaningfully nourish the human need for SOC/SOVC in our modern world.
Gale H. Prinster, C. Estelle Smith, Chenhao Tan, Brian Keegan
Proc. ACM Hum. Comput. Interact.3
2024 Governance of the Black Experience on Reddit: r/BlackPeopleTwitter as a Case Study in Supporting Sense of Virtual Community for Black Users
abstract
Despite frequent efforts to combat racism, almost no research has explored how to cultivate positive experiences of thriving Black culture on Reddit. In this case study, we surveyed users of r/BlackPeopleTwitter (BPT)--a large, popular subreddit that showcases screenshots of hilarious or insightful social media posts made by Black people (mainly from Black Twitter). Our research questions seek to understand users' motivations for visiting BPT, how they experience a sense of virtual community (SOVC) and membership in BPT, and how BPT's governance influences these experiences. We find that that users come to BPT primarily for excellent humor and entertainment, sociopolitical context on issues relevant to Black people, and/or partaking in the shared Black experience. Black users are more likely to report higher SOVC and to identify as members, whereas non-Black users are more likely to identify as guests or visitors to the community. To protect Black expression, the BPT moderation team implemented a governance strategy for verifying racial identity and limiting participation to only verified users in certain threads. Our data suggest that this policy is a contentious but influential aspect of SOVC that simultaneously constructs and challenges the sense of the subreddit existing as a safe space for Black people. We synthesize these results by discussing how: differing platform affordances across Twitter and Reddit combine to cultivate a thriving Black community on Reddit; the need for Black authenticity on an otherwise anonymous platform can guide future research in identity verification; and the limitations of this study motivate future work to support all marginalized communities online.
C. Estelle Smith, Shamika Klassen, Gale H. Prinster, Chenhao Tan, Brian Keegan
Proc. ACM Hum. Comput. Interact.4
2023 Language of Bargaining
abstract
Leveraging an established exercise in negotiation education, we build a novel dataset for studying how the use of language shapes bilateral bargaining.Our dataset extends existing work in two ways: 1) we recruit participants via behavioral labs instead of crowdsourcing platforms and allow participants to negotiate through audio, enabling more naturalistic interactions; 2) we add a control setting where participants negotiate only through alternating, written numeric offers.Despite the two contrasting forms of communication, we find that the average agreed prices of the two treatments are identical.But when subjects can talk, fewer offers are exchanged, negotiations finish faster, the likelihood of reaching agreement rises, and the variance of prices at which subjects agree drops substantially.We further propose a taxonomy of speech acts in negotiation and enrich the dataset with annotated speech acts.We set up prediction tasks to predict negotiation success and find that being reactive to the arguments of the other party is advantageous over driving the negotiation.
Mourad Heddaya, Solomon Dworkin, Chenhao Tan, Rob Voigt, Alexander Zentefis
ACL (1)3
2023 FLamE: Few-shot Learning from Natural Language Explanations
abstract
Natural language explanations have the potential to provide rich information that in principle guides model reasoning.Yet, recent work by Lampinen et al. (2022) has shown limited utility of natural language explanations in improving classification.To effectively learn from explanations, we present FLamE, a two-stage few-shot learning framework that first generates explanations using GPT-3, and then finetunes a smaller model (e.g., RoBERTa) with generated explanations.Our experiments on natural language inference demonstrate effectiveness over strong baselines, increasing accuracy by 17.6% over GPT-3 Babbage and 5.7% over GPT-3 Davinci in e-SNLI.Despite improving classification performance, human evaluation surprisingly reveals that the majority of generated explanations does not adequately justify classification decisions.Additional analyses point to the important role of label-specific cues (e.g., "not know" for the neutral label) in generated explanations.
Yangqiaoyu Zhou, Yiming Zhang 0022, Chenhao Tan
ACL (1)3
2023 Learning to Ignore Adversarial Attacks
abstract
Despite the strong performance of current NLP models, they can be brittle against adversarial attacks.To enable effective learning against adversarial inputs, we introduce the use of rationale models that can explicitly learn to ignore attack tokens.We find that the rationale models can successfully ignore over 90% of attack tokens.This approach leads to consistent and sizable improvements (∼10%) over baseline models in robustness on three datasets for both BERT and RoBERTa, and also reliably outperforms data augmentation with adversarial examples alone.In many cases, we find that our method is able to close the gap between model performance on a clean test set and an attacked test set and hence reduce the effect of adversarial attacks.
Yiming Zhang 0022, Yangqiaoyu Zhou, Samuel Carton, Chenhao Tan
EACL4
2023 Learning Human-Compatible Representations for Case-Based Decision Support
Yizhou Tian, Chacha Chen, Shi Feng 0005, Yuxin Chen 0001, Chenhao Tan
ICLR6
2023 Language Models Can Improve Event Prediction by Few-Shot Abductive Reasoning
abstract
Large language models have shown astonishing performance on a wide range of reasoning tasks. In this paper, we investigate whether they could reason about real-world events and help improve the prediction performance of event sequence models. We design LAMP, a framework that integrates a large language model in event prediction. Particularly, the language model performs abductive reasoning to assist an event sequence model: the event model proposes predictions on future events given the past; instructed by a few expert-annotated demonstrations, the language model learns to suggest possible causes for each proposal; a search module finds out the previous events that match the causes; a scoring function learns to examine whether the retrieved events could actually cause the proposal. Through extensive experiments on several challenging real-world datasets, we demonstrate that our framework---thanks to the reasoning capabilities of large language models---could significantly outperform the state-of-the-art event sequence models.
Xiaoming Shi 0001, Siqiao Xue, Kangrui Wang, Fan Zhou 0012, James Y. Zhang, Jun Zhou 0011, Chenhao Tan, Hongyuan Mei
NeurIPS7
2023 Selective Explanations: Leveraging Human Input to Align Explainable AI
abstract
While a vast collection of explainable AI (XAI) algorithms has been developed in recent years, they have been criticized for significant gaps with how humans produce and consume explanations. As a result, current XAI techniques are often found to be hard to use and lack effectiveness. In this work, we attempt to close these gaps by making AI explanations selective ---a fundamental property of human explanations---by selectively presenting a subset of model reasoning based on what aligns with the recipient's preferences. We propose a general framework for generating selective explanations by leveraging human input on a small dataset. This framework opens up a rich design space that accounts for different selectivity goals, types of input, and more. As a showcase, we use a decision-support task to explore selective explanations based on what the decision-maker would consider relevant to the decision task. We conducted two experimental studies to examine three paradigms based on our proposed framework: in Study 1, we ask the participants to provide critique-based or open-ended input to generate selective explanations (self-input). In Study 2, we show the participants selective explanations based on input from a panel of similar users (annotator input). Our experiments demonstrate the promise of selective explanations in reducing over-reliance on AI and improving collaborative decision making and subjective perceptions of the AI system, but also paint a nuanced picture that attributes some of these positive effects to the opportunity to provide one's own input to augment AI explanations. Overall, our work proposes a novel XAI framework inspired by human communication behaviors and demonstrates its potential to encourage future work to make AI explanations more human-compatible.
Vivian Lai, Yiming Zhang 0022, Chacha Chen, Qingzi Vera Liao, Chenhao Tan
Proc. ACM Hum. Comput. Interact.5
2023 How Large Language Models Will Disrupt Data Management
abstract
Large language models (LLMs), such as GPT-4, are revolutionizing software's ability to understand, process, and synthesize language. The authors of this paper believe that this advance in technology is significant enough to prompt introspection in the data management community, similar to previous technological disruptions such as the advents of the world wide web, cloud computing, and statistical machine learning. We argue that the disruptive influence that LLMs will have on data management will come from two angles. (1) A number of hard database problems, namely, entity resolution, schema matching, data discovery, and query synthesis, hit a ceiling of automation because the system does not fully understand the semantics of the underlying data. Based on large training corpora of natural language, structured data, and code, LLMs have an unprecedented ability to ground database tuples, schemas, and queries in real-world concepts. We will provide examples of how LLMs may completely change our approaches to these problems. (2) LLMs blur the line between predictive models and information retrieval systems with their ability to answer questions. We will present examples showing how large databases and information retrieval systems have complementary functionality.
Raul Castro Fernandez, Aaron J. Elmore, Michael J. Franklin, Sanjay Krishnan, Chenhao Tan
Proc. VLDB Endow.5
2022 Human-AI Collaboration via Conditional Delegation: A Case Study of Content Moderation
abstract
Despite impressive performance in many benchmark datasets, AI models can still make mistakes, especially among out-of-distribution examples. It remains an open question how such imperfect models can be used effectively in collaboration with humans. Prior work has focused on AI assistance that helps people make individual high-stakes decisions, which is not scalable for a large amount of relatively low-stakes decisions, e.g., moderating social media comments. Instead, we propose conditional delegation as an alternative paradigm for human-AI collaboration where humans create rules to indicate trustworthy regions of a model. Using content moderation as a testbed, we develop novel interfaces to assist humans in creating conditional delegation rules and conduct a randomized experiment with two datasets to simulate in-distribution and out-of-distribution scenarios. Our study demonstrates the promise of conditional delegation in improving model performance and provides insights into design for this novel paradigm, including the effect of AI explanations.
Vivian Lai, Samuel Carton, Rajat Bhatnagar, Qingzi Vera Liao, Chenhao Tan
CHI6
2022 Active Example Selection for In-Context Learning
abstract
With a handful of demonstration examples, large-scale language models show strong capability to perform various tasks by in-context learning from these examples, without any finetuning.We demonstrate that in-context learning performance can be highly unstable across samples of examples, indicating the idiosyncrasies of how language models acquire information.We formulate example selection for in-context learning as a sequential decision problem, and propose a reinforcement learning algorithm for identifying generalizable policies to select demonstration examples.For GPT-2, our learned policies demonstrate strong abilities of generalizing to unseen tasks in training, with a 5.8% improvement on average.Examples selected from our learned policies can even achieve a small improvement on GPT-3 Ada.However, the improvement diminishes on larger GPT-3 models, suggesting emerging capabilities of large language models.
Yiming Zhang 0022, Shi Feng 0005, Chenhao Tan
EMNLP3
2022 Explaining Why: How Instructions and User Interfaces Impact Annotator Rationales When Labeling Text Data
abstract
Jamar Sullivan Jr., Will Brackenbury, Andrew McNutt, Kevin Bryson, Kwam Byll, Yuxin Chen, Michael Littman, Chenhao Tan, Blase Ur. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Jamar L. Sullivan Jr., Will Brackenbury, Andrew McNut, Kevin Bryson 0002, Kwam Byll, Yuxin Chen 0001, Michael L. Littman, Chenhao Tan, Blase Ur
NAACL-HLT8
2022 On the Diversity and Limits of Human Explanations
abstract
A growing effort in NLP aims to build datasets of human explanations.However, it remains unclear whether these datasets serve their intended goals.This problem is exacerbated by the fact that the term explanation is overloaded and refers to a broad range of notions with different properties and ramifications.Our goal is to provide an overview of the diversity of explanations, discuss human limitations in providing explanations, and ultimately provide implications for collecting and using human explanations in NLP.Inspired by prior work in psychology and cognitive sciences, we group existing human explanations in NLP into three categories: proximal mechanism, evidence, and procedure.These three types differ in nature and have implications for the resultant explanations.For instance, procedure is not considered explanation in psychology and connects with a rich body of work on learning from instructions.The diversity of explanations is further evidenced by proxy questions that are needed for annotators to interpret and answer "why is [input] assigned [label]".Finally, giving explanations may require different, often deeper, understandings than predictions, which casts doubt on whether humans can provide valid explanations in some tasks.
Chenhao Tan
NAACL-HLT1
2022 Probing Classifiers are Unreliable for Concept Removal and Detection
abstract
Neural network models trained on text data have been found to encode undesirable linguistic or sensitive concepts in their representation. Removing such concepts is non-trivial because of a complex relationship between the concept, text input, and the learnt representation. Recent work has proposed post-hoc and adversarial methods to remove such unwanted concepts from a model's representation. Through an extensive theoretical and empirical analysis, we show that these methods can be counter-productive: they are unable to remove the concepts entirely, and in the worst case may end up destroying all task-relevant features. The reason is the methods' reliance on a probing classifier as a proxy for the concept. Even under the most favorable conditions for learning a probing classifier when a concept's relevant features in representation space alone can provide 100% accuracy, we prove that a probing classifier is likely to use non-concept features and thus post-hoc or adversarial methods will fail to remove the concept correctly. These theoretical implications are confirmed by experiments on models trained on synthetic, Multi-NLI, and Twitter datasets. For sensitive applications of concept removal such as fairness, we recommend caution against using these methods and propose a spuriousness metric to gauge the quality of the final classifier.
Abhinav Kumar 0001, Chenhao Tan, Amit Sharma 0007
NeurIPS2
2022 The Impact of Governance Bots on Sense of Virtual Community: Development and Validation of the GOV-BOTs Scale
abstract
Bots are increasingly being used for governance-related purposes in online communities, yet no instrumentation exists for measuring how users assess their beneficial or detrimental impacts. In order to support future human-centered and community-based research, we developed a new scale called GOVernance Bots in Online communiTies (GOV-BOTs) across two rounds of surveys on Reddit (N=820). We applied rigorous psychometric criteria to demonstrate the validity of GOV-BOTs, which contains two subscales: bot governance (4 items) and bot tensions (3 items). Whereas humans have historically expected communities to be composed entirely of humans, the social participation of bots as non-human agents now raises fundamental questions about psychological, philosophical, and ethical implications. Addressing psychological impacts, our data show that perceptions of effective bot governance positively contribute to users' sense of virtual community (SOVC), whereas perceived bot tensions may only impact SOVC if users are more aware of bots. Finally, we show that users tend to experience the greatest SOVC across groups of subreddits, rather than individual subreddits, suggesting that future research should carefully re-consider uses and operationalizations of the term "community."
C. Estelle Smith, Irfanul Alam, Chenhao Tan, Brian Keegan, Anita L. Blanchard
Proc. ACM Hum. Comput. Interact.3
2021 Using AI to Promote Equitable Classroom Discussions: The TalkMoves Application
Abhijit Suresh, Jennifer Jacobs 0002, Charis Clevenger, Vivian Lai, Chenhao Tan, James H. Martin, Tamara Sumner
AIED (2)5
2021 Towards Unifying Feature Attribution and Counterfactual Explanations: Different Means to the Same End
abstract
Feature attributions and counterfactual explanations are popular approaches to explain a ML model. The former assigns an importance score to each input feature, while the latter provides input examples with minimal changes to alter the model's predictions. To unify these approaches, we provide an interpretation based on the actual causality framework and present two key results in terms of their use. First, we present a method to generate feature attribution explanations from a set of counterfactual examples. These feature attributions convey how important a feature is to changing the classification outcome of a model, especially on whether a subset of features is necessary and/or sufficient for that change, which attribution-based methods are unable to provide. Second, we show how counterfactual examples can be used to evaluate the goodness of an attribution-based explanation in terms of its necessity and sufficiency. As a result, we highlight the complimentary of these two approaches. Our evaluation on three benchmark datasets --- Adult-Income, LendingClub, and German-Credit --- confirms the complimentary. Feature attribution methods like LIME and SHAP and counterfactual explanation methods like Wachter et al. and DiCE often do not agree on feature importance rankings. In addition, by restricting the features that can be modified for generating counterfactual examples, we find that the top-k features from LIME or SHAP are often neither necessary nor sufficient explanations of a model's prediction. Finally, we present a case study of different explanation methods on a real-world hospital triage problem.
Ramaravind Kommiya Mothilal, Divyat Mahajan, Chenhao Tan, Amit Sharma 0007
AIES3
2021 Decision-Focused Summarization
abstract
Relevance in summarization is typically de- fined based on textual information alone, without incorporating insights about a particular decision. As a result, to support risk analysis of pancreatic cancer, summaries of medical notes may include irrelevant information such as a knee injury. We propose a novel problem, decision-focused summarization, where the goal is to summarize relevant information for a decision. We leverage a predictive model that makes the decision based on the full text to provide valuable insights on how a decision can be inferred from text. To build a summary, we then select representative sentences that lead to similar model decisions as using the full text while accounting for textual non-redundancy. To evaluate our method (DecSum), we build a testbed where the task is to summarize the first ten reviews of a restaurant in support of predicting its future rating on Yelp. DecSum substantially outperforms text-only summarization methods and model-based explanation methods in decision faithfulness and representativeness. We further demonstrate that DecSum is the only method that enables humans to outperform random chance in predicting which restaurant will be better rated in the future.
Chao-Chun Hsu, Chenhao Tan
EMNLP (1)2
2021 Understanding the Diverging User Trajectories in Highly-related Online Communities during the COVID-19 Pandemic
Jason Shuo Zhang, Brian Keegan, Qin Lv, Chenhao Tan
ICWSM4
2021 Understanding the Effect of Out-of-distribution Examples and Interactive Explanations on Human-AI Decision Making
abstract
Although AI holds promise for improving human decision making in societally critical domains, it remains an open question how human-AI teams can reliably outperform AI alone and human alone in challenging prediction tasks (also known as complementary performance). We explore two directions to understand the gaps in achieving complementary performance. First, we argue that the typical experimental setup limits the potential of human-AI teams. To account for lower AI performance out-of-distribution than in-distribution because of distribution shift, we design experiments with different distribution types and investigate human performance for both in-distribution and out-of-distribution examples. Second, we develop novel interfaces to support interactive explanations so that humans can actively engage with AI assistance. Using virtual pilot studies and large-scale randomized experiments across three tasks, we demonstrate a clear difference between in-distribution and out-of-distribution, and observe mixed results for interactive explanations: while interactive explanations improve human perception of AI assistance's usefulness, they may reinforce human biases and lead to limited performance improvement. Overall, our work points out critical challenges and future directions towards enhancing human performance with AI assistance.
Vivian Lai, Chenhao Tan
Proc. ACM Hum. Comput. Interact.3
2020 "Why is 'Chicago' deceptive?" Towards Building Model-Driven Tutorials for Humans
abstract
To support human decision making with machine learning models, we often need to elucidate patterns embedded in the models that are unsalient, unknown, or counterintuitive to humans. While existing approaches focus on explaining machine predictions with real-time assistance, we explore model-driven tutorials to help humans understand these patterns in a train- ing phase. We consider both tutorials with guidelines from scientific papers, analogous to current practices of science communication, and automatically selected examples from training data with explanations. We use deceptive review detection as a testbed and conduct large-scale, randomized human-subject experiments to examine the effectiveness of such tutorials. We find that tutorials indeed improve human performance, with and without real-time assistance. In particular, although deep learning provides superior predictive performance than simple models, tutorials and explanations from simple models are more useful to humans. Our work suggests future directions for human-centered tutorials and explanations towards a synergy between humans and AI.
Vivian Lai, Chenhao Tan
CHI3
2020 Evaluating and Characterizing Human Rationales
abstract
Two main approaches for evaluating the quality of machine-generated rationales are: 1) using human rationales as a gold standard; and 2) automated metrics based on how rationales affect model behavior.An open question, however, is how human rationales fare with these automatic metrics.Analyzing a variety of datasets and models, we find that human rationales do not necessarily perform well on these metrics.To unpack this finding, we propose improved metrics to account for modeldependent baseline performance.We then propose two methods to further characterize rationale quality, one based on model retraining and one on using "fidelity curves" to reveal properties such as irrelevance and redundancy.Our work leads to actionable suggestions for evaluating and characterizing rationales.
Samuel Carton, Anirudh Rathore, Chenhao Tan
EMNLP (1)3
2019 What Gets Echoed? Understanding the "Pointers" in Explanations of Persuasive Arguments
abstract
David Atkinson, Kumar Bhargav Srinivasan, Chenhao Tan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
David Atkinson, Kumar Bhargav Srinivasan, Chenhao Tan
EMNLP/IJCNLP (1)3
2019 Many Faces of Feature Importance: Comparing Built-in and Post-hoc Feature Importance in Text Classification
abstract
Vivian Lai, Zheng Cai, Chenhao Tan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Vivian Lai, Zheng Cai, Chenhao Tan
EMNLP/IJCNLP (1)3
2019 Ask not what AI can do, but what AI should do: Towards a framework of task delegability
abstract
While artificial intelligence (AI) holds promise for addressing societal challenges, issues of exactly which tasks to automate and to what extent to do so remain understudied. We approach this problem of task delegability from a human-centered perspective by developing a framework on human perception of task delegation to AI. We consider four high-level factors that can contribute to a delegation decision: motivation, difficulty, risk, and trust. To obtain an empirical understanding of human preferences in different tasks, we build a dataset of 100 tasks from academic papers, popular media portrayal of AI, and everyday life, and administer a survey based on our proposed framework. We find little preference for full AI control and a strong preference for machine-in-the-loop designs, in which humans play the leading role. Among the four factors, trust is the most correlated with human preferences of optimal human-machine delegation. This framework represents a first step towards characterizing human preferences of AI automation across tasks. We hope this work encourages future efforts towards understanding such individual attitudes; our goal is to inform the public and the AI research community rather than dictating any direction in technology development.
Brian Lubars, Chenhao Tan
NeurIPS2
2019 What Makes a Good Team? A Large-scale Study on the Effect of Team Composition in Honor of Kings
abstract
Team composition is a central factor in determining the effectiveness of a team. In this paper, we present a large-scale study on the effect of team composition on multiple measures of team effectiveness. We use a dataset from the largest multiplayer online battle arena (MOBA) game, Honor of Kings, with 96 million matches involving 100 million players. We measure team effectiveness based on team performance (whether a team is going to win), team tenacity (whether a team is going to surrender), and team rapport (whether a team uses abusive language). Our results confirm the importance of team diversity with respect to player roles, and show that diversity has varying effects on team effectiveness: although diverse teams perform well and show tenacity in adversity, they are more likely to abuse when losing than less diverse teams. Our study also contributes to the situation vs. personality debate and show that abusive players tend to choose the leading role and players do not become more abusive when taking such roles.
Ziqiang Cheng, Yang Yang 0009, Chenhao Tan, Denny Cheng, Yueting Zhuang
WWW3
2019 Are All Successful Communities Alike? Characterizing and Predicting the Success of Online Communities
abstract
The proliferation of online communities has created exciting opportunities to study the mechanisms that explain group success. While a growing body of research investigates community success through a single measure - typically, the number of members - we argue that there are multiple ways of measuring success. Here, we present a systematic study to understand the relations between these success definitions and test how well they can be predicted based on community properties and behaviors from the earliest period of a community's lifetime. We identify four success measures that are desirable for most communities: (i) growth in the number of members; (ii) retention of members; (iii) long term survival of the community; and (iv) volume of activities within the community. Surprisingly, we find that our measures do not exhibit very high correlations, suggesting that they capture different types of success. Additionally, we find that different success measures are predicted by different attributes of online communities, suggesting that success can be achieved through different behaviors. Our work sheds light on the basic understanding on what success represents in online communities and what predicts it. Our results suggest that success is multi-faceted and cannot be measured nor predicted by a single measurement. This insight has practical implications for the creation of new online communities and the design of platforms that facilitate such communities.
David Jurgens, Chenhao Tan, Daniel M. Romero
WWW3
2019 Long-term prediction of time series based on stepwise linear division algorithm and time-variant zonary fuzzy information granules
Chao Luo 0001, Chenhao Tan, Yuanjie Zheng
Int. J. Approx. Reason.2
2019 Content Removal as a Moderation Strategy: Compliance and Other Outcomes in the ChangeMyView Community
abstract
Moderators of online communities often employ comment deletion as a tool. We ask here whether, beyond the positive effects of shielding a community from undesirable content, does comment removal actually cause the behavior of the comment's author to improve? We examine this question in a particularly well-moderated community, the ChangeMyView subreddit. The standard analytic approach of interrupted time-series analysis unfortunately cannot answer this question of causality because it fails to distinguish the effect of having made a non-compliant comment from the effect of being subjected to moderator removal of that comment. We therefore leverage a "delayed feedback" approach based on the observation that some users may remain active between the time when they posted the non-compliant comment and the time when that comment is deleted. Applying this approach to such users, we reveal the causal role of comment deletion in reducing immediate noncompliance rates, although we do not find evidence of it having a causal role in inducing other behavior improvements. Our work thus empirically demonstrates both the promise and some potential limits of content removal as a positive moderation strategy, and points to future directions for identifying causal effects from observational data.
Kumar Bhargav Srinivasan, Cristian Danescu-Niculescu-Mizil, Lillian Lee, Chenhao Tan
Proc. ACM Hum. Comput. Interact.4
2019 Intergroup Contact in the Wild: Characterizing Language Differences between Intergroup and Single-group Members in NBA-related Discussion Forums
abstract
Intergroup contact has long been considered as an effective strategy to reduce prejudice between groups. However, recent studies suggest that exposure to opposing groups in online platforms can exacerbate polarization. To further understand the behavior of individuals who actively engage in intergroup contact in practice, we provide a large-scale observational study of intragroup behavioral differences between members with and without intergroup contact. We leverage the existing structure of NBA-related discussion forums on Reddit to study the context of professional sports. We identify fans of each NBA team as members of a group and trace whether they have intergroup contact. Our results show that members with intergroup contact use more negative and abusive language in their affiliated group than those without such contact, after controlling for activity levels. We further quantify different levels of intergroup contact and show that there may exist nonlinear mechanisms regarding how intergroup contact relates to intragroup behavior. Our findings provide complementary evidence to experimental studies in a novel context and also shed light on possible reasons for the different outcomes in prior studies.
Jason Shuo Zhang, Chenhao Tan, Qin Lv
Proc. ACM Hum. Comput. Interact.2
2019 Measuring Online Debaters' Persuasive Skill from Text over Time
abstract
Online debates allow people to express their persuasive abilities and provide exciting opportunities for understanding persuasion. Prior studies have focused on studying persuasion in debate content, but without accounting for each debater’s history or exploring the progression of a debater’s persuasive ability. We study debater skill by modeling how participants progress over time in a collection of debates from Debate.org . We build on a widely used model of skill in two-player games and augment it with linguistic features of a debater’s content. We show that online debaters’ skill levels do tend to improve over time. Incorporating linguistic profiles leads to more robust skill estimation than winning records alone. Notably, we find that an interaction feature combining uncertainty cues (hedging) with terms strongly associated with either side of a particular debate (fightin’ words) is more predictive than either feature on its own, indicating the importance of fine- grained linguistic features.
Kelvin Luu, Chenhao Tan, Noah A. Smith
Trans. Assoc. Comput. Linguistics2
2018 Urban Dreams of Migrants: A Case Study of Migrant Integration in Shanghai
abstract
Unprecedented human mobility has driven the rapid urbanization around the world. In China, the fraction of population dwelling in cities increased from 17.9% to 52.6% between 1978 and 2012. Such large-scale migration poses challenges for policymakers and important questions for researchers. To investigate the process of migrant integration, we employ a one-month complete dataset of telecommunication metadata in Shanghai with 54 million users and 698 million call logs. We find systematic differences between locals and migrants in their mobile communication networks and geographical locations. For instance, migrants have more diverse contacts and move around the city with a larger radius than locals after they settle down. By distinguishing new migrants (who recently moved to Shanghai) from settled migrants (who have been in Shanghai for a while), we demonstrate the integration process of new migrants in their first three weeks. Moreover, we formulate classification problems to predict whether a person is a migrant. Our classifier is able to achieve an F1-score of 0.82 when distinguishing settled migrants from locals, but it remains challenging to identify new migrants because of class imbalance. This classification setup holds promise for identifying new migrants who will successfully integrate into locals (new migrants that misclassified as locals).
Yang Yang 0009, Chenhao Tan, Zongtao Liu, Fei Wu 0001, Yueting Zhuang
AAAI2
2018 Neural Models for Documents with Metadata
abstract
Most real-world document collections involve various types of metadata, such as author, source, and date, and yet the most commonly-used approaches to modeling text corpora ignore this information.While specialized models have been developed for particular applications, few are widely used in practice, as customization typically requires derivation of a custom inference algorithm.In this paper, we build on recent advances in variational inference methods and propose a general neural framework, based on topic models, to enable flexible incorporation of metadata and allow for rapid exploration of alternative models.Our approach achieves strong performance, with a manageable tradeoff between perplexity, coherence, and sparsity.Finally, we demonstrate the potential of our framework through an exploration of a corpus of articles about US immigration.
Dallas Card, Chenhao Tan, Noah A. Smith
ACL (1)2
2018 Tracing Community Genealogy: How New Communities Emerge from the Old
Chenhao Tan
ICWSM1
2018 Creative Writing with a Machine in the Loop: Case Studies on Slogans and Stories
abstract
As the quality of natural language generated by artificial intelligence systems improves, writing interfaces can support interventions beyond grammar-checking and spell-checking, such as suggesting content to spark new ideas. To explore the possibility of machine-in-the-loop creative writing, we performed two case studies using two system prototypes, one for short story writing and one for slogan writing. Participants in our studies were asked to write with a machine in the loop or alone (control condition). They assessed their writing and experience through surveys and an open-ended interview. We collected additional assessments of the writing from Amazon Mechanical Turk crowdworkers. Our findings indicate that participants found the process fun and helpful and could envision use cases for future systems. At the same time, machine suggestions do not necessarily lead to better written artifacts. We therefore suggest novel natural language models and design choices that may better support creative writing.
Elizabeth Clark, Anne Spencer Ross, Chenhao Tan, Yangfeng Ji, Noah A. Smith
IUI3
2018 To Stay or to Leave: Churn Prediction for Urban Migrants in the Initial Period
abstract
In China, 260 million people migrate to cities to realize their urban dreams. Despite that these migrants play an important role in the rapid urbanization process, many of them fail to settle down and eventually leave the city. The integration process of migrants thus raises an important issue for scholars and policymakers. In this paper, we use Shanghai as an example to investigate migrants' behavior in their first weeks and in particular, how their behavior relates to early departure. Our dataset consists of a one-month complete dataset of 698 telecommunication logs between 54 million users, plus a novel and publicly available housing price data for 18K real estates in Shanghai. We find that migrants who end up leaving early tend to neither develop diverse connections in their first weeks nor move around the city. Their active areas also have higher housing prices than that of staying migrants. We formulate a churn prediction problem to determine whether a migrant is going to leave based on her behavior in the first few days. The prediction performance improves as we include data from more days. Interestingly, when using the same features, the classifier trained from only the first few days is already as good as the classifier trained using full data, suggesting that the performance difference mainly lies in the difference between features.
Yang Yang 0009, Zongtao Liu, Chenhao Tan, Fei Wu 0001, Yueting Zhuang
WWW3
2018 "You are no Jack Kennedy": On Media Selection of Highlights from Presidential Debates
abstract
Political speeches and debates play an important role in shaping the images of politicians, and the public often relies on media outlets to select bits of political communication from a large pool of utterances. It is an important research question to understand what factors impact this selection process. To quantitatively explore the selection process, we build a three- decade dataset of presidential debate transcripts and post-debate coverage. We first examine the effect of wording and propose a binary classification framework that controls for both the speaker and the debate situation. We find that crowdworkers can only achieve an accuracy of 60% in this task, indicating that media choices are not entirely obvious. Our classifiers outperform crowdworkers on average, mainly in primary debates. We also compare important factors from crowdworkers» free-form explanations with those from data-driven methods and find interesting differences. Few crowdworkers mentioned that "context matters", whereas our data show that well-quoted sentences are more distinct from the previous utterance by the same speaker than less-quoted sentences. Finally, we examine the aggregate effect of media preferences towards different wordings to understand the extent of fragmentation among media outlets. By analyzing a bipartite graph built from quoting behavior in our data, we observe a decreasing trend in bipartisan coverage.
Chenhao Tan, Hao Peng 0009, Noah A. Smith
WWW1
2018 Framing Effects: Choice of Slogans Used to Advertise Online Experiments Can Boost Recruitment and Lead to Sample Biases
abstract
Online experimentation with volunteers relies on participants' non-financial motivations to complete a study, such as to altruistically support science or to compare oneself to others. Researchers rely on these motivations to attract study participants and often use incentives, like performance comparisons, to encourage participation. Often, these study incentives are advertised using a slogan (e.g., "What is your thinking style?''). Research on framing effects suggests that advertisement slogans attract people with varying demographics and motivations. Could the slogan advertisements for studies risk attracting only specific users? To investigate the existence of potential sample biases, we measured how different slogan frames affected which participants self-selected into studies. We found that slogan frames impact recruitment significantly; changing the slogan frame from a 'supporting science' frame to a 'comparing oneself to others' frame lead to a 9% increase in recruitment for some studies. Additionally, slogans framed as learning more about oneself attract participants significantly more motivated by boredom compared to other slogan frames. We discuss design implications for using frames to improve recruitment and mitigate sources of sample bias in online research with volunteers.
Tal August, Nigini Oliveira, Chenhao Tan, Noah A. Smith, Katharina Reinecke
Proc. ACM Hum. Comput. Interact.3
2018 "This is why we play": Characterizing Online Fan Communities of the NBA Teams
abstract
Professional sports constitute an important part of people's modern life. People spend substantial amounts of time and money supporting their favorite players and teams, and sometimes even riot after games. However, how team performance affects fan behavior remains understudied at a large scale. As almost every notable professional team has its own online fan community, these communities provide great opportunities for investigating this research question. In this work, we provide the first large-scale characterization of online fan communities of professional sports teams. Since user behavior in these online fan communities is inherently connected to game events and team performance, we construct a unique dataset that combines 1.5M posts and 43M comments in NBA-related communities on Reddit with statistics that document team performance in the NBA. We analyze the impact of team performance on fan behavior both at the game level and the season level. First, we study how team performance in a game relates to user activity during that game. We find that surprise plays an important role: the fans of the top teams are more active when their teams lose and so are the fans of the bottom teams in an unexpected win. Second, we study fan behavior over consecutive seasons and show that strong team performance is associated with fans of low loyalty, likely due to "bandwagon fans." Fans of the bottom teams tend to discuss their team's future such as young talents in the roster, which may help them stay optimistic during adversity. Our results not only contribute to understanding the interplay between online sports communities and offline context but also provide significant insights into sports management.
Jason Shuo Zhang, Chenhao Tan, Qin Lv
Proc. ACM Hum. Comput. Interact.2
2017 Friendships, Rivalries, and Trysts: Characterizing Relations between Ideas in Texts
abstract
Understanding how ideas relate to each other is a fundamental question in many domains, ranging from intellectual history to public communication.Because ideas are naturally embedded in texts, we propose the first framework to systematically characterize the relations between ideas based on their occurrence in a corpus of documents, independent of how these ideas are represented.Combining two statistics-cooccurrence within documents and prevalence correlation over time-our approach reveals a number of different ways in which ideas can cooperate and compete.For instance, two ideas can closely track each other's prevalence over time, and yet rarely cooccur, almost like a "cold war" scenario.We observe that pairwise cooccurrence and prevalence correlation exhibit different distributions.We further demonstrate that our approach is able to uncover intriguing relations between ideas through in-depth case studies on news articles and research papers.
Chenhao Tan, Dallas Card, Noah A. Smith
ACL (1)1
2017 Dynamic Entity Representations in Neural Language Models
abstract
Understanding a long document requires tracking how entities are introduced and evolve over time.We present a new type of language model, ENTITYNLM, that can explicitly model entities, dynamically update their representations, and contextually generate their mentions.Our model is generative and flexible; it can model an arbitrary number of entities in context while generating each entity mention at an arbitrary length.In addition, it can be used for several different tasks such as language modeling, coreference resolution, and entity prediction.Experimental results with all these tasks demonstrate that our model consistently outperforms strong baselines and prior work.
Yangfeng Ji, Chenhao Tan, Sebastian Martschat, Yejin Choi 0001, Noah A. Smith
EMNLP2
2016 Science, AskScience, and BadScience: On the Coexistence of Highly Related Communities
Jack Hessel, Chenhao Tan, Lillian Lee
ICWSM2
2016 Lost in Propagation? Unfolding News Cycles from the Source
Chenhao Tan, Adrien Friggeri, Lada A. Adamic
ICWSM1
2016 Internet Collaboration on Extremely Difficult Problems: Research versus Olympiad Questions on the Polymath Site
abstract
Despite the existence of highly successful Internet collaborations on complex projects, including open-source software, little is known about how Internet collaborations work for solving "extremely" difficult problems, such as open-ended research questions. We quantitatively investigate a series of efforts known as the Polymath projects, which tackle mathematical research problems through open online discussion. A key analytical insight is that we can contrast the polymath projects with mini-polymaths -- spinoffs that were conducted in the same manner as the polymaths but aimed at addressing math Olympiad questions, which, while quite difficult, are known to be feasible. Our comparative analysis shifts between three elements of the projects: the roles and relationships of the authors, the temporal dynamics of how the projects evolved, and the linguistic properties of the discussions themselves. We find interesting differences between the two domains through each of these analyses, and present these analyses as a template to facilitate comparison between Polymath and other domains for collaboration and communication. We also develop models that have strong performance in distinguishing research-level comments based on any of our groups of features. Finally, we examine whether comments representing research breakthroughs can be recognized more effectively based on their intrinsic features, or by the (re-)actions of others, and find good predictive power in linguistic features.
Isabel M. Kloumann, Chenhao Tan, Jon M. Kleinberg, Lillian Lee
WWW2
2016 Winning Arguments: Interaction Dynamics and Persuasion Strategies in Good-faith Online Discussions
abstract
Changing someone's opinion is arguably one of the most important challenges of social interaction. The underlying process proves difficult to study: it is hard to know how someone's opinions are formed and whether and how someone's views shift. Fortunately, ChangeMyView, an active community on Reddit, provides a platform where users present their own opinions and reasoning, invite others to contest them, and acknowledge when the ensuing discussions change their original views. In this work, we study these interactions to understand the mechanisms behind persuasion.
Chenhao Tan, Vlad Niculae, Cristian Danescu-Niculescu-Mizil, Lillian Lee
WWW1
2015 All Who Wander: On the Prevalence and Characteristics of Multi-community Engagement
abstract
Although analyzing user behavior within individual communities is an active and rich research domain, people usually interact with multiple communities both on- and off-line. How do users act in such multi-community environments? Although there are a host of intriguing aspects to this question, it has received much less attention in the research community in comparison to the intra-community case. In this paper, we examine three aspects of multi-community engagement: the sequence of communities that users post to, the language that users employ in those communities, and the feedback that users receive, using longitudinal posting behavior on Reddit as our main data source, and DBLP for auxiliary experiments. We also demonstrate the effectiveness of features drawn from these aspects in predicting users' future level of activity. One might expect that a user's trajectory mimics the "settling-down" process in real life: an initial exploration of sub-communities before settling down into a few niches. However, we find that the users in our data continually post in new communities; moreover, as time goes on, they post increasingly evenly among a more diverse set of smaller communities. Interestingly, it seems that users that eventually leave the community are "destined" to do so from the very beginning, in the sense of showing significantly different "wandering" patterns very early on in their trajectories; this finding has potentially important design implications for community maintainers. Our multi-community perspective also allows us to investigate the "situation vs. personality" debate from language usage across different communities.
Chenhao Tan, Lillian Lee
WWW1
2014 The effect of wording on message propagation: Topic- and author-controlled natural experiments on Twitter
abstract
Consider a person trying to spread an important message on a social network.He/she can spend hours trying to craft the message.Does it actually matter?While there has been extensive prior work looking into predicting popularity of socialmedia content, the effect of wording per se has rarely been studied since it is often confounded with the popularity of the author and the topic.To control for these confounding factors, we take advantage of the surprising fact that there are many pairs of tweets containing the same url and written by the same user but employing different wording.Given such pairs, we ask: which version attracts more retweets?This turns out to be a more difficult task than predicting popular topics.Still, humans can answer this question better than chance (but far from perfectly), and the computational methods we develop can do better than both an average human and a strong competing method trained on noncontrolled data.
Chenhao Tan, Lillian Lee, Bo Pang 0001
ACL (1)1
2013 Instant foodie: predicting expert ratings from grassroots
abstract
Consumer review sites and recommender systems typically rely on a large volume of user-contributed ratings, which makes rating acquisition an essential component in the design of such systems. User ratings are then summarized to provide an aggregate score representing a popular evaluation of an item. An inherent problem in such summarization is potential bias due to raters self-selection and heterogeneity in terms of experience, tastes and rating scale interpretation. There are two major approaches to collecting ratings, which have different advantages and disadvantages. One is to allow a large number of volunteers to choose and rate items directly (a method employed by e.g. Yelp and Google Places). Alternatively, a panel of raters may be maintained and invited to rate a predefined set of items at regular intervals (such as in Zagat Survey). The latter approach arguably results in more consistent reviews and reduced selection bias, however, at the expense of much smaller coverage (fewer rated items).
Chenhao Tan, Ed H. Chi, David A. Huffaker, Gueorgi Kossinets, Alexander J. Smola
CIKM1
2013 On the Interplay between Social and Topical Structure
Daniel M. Romero, Chenhao Tan, Johan Ugander
ICWSM2
2013 Query-dependent cross-domain ranking in heterogeneous network
Bo Wang 0022, Jie Tang 0001, Wei Fan 0001, Songcan Chen, Chenhao Tan
Knowl. Inf. Syst.5
2012 To each his own: personalized content selection based on text comprehensibility
abstract
Imagine a physician and a patient doing a search on antibiotic resistance. Or a chess amateur and a grandmaster conducting a search on Alekhine's Defence. Although the topic is the same, arguably the two users in each case will satisfy their information needs with very different texts. Yet today search engines mostly adopt the one-size-fits-all solution, where personalization is restricted to topical preference. We found that users do not uniformly prefer simple texts, and that the text comprehensibility level should match the user's level of preparedness. Consequently, we propose to model the comprehensibility of texts as well as the users' reading proficiency in order to better explain how different users choose content for further exploration. We also model topic-specific reading proficiency, which allows us to better explain why a physician might choose to read sophisticated medical articles yet simple descriptions of SLR cameras. We explore different ways to build user profiles, and use collaborative filtering techniques to overcome data sparsity. We conducted experiments on large-scale datasets from a major Web search engine and a community question answering forum. Our findings confirm that explicitly modeling text comprehensibility can significantly improve content ranking (search results or answers, respectively).
Chenhao Tan, Evgeniy Gabrilovich, Bo Pang 0001
WSDM1
2012 On optimization of expertise matching with various constraints
Jie Tang 0001, Chenhao Tan
Neurocomputing4
2011 Joint Bilingual Sentiment Classification with Unlabeled Parallel Corpora
Bin Lu 0001, Chenhao Tan, Claire Cardie, Benjamin Ka-Yin T'sou
ACL2
2011 Does Bad News Go Away Faster?
Shaomei Wu, Chenhao Tan, Jon M. Kleinberg, Michael W. Macy
ICWSM2
2011 User-level sentiment analysis incorporating social networks
abstract
We show that information about social relationships can be used to improve user-level sentiment analysis. The main motivation behind our approach is that users that are somehow "connected" may be more likely to hold similar opinions; therefore, relationship information can complement what we can extract about a user's viewpoints from their utterances. Employing Twitter as a source for our experimental data, and working within a semi-supervised framework, we propose models that are induced either from the Twitter follower/followee network or from the network in Twitter formed by users referring to each other using "@" mentions. Our transductive learning results reveal that incorporating social-network information can indeed lead to statistically significant sentiment classification improvements over the performance of an approach based on Support Vector Machines having access only to textual features.
Chenhao Tan, Lillian Lee, Jie Tang 0001, Long Jiang, Ming Zhou 0001, Ping Li 0001
KDD1
2010 Social action tracking via noise tolerant time-varying factor graphs
abstract
It is well known that users' behaviors (actions) in a social network are influenced by various factors such as personal interests, social influence, and global trends. However, few publications systematically study how social actions evolve in a dynamic social network and to what extent different factors affect the user actions.
Chenhao Tan, Jie Tang 0001, Jimeng Sun 0001, Quan Lin, Fengjiao Wang
KDD1
2010 Expertise Matching via Constraint-Based Optimization
abstract
Expertise matching, aiming to find the alignment between experts and queries, is a common problem in many real applications such as conference paper-reviewer assignment, product-reviewer alignment, and product-endorser matching. Most of existing methods for this problem usually find “relevant” experts for each query independently by using, e.g., an information retrieval method. However, in real-world systems, various domain-specific constraints must be considered. For example, to review a paper, it is desirable that there is at least one senior reviewer to guide the reviewing process. An important question is: “Can we design a framework to efficiently find the optimal solution for expertise matching under various constraints?” This paper explores such an approach by formulating the expertise matching problem in a constraint based optimization framework. Interestingly, the problem can be linked to a convex cost flow problem, which guarantees an optimal solution under given constraints. We also present an online matching algorithm to support incorporating user feedbacks in real time. The proposed approach has been evaluated on two different genres of expertise matching problems. Experimental results validate the effectiveness of the proposed approach.
Jie Tang 0001, Chenhao Tan
Web Intelligence3