EDBT 2026 Demo / reviewers in the wild / expert
Jilin Chen
dblp:50/6953
· DBLP profile ↗
42ranked-venue papers
9as first author
7since 2021 · last 2024
0000-0002-3359-0938ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 25 · 8 first-author · 2 since 2021Databases, data management, data science and information retrieval · 16 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 14 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Trustworthy machine learning · 41% Language models and text generation · 32% Learning paradigms · 11% | |
| Human-computer interaction and pervasive computing
13 papers |
Collaborative and social computing · 65% Human-AI interaction · 23% Usability and user experience research · 7% | |
| Databases, data mining, and information retrieval
9 papers |
Recommender systems · 67% Web and social media mining · 22% Information retrieval · 9% |
Topics — the 30 heaviest of 53, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
2.2 | 5 | 2023 | Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-Voting · EMNLP 2023 Practical Compositional Fairness: Understanding Fairness in Multi-Component Recommender Systems · WSDM 2021 Understanding and Improving Fairness-Accuracy Trade-offs in Multi-Task Learning · KDD 2021 |
Machine learning › Learning paradigms
multi-task learning |
1.2 | 3 | 2021 | Understanding and Improving Fairness-Accuracy Trade-offs in Multi-Task Learning · KDD 2021 SNR: Sub-Network Routing for Flexible Parameter Sharing in Multi-Task Learning · AAAI 2019 Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-Experts · KDD 2018 |
Recommender systems
fairness-aware recommendation |
0.9 | 2 | 2021 | Practical Compositional Fairness: Understanding Fairness in Multi-Component Recommender Systems · WSDM 2021 Fairness in Recommendation Ranking through Pairwise Comparisons · KDD 2019 |
Collaborative and social computing
online communities |
0.8 | 5 | 2015 | They Said What?: Exploring the Relationship Between Language Use and Member Satisfaction in Communities · CSCW 2015 Selecting an effective niche: an ecological view of the success of online communities · CHI 2014 Goals and perceived success of online enterprise communities: what is important to leaders & members? · CHI 2014 |
Natural language and speech › Language models and text generation
alignment |
0.8 | 1 | 2024 | Controlled Decoding from Language Models · ICML 2024 |
Machine learning › Trustworthy machine learning
calibration |
0.8 | 1 | 2024 | Batch Calibration: Rethinking Calibration for In-Context Learning and Prompt Engineering · ICLR 2024 |
Machine learning › Trustworthy machine learning › fairness and bias
contextual bias |
0.8 | 1 | 2024 | Batch Calibration: Rethinking Calibration for In-Context Learning and Prompt Engineering · ICLR 2024 |
Natural language and speech › Language models and text generation › decoding
controlled decoding |
0.8 | 1 | 2024 | Controlled Decoding from Language Models · ICML 2024 |
Natural language and speech › Language models and text generation
in-context learning |
0.8 | 1 | 2024 | Batch Calibration: Rethinking Calibration for In-Context Learning and Prompt Engineering · ICLR 2024 |
Machine learning › Reinforcement learning › policy optimization
KL-regularized RL |
0.8 | 1 | 2024 | Controlled Decoding from Language Models · ICML 2024 |
Natural language and speech › Language models and text generation
prompting |
0.8 | 1 | 2024 | Batch Calibration: Rethinking Calibration for In-Context Learning and Prompt Engineering · ICLR 2024 |
Machine learning › Trustworthy machine learning › fairness
demographic representation |
0.7 | 1 | 2023 | Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-Voting · EMNLP 2023 |
Human-AI interaction › voice assistants
voice assistant interaction |
0.7 | 1 | 2023 | A Mixed-Methods Approach to Understanding User Trust after Voice Assistant Failures · CHI 2023 |
Machine learning › Trustworthy machine learning › fairness › fairness trade-off
fairness-accuracy trade-off |
0.5 | 1 | 2021 | Understanding and Improving Fairness-Accuracy Trade-offs in Multi-Task Learning · KDD 2021 |
Natural language and speech › Language models and text generation › text generation › synthetic text generation
adversarial text generation |
0.4 | 1 | 2020 | CAT-Gen: Improving Robustness in NLP Models via Controlled Adversarial Text Generation · EMNLP (1) 2020 |
Machine learning › Trustworthy machine learning › fairness › algorithmic fairness
fairness without sensitive attributes |
0.4 | 1 | 2020 | Fairness without Demographics through Adversarially Reweighted Learning · NeurIPS 2020 |
Machine learning › Trustworthy machine learning
robustness |
0.4 | 1 | 2020 | CAT-Gen: Improving Robustness in NLP Models via Controlled Adversarial Text Generation · EMNLP (1) 2020 |
Machine learning › Efficient and distributed learning
model compression |
0.4 | 1 | 2019 | SNR: Sub-Network Routing for Flexible Parameter Sharing in Multi-Task Learning · AAAI 2019 |
Machine learning › Efficient and distributed learning
parameter sharing |
0.4 | 1 | 2019 | SNR: Sub-Network Routing for Flexible Parameter Sharing in Multi-Task Learning · AAAI 2019 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.3 | 1 | 2018 | Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-Experts · KDD 2018 |
Machine learning › Learning paradigms › multi-task learning
task relationship modeling |
0.3 | 1 | 2018 | Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-Experts · KDD 2018 |
Computer vision › Image recognition and object detection
image classification |
0.2 | 1 | 2024 | Batch Calibration: Rethinking Calibration for In-Context Learning and Prompt Engineering · ICLR 2024 |
Natural language and speech › Language models and text generation › large language model inference
inference-time decoding |
0.2 | 1 | 2024 | Controlled Decoding from Language Models · ICML 2024 |
Recommender systems
content recommendation |
0.2 | 2 | 2018 | Short and tweet: experiments on recommending content from information streams · CHI 2010 Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-Experts · KDD 2018 |
Computational social science and digital humanities
social media analysis |
0.2 | 1 | 2014 | Understanding individuals' personal values from social media word use · CSCW 2014 |
Web and social media mining › social media analysis
reddit |
0.2 | 1 | 2014 | Understanding individuals' personal values from social media word use · CSCW 2014 |
Collaborative and social computing › social media
enterprise social media |
0.2 | 1 | 2014 | Social media participation and performance at work: a longitudinal study · CHI 2014 |
Empirical software engineering
developer studies |
0.2 | 1 | 2014 | Social media participation and performance at work: a longitudinal study · CHI 2014 |
Visualization and visual analytics
visual comparison |
0.2 | 1 | 2013 | CommunityCompare: visually comparing communities for online community leaders in the enterprise · CHI 2013 |
Natural language and speech › Language models and text generation › trustworthy language model
natural language processing robustness |
0.1 | 1 | 2020 | CAT-Gen: Improving Robustness in NLP Models via Controlled Adversarial Text Generation · EMNLP (1) 2020 |
Methods — techniques the papers use, named apart from their topics
survey · 1.4interviews · 1.0fairness analysis · 1.0value function · 0.8reinforcement learning · 0.8prompt engineering · 0.8prefix scorer · 0.8batch calibration · 0.8self-voting · 0.7mixed methods · 0.7crowdsourced dataset · 0.7collective-critiques · 0.7longitudinal study · 0.6quantitative analysis · 0.5pareto optimization · 0.5adversarial training · 0.4regularization · 0.4pairwise comparison · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Batch Calibration: Rethinking Calibration for In-Context Learning and Prompt EngineeringabstractPrompting and in-context learning (ICL) have become efficient learning paradigms for large language models (LLMs). However, LLMs suffer from prompt brittleness and various bias factors in the prompt, including but not limited to the formatting, the choice verbalizers, and the ICL examples. To address this problem that results in unexpected performance degradation, calibration methods have been developed to mitigate the effects of these biases while recovering LLM performance. In this work, we first conduct a systematic analysis of the existing calibration methods, where we both provide a unified view and reveal the failure cases. Inspired by these analyses, we propose Batch Calibration (BC), a simple yet intuitive method that controls the contextual bias from the batched input, unifies various prior approaches and effectively addresses the aforementioned issues. BC is zero-shot, inference-only, and incurs negligible additional costs. In the few-shot setup, we further extend BC to allow it to learn the contextual bias from labeled data. We validate the effectiveness of BC with PaLM 2-(S, M, L) and CLIP models and demonstrate state-of-the-art performance over previous calibration baselines across more than 10 natural language understanding and image classification tasks. Han Zhou 0010, Xingchen Wan, Lev Proleev, Diana Mincu, Jilin Chen, Katherine A. Heller, Subhrajit Roy |
ICLR | 5 |
| 2024 | Controlled Decoding from Language ModelsabstractKL-regularized reinforcement learning (RL) is a popular alignment framework to control the language model responses towards high reward outcomes. We pose a tokenwise RL objective and propose a modular solver for it, called *controlled decoding (CD)*. CD exerts control through a separate *prefix scorer* module, which is trained to learn a value function for the reward. The prefix scorer is used at inference time to control the generation from a frozen base model, provably sampling from a solution to the RL objective. We empirically demonstrate that CD is effective as a control mechanism on popular benchmarks. We also show that prefix scorers for multiple rewards may be combined at inference time, effectively solving a multi-objective RL problem with no additional training. We show that the benefits of applying CD transfer to an unseen base model with no further tuning as well. Finally, we show that CD can be applied in a blockwise decoding fashion at inference-time, essentially bridging the gap between the popular best-of-$K$ strategy and tokenwise control through reinforcement learning. This makes CD a promising approach for alignment of language models. Sidharth Mudgal, Jong Lee, Harish Ganapathy, YaGuang Li, Yanping Huang, Heng-Tze Cheng, Trevor Strohman, Jilin Chen, Alex Beutel, Ahmad Beirami |
ICML | 11 |
| 2023 | A Mixed-Methods Approach to Understanding User Trust after Voice Assistant FailuresabstractDespite huge gains in performance in natural language understanding via large language models in recent years, voice assistants still often fail to meet user expectations. In this study, we conducted a mixed-methods analysis of how voice assistant failures affect users’ trust in their voice assistants. To illustrate how users have experienced these failures, we contribute a crowdsourced dataset of 199 voice assistant failures, categorized across 12 failure sources. Relying on interview and survey data, we find that certain failures, such as those due to overcapturing users’ input, derail user trust more than others. We additionally examine how failures impact users’ willingness to rely on voice assistants for future tasks. Users often stop using their voice assistants for specific tasks that result in failures for a short period of time before resuming similar usage. We demonstrate the importance of low stakes tasks, such as playing music, towards building trust after failures. Amanda Baughan, Xuezhi Wang 0002, Ariel Liu, Allison Mercurio, Jilin Chen, Xiao Ma 0010 |
CHI | 5 |
| 2023 | Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-VotingabstractPreethi Lahoti, Nicholas Blumm, Xiao Ma, Raghavendra Kotikalapudi, Sahitya Potluri, Qijun Tan, Hansa Srinivasan, Ben Packer, Ahmad Beirami, Alex Beutel, Jilin Chen. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Preethi Lahoti, Nicholas Blumm, Xiao Ma 0010, Raghavendra Kotikalapudi, Sahitya Potluri, Qijun Tan, Hansa Srinivasan, Ben Packer, Ahmad Beirami, Alex Beutel, Jilin Chen |
EMNLP | 11 |
| 2021 | Measuring Model Fairness under Noisy Covariates: A Theoretical PerspectiveabstractIn this work we study the problem of measuring the fairness of a machine learning model under noisy information. Focusing on group fairness metrics, we investigate the particular but common situation when the evaluation requires controlling for the confounding effect of covariate variables. In a practical setting, we might not be able to jointly observe the covariate and group information, and a standard workaround is to then use proxies for one or more of these variables. Prior works have demonstrated the challenges with using a proxy for sensitive attributes, and strong independence assumptions are needed to provide guarantees on the accuracy of the noisy estimates. In contrast, in this work we study using a proxy for the covariate variable and present a theoretical analysis that aims to characterize weaker conditions under which accurate fairness evaluation is possible. Furthermore, our theory identifies potential sources of errors and decouples them into two interpretable parts y and E. The first part y depends solely on the performance of the proxy such as precision and recall, whereas the second part E captures correlations between all the variables of interest. We show that in many scenarios the error in the estimates is dominated by y via a linear dependence, whereas the dependence on the correlations E only constitutes a lower order term. As a result we expand the understanding of scenarios where measuring model fairness via proxies can be an effective approach. Finally, we compare, via simulations, the theoretical upper-bounds to the distribution of simulated estimation errors and show that assuming some structure on the data, even weak, is key to significantly improve both theoretical guarantees and empirical results. Flavien Prost, Pranjal Awasthi, Nick Blumm, Aditee Kumthekar, Trevor Potter, Xuezhi Wang 0002, Ed H. Chi, Jilin Chen, Alex Beutel |
AIES | 9 |
| 2021 | Understanding and Improving Fairness-Accuracy Trade-offs in Multi-Task LearningabstractAs multi-task models gain popularity in a wider range of machine learning applications, it is becoming increasingly important for practitioners to understand the fairness implications associated with those models. Most existing fairness literature focuses on learning a single task more fairly, while how ML fairness interacts with multiple tasks in the joint learning setting is largely under-explored. In this paper, we are concerned with how group fairness (e.g., equal opportunity, equalized odds) as an ML fairness concept plays out in the multi-task scenario. In multi-task learning, several tasks are learned jointly to exploit task correlations for a more efficient inductive transfer. This presents a multi-dimensional Pareto frontier on (1) the trade-off between group fairness and accuracy with respect to each task, as well as (2) the trade-offs across multiple tasks. We aim to provide a deeper understanding on how group fairness interacts with accuracy in multi-task learning, and we show that traditional approaches that mainly focus on optimizing the Pareto frontier of multi-task accuracy might not perform well on fairness goals. We propose a new set of metrics to better capture the multi-dimensional Pareto frontier of fairness-accuracy trade-offs uniquely presented in a multi-task learning setting. We further propose a Multi-Task-Aware Fairness (MTA-F) approach to improve fairness in multi-task learning. Experiments on several real-world datasets demonstrate the effectiveness of our proposed approach. Xuezhi Wang 0002, Alex Beutel, Flavien Prost, Jilin Chen, Ed H. Chi |
KDD | 5 |
| 2021 | Practical Compositional Fairness: Understanding Fairness in Multi-Component Recommender SystemsabstractHow can we build recommender systems to take into account fairness? Real-world recommender systems are often composed of multiple models, built by multiple teams. However, most research on fairness focuses on improving fairness in a single model. Further, recent research on classification fairness has shown that combining multiple "fair" classifiers can still result in an "unfair" classification system. This presents a significant challenge: how do we understand and improve fairness in recommender systems composed of multiple components? Xuezhi Wang 0002, Nithum Thain, Anu Sinha, Flavien Prost, Ed H. Chi, Jilin Chen, Alex Beutel |
WSDM | 6 |
| 2020 | CAT-Gen: Improving Robustness in NLP Models via Controlled Adversarial Text GenerationabstractNLP models are shown to suffer from robustness issues, i.e., a model's prediction can be easily changed under small perturbations to the input.In this work, we present a Controlled Adversarial Text Generation (CAT-Gen) model that, given an input text, generates adversarial texts through controllable attributes that are known to be irrelevant to task labels.For example, in order to attack a model for sentiment classification over product reviews, we can use the product categories as the controllable attribute which should not change the sentiment of the reviews.Experiments on real-world NLP datasets demonstrate that our method can generate more diverse and fluent adversarial texts, compared to many existing adversarial text generation approaches.We further use our generated adversarial examples to improve models through adversarial training, and we demonstrate that our generated attacks are more robust against model retraining and different model architectures. Xuezhi Wang 0002, Yao Qin 0001, Ben Packer, Jilin Chen, Alex Beutel, Ed H. Chi |
EMNLP (1) | 6 |
| 2020 | Fairness without Demographics through Adversarially Reweighted LearningabstractMuch of the previous machine learning (ML) fairness literature assumes that protected features such as race and sex are present in the dataset, and relies upon them to mitigate fairness concerns. However, in practice factors like privacy and regulation often preclude the collection of protected features, or their use for training or inference, severely limiting the applicability of traditional fairness research. Therefore, we ask: How can we train a ML model to improve fairness when we do not even know the protected group memberships? In this work we address this problem by proposing Adversarially Reweighted Learning (ARL). In particular, we hypothesize that non-protected features and task labels are valuable for identifying fairness issues, and can be used to co-train an adversarial reweighting approach for improving fairness. Our results show that ARL improves Rawlsian Max-Min fairness, with notable AUC improvements for worst-case protected groups in multiple datasets, outperforming state-of-the-art alternatives. Preethi Lahoti, Alex Beutel, Jilin Chen, Kang Lee, Flavien Prost, Nithum Thain, Xuezhi Wang 0002, Ed H. Chi |
NeurIPS | 3 |
| 2019 | SNR: Sub-Network Routing for Flexible Parameter Sharing in Multi-Task LearningabstractMachine learning applications, such as object detection and content recommendation, often require training a single model to predict multiple targets at the same time. Multi-task learning through neural networks became popular recently, because it not only helps improve the accuracy of many prediction tasks when they are related, but also saves computation cost by sharing model architectures and low-level representations. The latter is critical for real-time large-scale machine learning systems. However, classic multi-task neural networks may degenerate significantly in accuracy when tasks are less related. Previous works (Misra et al. 2016; Yang and Hospedales 2016; Ma et al. 2018) showed that having more flexible architectures in multi-task models, either manually-tuned or softparameter-sharing structures like gating networks, helps improve the prediction accuracy. However, manual tuning is not scalable, and the previous soft-parameter sharing models are either not flexible enough or computationally expensive. In this work, we propose a novel framework called SubNetwork Routing (SNR) to achieve more flexible parameter sharing while maintaining the computational advantage of the classic multi-task neural-network model. SNR modularizes the shared low-level hidden layers into multiple layers of subnetworks, and controls the connection of sub-networks with learnable latent variables to achieve flexible parameter sharing. We demonstrate the effectiveness of our approach on a large-scale dataset YouTube8M. We show that the proposed method improves the accuracy of multi-task models while maintaining their computation efficiency. Jiaqi W. Ma, Zhe Zhao 0001, Jilin Chen, Lichan Hong, Ed H. Chi |
AAAI | 3 |
| 2019 | Putting Fairness Principles into Practice: Challenges, Metrics, and ImprovementsabstractAs more researchers have become aware of and passionate about algorithmic fairness, there has been an explosion in papers laying out new metrics, suggesting algorithms to address issues, and calling attention to issues in existing applications of machine learning. This research has greatly expanded our understanding of the concerns and challenges in deploying machine learning, but there has been much less work in seeing how the rubber meets the road. In this paper we provide a case-study on the application of fairness in machine learning research to a production classification system, and offer new insights in how to measure and address algorithmic fairness issues. We discuss open questions in implementing equality of opportunity and describe our fairness metric, conditional equality, that takes into account distributional differences. Further, we provide a new approach to improve on the fairness metric during model training and demonstrate its efficacy in improving performance for a real-world product. Alex Beutel, Jilin Chen, Tulsee Doshi, Hai Qian, Allison Woodruff, Christine Luu, Pierre Kreitmann, Jonathan Bischof, Ed H. Chi |
AIES | 2 |
| 2019 | Fairness in Recommendation Ranking through Pairwise ComparisonsabstractRecommender systems are one of the most pervasive applications of machine learning in industry, with many services using them to match users to products or information. As such it is important to ask: what are the possible fairness risks, how can we quantify them, and how should we address them? In this paper we offer a set of novel metrics for evaluating algorithmic fairness concerns in recommender systems. In particular we show how measuring fairness based on pairwise comparisons from randomized experiments provides a tractable means to reason about fairness in rankings from recommender systems. Building on this metric, we offer a new regularizer to encourage improving this metric during model training and thus improve fairness in the resulting rankings. We apply this pairwise regularization to a large-scale, production recommender system and show that we are able to significantly improve the system's pairwise fairness. Alex Beutel, Jilin Chen, Tulsee Doshi, Hai Qian, Lukasz Heldt, Zhe Zhao 0001, Lichan Hong, Ed H. Chi, Cristos Goodrow |
KDD | 2 |
| 2019 | Recommending what video to watch next: a multitask ranking systemabstractIn this paper, we introduce a large scale multi-objective ranking system for recommending what video to watch next on an industrial video sharing platform. The system faces many real-world challenges, including the presence of multiple competing ranking objectives, as well as implicit selection biases in user feedback. To tackle these challenges, we explored a variety of soft-parameter sharing techniques such as Multi-gate Mixture-of-Experts so as to efficiently optimize for multiple ranking objectives. Additionally, we mitigated the selection biases by adopting a Wide & Deep framework. We demonstrated that our proposed techniques can lead to substantial improvements on recommendation quality on one of the world's largest video sharing platforms. Zhe Zhao 0001, Lichan Hong, Jilin Chen, Aniruddh Nath, Shawn Andrews, Aditee Kumthekar, Maheswaran Sathiamoorthy, Xinyang Yi, Ed H. Chi |
RecSys | 4 |
| 2018 | Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-ExpertsabstractNeural-based multi-task learning has been successfully used in many real-world large-scale applications such as recommendation systems. For example, in movie recommendations, beyond providing users movies which they tend to purchase and watch, the system might also optimize for users liking the movies afterwards. With multi-task learning, we aim to build a single model that learns these multiple goals and tasks simultaneously. However, the prediction quality of commonly used multi-task models is often sensitive to the relationships between tasks. It is therefore important to study the modeling tradeoffs between task-specific objectives and inter-task relationships. In this work, we propose a novel multi-task learning approach, Multi-gate Mixture-of-Experts (MMoE), which explicitly learns to model task relationships from data. We adapt the Mixture-of-Experts (MoE) structure to multi-task learning by sharing the expert submodels across all tasks, while also having a gating network trained to optimize each task. To validate our approach on data with different levels of task relatedness, we first apply it to a synthetic dataset where we control the task relatedness. We show that the proposed approach performs better than baseline methods when the tasks are less related. We also show that the MMoE structure results in an additional trainability benefit, depending on different levels of randomness in the training data and model initialization. Furthermore, we demonstrate the performance improvements by MMoE on real tasks including a binary classification benchmark, and a large-scale content recommendation system at Google. Jiaqi W. Ma, Zhe Zhao 0001, Xinyang Yi, Jilin Chen, Lichan Hong, Ed H. Chi |
KDD | 4 |
| 2018 | Categorical-attributes-based item classification for recommender systemsabstractMany techniques to utilize side information of users and/or items as inputs to recommenders to improve recommendation, especially on cold-start items/users, have been developed over the years. In this work, we test the approach of utilizing item side information, specifically categorical attributes, in the output of recommendation models either through multi-task learning or hierarchical classification. We first demonstrate the efficacy of these approaches for both matrix factorization and neural networks with a medium-size real-word data set. We then show that they improve a neural-network based production model in an industrial-scale recommender system. We demonstrate the robustness of the hierarchical classification approach by introducing noise in building the hierarchy. Lastly, we investigate the generalizability of hierarchical classification on a simulated dataset by building two user models in which we can fully control the generative process of user-item interactions. Jilin Chen, Minmin Chen, Sagar Jain, Alex Beutel, Francois Belletti, Ed H. Chi |
RecSys | 2 |
| 2018 | Evaluation and Refinement of Clustered Search Results with the CrowdabstractWhen searching on the web or in an app, results are often returned as lists of hundreds to thousands of items, making it difficult for users to understand or navigate the space of results. Research has demonstrated that using clustering to partition search results into coherent, topical clusters can aid in both exploration and discovery. Yet clusters generated by an algorithm for this purpose are often of poor quality and do not satisfy users. To achieve acceptable clustered search results, experts must manually evaluate and refine the clustered results for each search query, a process that does not scale to large numbers of search queries. In this article, we investigate using crowd-based human evaluation to inspect, evaluate, and improve clusters to create high-quality clustered search results at scale. We introduce a workflow that begins by using a collection of well-known clustering algorithms to produce a set of clustered search results for a given query. Then, we use crowd workers to holistically assess the quality of each clustered search result to find the best one. Finally, the workflow has the crowd spot and fix problems in the best result to produce a final output. We evaluate this workflow on 120 top search queries from the Google Play Store, some of whom have clustered search results as a result of evaluations and refinements by experts. Our evaluations demonstrate that the workflow is effective at reproducing the evaluation of expert judges and also improves clusters in a way that agrees with experts and crowds alike. Amy X. Zhang, Jilin Chen, Wei Chai, Jinjun Xu, Lichan Hong, Ed H. Chi |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2017 | 'Just the Facts: ' Exploring the Relationship Between Emotional Language and Member Satisfaction in Enterprise Online Communities
Ryan Compton 0002, Jilin Chen, Eben M. Haber, Hernan Badenes, Steve Whittaker 0001 |
ICWSM | 2 |
| 2015 | InkWell: A Creative Writer's Creative AssistantabstractInkWell is a writer's assistant---a natural language revision program designed to assist creative writers by producing stylistic variations on texts based on craft-based facets of creative writing and by mimicking aspects of specified writers and their personality traits. It is built on top of an optimization process that produces variations on a supplied text, evaluates those variations quantitatively, and selects variations that best satisfy the goals of writing craft and writer mimicry. We describe the design and capabilities of InkWell, and present an early evaluation of its effectiveness and uses with two established literary writers along with an experiment using InkWell to write haiku on its own. Richard P. Gabriel, Jilin Chen, Jeffrey Nichols 0001 |
Creativity & Cognition | 2 |
| 2015 | They Said What?: Exploring the Relationship Between Language Use and Member Satisfaction in CommunitiesabstractIn online communities, satisfied members are essential to community success, since they are more likely to contribute and consume content, engage with other members, and feel committed to the community. However, it is difficult for community leaders to know, on an on-going basis, whether members are satisfied. In this paper, we explore the relationship between member satisfaction and language use within content posted in workplace online communities. We hope to find patterns of language use that are associated with satisfied members. We employ linguistic analysis based on LIWC, and a survey to directly measure member satisfaction in 142 workplace communities. We contribute a better understanding of how members interact in effective workplace communities, and show that linguistic analysis could be a useful part of future methods to automatically assess community member satisfaction. Tara Matthews, Jalal Mahmud, Jilin Chen, Michael J. Muller, Eben M. Haber, Hernan Badenes |
CSCW | 3 |
| 2015 | Making Use of Derived Personality: The Case of Social Media Ad Targeting
Jilin Chen, Eben M. Haber, Ruogu Kang, Gary Hsieh, Jalal Mahmud |
ICWSM | 1 |
| 2015 | Who Will Retweet This? Detecting Strangers from Twitter to Retweet InformationabstractThere has been much effort on studying how social media sites, such as Twitter, help propagate information in different situations, including spreading alerts and SOS messages in an emergency. However, existing work has not addressed how to actively identify and engage the right strangers at the right time on social media to help effectively propagate intended information within a desired time frame. To address this problem, we have developed three models: (1) a feature-based model that leverages people's exhibited social behavior, including the content of their tweets and social interactions, to characterize their willingness and readiness to propagate information on Twitter via the act of retweeting; (2) a wait-time model based on a user's previous retweeting wait times to predict his or her next retweeting time when asked; and (3) a subset selection model that automatically selects a subset of people from a set of available people using probabilities predicted by the feature-based model and maximizes retweeting rate. Based on these three models, we build a recommender system that predicts the likelihood of a stranger to retweet information when asked, within a specific time window, and recommends the top-N qualified strangers to engage with. Our experiments, including live studies in the real world, demonstrate the effectiveness of our work. Kyumin Lee, Jalal Mahmud, Jilin Chen, Michelle X. Zhou, Jeffrey Nichols 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2014 | You read what you value: understanding personal values and reading interestsabstractThis paper presents an experiment on the relationship between personal values and reading interests of online articles. Results suggest that individuals' values can predict their topical interests. For example, holding stronger universalism values predict interests towards environmental articles, whereas holding stronger achievement values predict interest towards work-related articles. Findings demonstrate the possibility of targeting based on individuals' personal values, but also highlight certain challenges and limitations when applying this approach for online content. Gary Hsieh, Jilin Chen, Jalal Mahmud, Jeffrey Nichols 0001 |
CHI | 2 |
| 2014 | Goals and perceived success of online enterprise communities: what is important to leaders & members?abstractOnline communities are successful only if they achieve their goals, but there has been little direct study of goals. We analyze novel data characterizing the goals of enterprise online communities, assessing the importance of goals for leaders, how goals influence member perceptions of community value, and how goals relate to success measures proposed in the literature. We find that most communities have multiple goals and common goals are learning, reuse of resources, collaboration, networking, influencing change, and innovation. Leaders and members agree that all of these goals are important, but their perceptions of success on goals do not align with each other, or with commonly used behavioral success measures. We conclude that simple behavioral measures and leader perceptions are not good success metrics, and propose alternatives based on specific goals members and leaders judge most important. Tara Matthews, Jilin Chen, Steve Whittaker 0001, Aditya Pal, Haiyi Zhu, Hernan Badenes, Barton A. Smith |
CHI | 2 |
| 2014 | Social media participation and performance at work: a longitudinal studyabstractThe use of social media at work is gaining traction, and there is evidence to suggest that various benefits accrue from its use. Yet the relationship between using social media at work and employee performance is not clear. Through a study of 75,747 employees of a large global company over the course of 3 years, we find that some social media usage (number of forum posts, forum post length, and status update length) was positively associated with performance ratings. This study is one of the first to show the relationship among different forms of social media use and employee performance ratings. N. Sadat Shami, Jeffrey Nichols 0001, Jilin Chen |
CHI | 3 |
| 2014 | Selecting an effective niche: an ecological view of the success of online communitiesabstractOnline communities serve various important functions, but many fail to thrive. Research on community success has traditionally focused on internal factors. In contrast, we take an ecological view to understand how the success of a community is influenced by other communities. We measured a community's relationship with other communities - its "niche" - through four dimensions: topic overlap, shared members, content linking, and shared offline organizational affiliation. We used a mixed-method approach, combining the quantitative analysis of 9495 online enterprise communities and interviews with community members. Our results show that too little or too much overlap in topic with other communities causes a community's activity to suffer. We also show that this main result is moderated in predictable ways by whether the community shares members with, links to content in, or shares an organizational affiliation with other communities. These findings provide new insight on community success, guiding online community designers on how to effectively position their community in relation to others. Haiyi Zhu, Jilin Chen, Tara Matthews, Aditya Pal, Hernan Badenes, Robert E. Kraut |
CHI | 2 |
| 2014 | Understanding individuals' personal values from social media word useabstractThe theory of values posits that each person has a set of values, or desirable and trans-situational goals, that motivate their actions. The Basic Human Values, a motivational construct that captures people's values, have been shown to influence a wide range of human behaviors. In this work, we analyze people's values and their word use on Reddit, an online social news sharing community. Through conducting surveys and analyzing text contributions of 799 Reddit users, we identify and interpret categories of words that are indicative of user's value orientations. Using the same data, we further report a preliminary exploration on word-based prediction of Basic Human Values. Jilin Chen, Gary Hsieh, Jalal Mahmud, Jeffrey Nichols 0001 |
CSCW | 1 |
| 2014 | Modeling User Attitude toward Controversial Topics in Online Social Media
Huiji Gao, Jalal Mahmud, Jilin Chen, Jeffrey Nichols 0001, Michelle X. Zhou |
ICWSM | 3 |
| 2014 | Who will retweet this?: Automatically Identifying and Engaging Strangers on Twitter to Spread InformationabstractThere has been much effort on studying how social media sites, such as Twitter, help propagate information in different situations, including spreading alerts and SOS messages in an emergency. However, existing work has not addressed how to actively identify and engage the right strangers at the right time on social media to help effectively propagate intended information within a desired time frame. To ad-dress this problem, we have developed two models: (i) a feature-based model that leverages peoplesfi exhibited social behavior, including the content of their tweets and social interactions, to characterize their willingness and readiness to propagate information on Twitter via the act of retweeting; and (ii) a wait-time model based on a user's previous retweeting wait times to predict her next retweeting time when asked. Based on these two models, we build a recommender system that predicts the likelihood of a stranger to retweet information when asked, within a specific time window, and recommends the top-N qualified strangers to engage with. Our experiments, including live studies in the real world, demonstrate the effectiveness of our work. Kyumin Lee, Jalal Mahmud, Jilin Chen, Michelle X. Zhou, Jeffrey Nichols 0001 |
IUI | 3 |
| 2014 | System U: automatically deriving personality traits from social media for people recommendationabstractThis paper presents a system, System U, which automatically derives people's personality traits from social media and recommends people for different tasks. The system leverages linguistic signals appearing in a person's social media activities to compute the personality portraits including Big Five personality, fundamental needs and basic human values. This system and technology can be used in a wide variety of personalized applications, such as recommending people to answer questions. Hernan Badenes, Mateo N. Bengualid, Jilin Chen, Liang Gou, Eben M. Haber, Jalal Mahmud, Jeffrey Nichols 0001, Aditya Pal, Jerald Schoudt, Barton A. Smith, Ying Xuan, Huahai Yang, Michelle X. Zhou |
RecSys | 3 |
| 2013 | CommunityCompare: visually comparing communities for online community leaders in the enterpriseabstractOnline communities are important in enterprises, helping workers to build skills and collaborate. Despite their unique and critical role fostering successful communities, community leaders have little direct support in existing technologies. We introduce CommunityCompare, an interactive visual analytic system to enable leaders to make sense of their community's activity with comparisons. Composed of a parallel coordinates plot, various control widgets, and a preview of example posts from communities, the system supports comparisons with hundreds of related communities on multiple metrics and the ability to learn by example. We motivate and inform the system design with formative interviews of community leaders. From additional interviews, a field deployment, and surveys of leaders, we show how the system enabled leaders to assess community performance in the context of other comparable communities, learn about community dynamics through data exploration, and identify examples of top performing communities from which to learn. We conclude by discussing how our system and design lessons generalize. Anbang Xu, Jilin Chen, Tara Matthews, Michael J. Muller, Hernan Badenes |
CHI | 2 |
| 2013 | CrowdE: Filtering Tweets for Direct Customer Engagements
Jilin Chen, Allen Cypher, Clemens Drews, Jeffrey Nichols 0001 |
ICWSM | 1 |
| 2013 | When Will You Answer This? Estimating Response Time in Twitter
Jalal Mahmud, Jilin Chen, Jeffrey Nichols 0001 |
ICWSM | 2 |
| 2012 | Searching for the goldilocks zone: trade-offs in managing online volunteer groupsabstractDedicated and productive members who actively contribute to community efforts are crucial to the success of online volunteer groups such as Wikipedia. What predicts member productivity? Do productive members stay longer? How does involvement in multiple projects affect member contribution to the community? In this paper, we analyze data from 648 WikiProjects to address these questions. Our results reveal two critical trade-offs in managing online volunteer groups. First, factors that increase member productivity, measured by the number of edits on Wikipedia articles, also increase likelihood of withdrawal from contributing, perhaps due to feelings of mission accomplished or burnout. Second, individual membership in multiple projects has mixed effects. It decreases the amount of work editors contribute to both the individual projects and Wikipedia as a whole. It increases withdrawal for each individual project yet reduces withdrawal from Wikipedia. We discuss how our findings expand existing theories to fit the online context and inform the design of new tools to improve online volunteer work. Loxley Sijia Wang, Jilin Chen, Yuqing Ren, John Riedl |
CSCW | 2 |
| 2012 | Why You Are More Engaged: Factors Influencing Twitter Engagement in Occupy Wall Street
Jilin Chen, Peter Pirolli |
ICWSM | 1 |
| 2011 | Speak little and well: recommending conversations in online social streamsabstractConversation is a key element in online social streams such as Twitter and Facebook. However, finding interesting conversations to read is often a challenge, due to information overload and differing user preferences. In this work we explored five algorithms that recommend conversations to Twitter users, utilizing thread length, topic and tie-strength as factors. We compared the algorithms through an online user study and gathered feedback from real Twitter users. In particular, we investigated how users' purposes of using Twitter affect user preferences for different types of conversations and the performance of different algorithms. Compared to a random baseline, all algorithms recommended more interesting conversations. Further, tie-strength based algorithms performed significantly better for people who use Twitter for social purposes than for people who use Twitter for informational purpose only. Jilin Chen, Rowan Nairn, Ed H. Chi |
CHI | 1 |
| 2010 | Short and tweet: experiments on recommending content from information streamsabstractMore and more web users keep up with newest information through information streams such as the popular micro-blogging website Twitter. In this paper we studied content recommendation on Twitter to better direct user attention. In a modular approach, we explored three separate dimensions in designing such a recommender: content sources, topic interest models for users, and social voting. We implemented 12 recommendation engines in the design space we formulated, and deployed them to a recommender service on the web to gather feedback from real Twitter users. The best performing algorithm improved the percentage of interesting content to 72% from a baseline of 33%. We conclude this work by discussing the implications of our recommender design and how our design can generalize to other information streams. Jilin Chen, Rowan Nairn, Les Nelson, Michael S. Bernstein, Ed H. Chi |
CHI | 1 |
| 2010 | The effects of diversity on group productivity and member withdrawal in online volunteer groupsabstractThe "wisdom of crowds" argument emphasizes the importance of diversity in online collaborations, such as open source projects and Wikipedia. However, decades of research on diversity in offline work groups have painted an inconclusive picture. On the one hand, the broader range of insights from a diverse group can lead to improved outcomes. On the other hand, individual differences can lead to conflict and diminished performance. In this paper, we examine the effects of group diversity on the amount of work accomplished and on member withdrawal behaviors in the context of WikiProjects. We find that increased diversity in experience with Wikipedia increases group productivity and decreases member withdrawal -- up to a point. Beyond that point, group productivity remains high, but members are more likely to withdraw. Strikingly, no such diminishing returns were observed for differences in member interest, which increases productivity and decreases member withdrawal in a linear fashion. Our results suggest that the low visibility of individual differences in online groups may allow them to harvest more of the benefits of diversity while bearing less of the cost. We discuss how our findings can inform further research of online collaboration. Jilin Chen, Yuqing Ren, John Riedl |
CHI | 1 |
| 2010 | Eddi: interactive topic-based browsing of social status streamsabstractTwitter streams are on overload: active users receive hundreds of items per day, and existing interfaces force us to march through a chronologically-ordered morass to find tweets of interest. We present an approach to organizing a user's own feed into coherently clustered trending topics for more directed exploration. Our Twitter client, called Eddi, groups tweets in a user's feed into topics mentioned explicitly or implicitly, which users can then browse for items of interest. To implement this topic clustering, we have developed a novel algorithm for discovering topics in short status updates powered by linguistic syntactic transformation and callouts to a search engine. An algorithm evaluation reveals that search engine callouts outperform other approaches when they employ simple syntactic transformation and backoff strategies. Active Twitter users evaluated Eddi and found it to be a more efficient and enjoyable way to browse an overwhelming status update feed than the standard chronological interface. Michael S. Bernstein, Bongwon Suh, Lichan Hong, Jilin Chen, Sanjay Kairam, Ed H. Chi |
UIST | 4 |
| 2009 | Make new friends, but keep the old: recommending people on social networking sitesabstractThis paper studies people recommendations designed to help users find known, offline contacts and discover new friends on social networking sites. We evaluated four recommender algorithms in an enterprise social networking site using a personalized survey of 500 users and a field study of 3,000 users. We found all algorithms effective in expanding users' friend lists. Algorithms based on social network information were able to produce better-received recommendations and find more known contacts for users, while algorithms using similarity of user-created content were stronger in discovering new friends. We also collected qualitative feedback from our survey users and draw several meaningful design implications. Jilin Chen, Werner Geyer, Casey Dugan, Michael J. Muller, Ido Guy |
CHI | 1 |
| 2007 | Creating, destroying, and restoring value in wikipediaabstractWikipedia’s brilliance and curse is that any user can edit any of the encyclopedia entries. We introduce the notion of the impact of an edit, measured by the number of times the edited version is viewed. Using several datasets, including recent logs of all article views, we show that frequent editors dominate what people see when they visit Wikipedia, and that this domination is increasing. Similarly, using the same impact measure, we show that the probability of a typical article view being damaged is small but increasing, and we present empirically grounded classes of damage. Finally, we make policy recommendations for Wikipedia and other wikis in light of these findings. Reid Priedhorsky, Jilin Chen, Shyong K. Lam, Katherine A. Panciera, Loren G. Terveen, John Riedl |
GROUP | 2 |
| 2007 | Techlens: a researcher's desktopabstractRapid and continuous growth of digital libraries, coupled with brisk advancements in technology, has driven users to seek tools and services that are not only customized to their specific needs, but are also helpful in keeping them stay abreast with the latest developments in their field. TechLens is a recommender system that learns about its users through implicit feedback, builds correlations among them, and uses that information to generate recommendations that match the user's profile. It gives users control over which parts of their profile of known citations are used in forming recommendations for new articles. This demonstration is a prototype that showcases some of the tools and services that TechLens offers to the users of digital libraries. Nishikant Kapoor, Jilin Chen, John T. Butler, Gary C. Fouty, James A. Stemper, John Riedl, Joseph A. Konstan |
RecSys | 2 |
| 2006 | Diverse Topic Phrase Extraction through Latent Semantic AnalysisabstractWe propose a novel algorithm for extracting diverse topic phrases in order to provide summary for large corpora. Previous works often ignore the importance of diversity and thus extract phrases crowded on some hot topics while failing to cover other less obvious but important topics. We solve this problem through document re-weighting and phrase diversification by using latent semantic analysis (LSA). Experiments on various datasets show that our new algorithm can improve relevance as well as diversity over different topics for topic phrase extraction problems. Jilin Chen, Jun Yan 0001, Benyu Zhang, Qiang Yang 0001, Zheng Chen 0001 |
ICDM | 1 |