VLDB 2026 Research / reviewers in the wild / expert
Lyle H. Ungar
dblp:u/LyleHUngar
· DBLP profile ↗
31ranked-venue papers in the field
0as first author
5since 2021 · last 2024
0000-0003-2047-1443ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 18Data Mining & Knowledge Discovery · 12Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Measuring Causal Effects of Civil Communication without RandomizationabstractUnderstanding the causal effects of civility is critical when analyzing online social communication, yet measuring causality is difficult. A/B tests and other randomized experiments are the gold standard for establishing causal effects but they are inapplicable in this setting due to 1) the inability to control civility levels in an experiment, and more importantly, 2) ethical constraints on intentionally randomizing civility levels. We develop a novel quasi-experimental approach to quantify the causal effect of civility in online communities on the Roblox social 3D platform without requiring explicit randomization. This method uses residual stochasticity in the "matchmaking" assignment of users to servers as a quasi-randomization mechanism in observational historical data. We find that assigning a user to a server with higher levels of civil communication could increase engagement time by as much as 1.5% in particular experiences. Given the 4.8B person hours spent monthly on the platform, this implies a potential increase of over 8,000 person years of social interaction every month. Furthermore, this effect is mis-estimated by non-causal methods. Quasi-experimental approaches promise new avenues for measuring the causal impact of user behavior in online communities without adversely affecting users through randomized experiments. Tony Liu 0004, Lyle H. Ungar, Konrad P. Kording, Morgan McGuire |
ICWSM | 2 |
| 2023 | Different Affordances on Facebook and SMS Text Messaging Do Not Impede Generalization of Language-Based Predictive ModelsabstractAdaptive mobile device-based health interventions often use machine learning models trained on non-mobile device data, such as social media text, due to the difficulty and high expense of collecting large text message (SMS) data. Therefore, understanding the differences and generalization of models between these platforms is crucial for proper deployment. We examined the psycho-linguistic differences between Facebook and text messages, and their impact on out-of-domain model performance, using a sample of 120 users who shared both. We found that users use Facebook for sharing experiences (e.g., leisure) and SMS for task-oriented and conversational purposes (e.g., plan confirmations), reflecting the differences in the affordances. To examine the downstream effects of these differences, we used pre-trained Facebook-based language models to estimate age, gender, depression, life satisfaction, and stress on both Facebook and SMS. We found no significant differences in correlations between the estimates and self-reports across 6 of 8 models. These results suggest using pre-trained Facebook language models to achieve better accuracy with just-in-time interventions. Salvatore Giorgi, Xiangyu Tao, Sharath Chandra Guntuku, Douglas Bellew, Brenda Curtis, Lyle H. Ungar |
ICWSM | 7 |
| 2022 | Social Media Reveals Urban-Rural Differences in Stress across China
Jesse Cui, Tingdan Zhang, Kokil Jaidka, Dandan Pang, Garrick Sherman, Vinit Jakhetiya, Lyle H. Ungar, Sharath Chandra Guntuku |
ICWSM | 7 |
| 2022 | Correcting Sociodemographic Selection Biases for Population Prediction from Social Media
Salvatore Giorgi, Veronica E. Lynn, Farhan Ahmed, Sandra Matz, Lyle H. Ungar, H. Andrew Schwartz |
ICWSM | 6 |
| 2021 | Well-Being Depends on Social Comparison: Hierarchical Models of Twitter Language Suggest That Richer Neighbors Make You Less Happy
Salvatore Giorgi, Sharath Chandra Guntuku, Johannes C. Eichstaedt, Claire Pajot, H. Andrew Schwartz, Lyle H. Ungar |
ICWSM | 6 |
| 2020 | Beyond Positive Emotion: Deconstructing Happy Moments Based on Writing Prompts
Kokil Jaidka, Niyati Chhaya, Saran Mumick, Matthew Killingsworth, Alon Y. Halevy, Lyle H. Ungar |
ICWSM | 6 |
| 2019 | Understanding and Measuring Psychological Stress Using Social Media
Sharath Chandra Guntuku, Anneke Buffone, Kokil Jaidka, Johannes C. Eichstaedt, Lyle H. Ungar |
ICWSM | 5 |
| 2019 | Studying Cultural Differences in Emoji Usage across the East and the West
Sharath Chandra Guntuku, Louis Tay, Lyle H. Ungar |
ICWSM | 4 |
| 2019 | What Twitter Profile and Posted Images Reveal about Depression and Anxiety
Sharath Chandra Guntuku, Daniel Preotiuc-Pietro, Johannes C. Eichstaedt, Lyle H. Ungar |
ICWSM | 4 |
| 2018 | Multi-Attribute Topic Feature Construction for Social Media-based PredictionabstractThe effectiveness of social media-based prediction highly depends on whether we can construct effective content-based features based on social media text data. Features constructed based on topics learned using a topic model are very attractive due to their expressiveness in semantic representation and accommodation of inexact matching of semantically related words. We develop a novel general framework for constructing multi-attribute topic features using multi-views of the text data defined according to metadata attributes and study their effectiveness for a text-based prediction task. Furthermore we propose and study multiple weighting strategies to align text-based features and prediction outcomes. We evaluate the proposed method on a Twitter corpus of over 100 million tweets collected over a seven year period in 2009-2015 to predict human immunodeficiency virus (HIV) new diagnosis and other sexually transmitted infections (STIs) new diagnosis in the United States at the zipcode-level and county-level resolutions. The results show that feature representations based on attributes such as authors, locations, and hashtags are generally more effective than the conventional topic feature representation. Alex Morales, Nupoor Gandhi, Man-pui Sally Chan, Sophie Lohmann, Travis Sanchez, Kathleen A. Brady, Lyle H. Ungar, Dolores Albarracin, ChengXiang Zhai |
IEEE BigData | 7 |
| 2018 | Modeling and Visualizing Locus of Control with Facebook Language
Kokil Jaidka, Anneke Buffone, Johannes C. Eichstaedt, Masoud Rouhizadeh, Lyle H. Ungar |
ICWSM | 5 |
| 2018 | Facebook versus Twitter: Differences in Self-Disclosure and Trait Prediction
Kokil Jaidka, Sharath Chandra Guntuku, Lyle H. Ungar |
ICWSM | 3 |
| 2017 | Recognizing Pathogenic Empathy in Social Media
Muhammad Abdul-Mageed, Anneke Buffone, Johannes C. Eichstaedt, Lyle H. Ungar |
ICWSM | 5 |
| 2016 | Studying the Dark Triad of Personality through Twitter BehaviorabstractResearch into the darker traits of human nature is growing in interest especially in the context of increased social media usage. This allows users to express themselves to a wider online audience. We study the extent to which the standard model of dark personality -- the dark triad -- consisting of narcissism, psychopathy and Machiavellianism, is related to observable Twitter behavior such as platform usage, posted text and profile image choice. Our results show that we can map various behaviors to psychological theory and study new aspects related to social media usage. Finally, we build a machine learning algorithm that predicts the dark triad of personality in out-of-sample users with reliable accuracy. Daniel Preotiuc-Pietro, Jordan Carpenter, Salvatore Giorgi, Lyle H. Ungar |
CIKM | 4 |
| 2016 | Analyzing Personality through Social Media Profile Picture Choice
Liu Leqi, Daniel Preotiuc-Pietro, Zahra Riahi Samani, Mohsen Ebrahimi Moghaddam, Lyle H. Ungar |
ICWSM | 5 |
| 2013 | Characterizing Geographic Variation in Well-Being Using Tweets
H. Andrew Schwartz, Johannes C. Eichstaedt, Margaret L. Kern, Lukasz Dziurzynski, Richard E. Lucas, Megha Agrawal, Gregory J. Park, Shrinidhi K. Lakshmikanth, Sneha Jha, Martin E. P. Seligman, Lyle H. Ungar |
ICWSM | 11 |
| 2010 | Discovery of significant emerging trendsabstractWe describe a system that monitors social and mainstream media to determine shifts in what people are thinking about a product or company. We process over 100,000 news articles, blog posts, review sites, and tweets a day for mentions of items (e.g., products) of interest, extract phrases that are mentioned near them, and determine which of the phrases are of greatest possible interest to, for example, brand managers. Case studies show a good ability to rapidly pinpoint emerging subjects buried deep in large volumes of data and then highlight those that are rising or falling in significance as they relate to the firms interests. The tool and algorithm improves the signal-to-noise ratio and pinpoints precisely the opportunities and risks that matter most to communications professionals and their organizations. Saurabh Goorha, Lyle H. Ungar |
KDD | 2 |
| 2010 | Analyzing knowledge communities using foreground and background clustersabstractInsight into the growth (or shrinkage) of “knowledge communities” of authors that build on each other's work can be gained by studying the evolution over time of clusters of documents. We cluster documents based on the documents they cite in common using the Streemer clustering method, which finds cohesive foreground clusters (the knowledge communities) embedded in a diffuse background. We build predictive models with features based on the citation structure, the vocabulary of the papers, and the affiliations and prestige of the authors and use these models to study the drivers of community growth and the predictors of how widely a paper will be cited. We find that scientific knowledge communities tend to grow more rapidly if their publications build on diverse information and use narrow vocabulary and that papers that lie on the periphery of a community have the highest impact, while those not in any community have the lowest impact. Vasileios Kandylas, S. Phineas Upham, Lyle H. Ungar |
ACM Trans. Knowl. Discov. Data | 3 |
| 2009 | Resolving Identity Uncertainty with Learned Random WalksabstractA pervasive problem in large relational databases is identity uncertainty which occurs when multiple entries in a database refer to the same underlying entity in the world. Relational databases exhibit rich graphical structure and are naturally modeled as graphs whose nodes represent entities and whose typed-edges represent relations between them. We propose using random walk models for resolving identity uncertainty since they have proven effective for finding points which are proximately located in a network. Because not all types of relations are equally helpful in alleviating identity uncertainty, we develop a supervised approach to learning the usefulness of different database relations from a training set of database entries whose true identities are known. When tested on the task of resolving uncertainty of ambiguously named authors in bibliographical data, the learned random walk models yield performance superior to support vector machines, and to a related spectral clustering method. Ted Sandler, Lyle H. Ungar, Koby Crammer |
ICDM | 2 |
| 2009 | Multi-task Feature Selection Using the Multiple Inclusion Criterion (MIC)
Paramveer S. Dhillon, Brian Tomasik, Dean P. Foster, Lyle H. Ungar |
ECML/PKDD (1) | 4 |
| 2008 | Using sequence classification for filtering web pagesabstractWeb pages often contain text that is irrelevant to their main content, such as advertisements, generic format elements, and references to other pages on the same site. When used by automatic content-processing systems, e.g., for Web indexing, text classification, or information extraction, this irrelevant text often produces substantial amount of noise. This paper describes a trainable filtering system based on a feature-rich sequence classifier that removes irrelevant parts from pages, while keeping the content intact. Most of the features the system uses are purely form-related: HTML tags and their positions, sizes of elements, etc. This keeps the system general and domain-independent. We also experiment with content words and show that while they perform very poorly alone, they can slightly improve the performance of pure-form features, without jeopardizing the domain-independence. Our system achieves very high accuracy (95% and above) on several collections of Web pages. We also do a series of tests with different features and different classifiers, comparing the contribution of different components to the system performance, and comparing two known sequence classifiers, Robust Risk Minimization (RRM) and Conditional Random Fields (CRF), in a novel setting. Binyamin Rosenfeld, Ronen Feldman, Lyle H. Ungar |
CIKM | 3 |
| 2008 | Web-scale named entity recognitionabstractAutomatic recognition of named entities such as people, places, organizations, books, and movies across the entire web presents a number of challenges, both of scale and scope. Data for training general named entity recognizers is difficult to come by, and efficient machine learning methods are required once we have found hundreds of millions of labeled observations. We present an implemented system that addresses these issues, including a method for automatically generating training data, and a multi-class online classification training method that learns to recognize not only high level categories such as place and person, but also more fine-grained categories such as soccer players, birds, and universities. The resulting system gives precision and recall performance comparable to that obtained for more limited entity types in much more structured domains such as company recognition in newswire, even though web documents often lack consistent capitalization and grammatical sentence construction. Casey Whitelaw, Alexander Kehlenbeck, Nemanja Petrovic, Lyle H. Ungar |
CIKM | 4 |
| 2008 | Efficient Feature Selection in the Presence of Multiple Feature ClassesabstractWe present an information theoretic approach to feature selection when the data possesses feature classes. Feature classes are pervasive in real data. For example, in gene expression data, the genes which serve as features may be divided into classes based on their membership in gene families or pathways. When doing word sense disambiguation or named entity extraction, features fall into classes including adjacent words, their parts of speech, and the topic and venue of the document the word is in. When predictive features occur predominantly in a small number of feature classes, our information theoretic approach significantly improves feature selection. Experiments on real and synthetic data demonstrate substantial improvement in predictive accuracy over the standard L0penalty-based stepwise and stream wise feature selection methods as well as over Lasso and Elastic Nets, all of which are oblivious to the existence of feature classes. Paramveer S. Dhillon, Dean P. Foster, Lyle H. Ungar |
ICDM | 3 |
| 2008 | Finding cohesive clusters for analyzing knowledge communities
Vasileios Kandylas, S. Phineas Upham, Lyle H. Ungar |
Knowl. Inf. Syst. | 3 |
| 2007 | Extracting Product Comparisons from Discussion BoardsabstractIn recent years, product discussion forums have become a rich environment in which consumers and potential adopters exchange views and information. Researchers and practitioners are starting to extract user sentiment about products from user product reviews. Users often compare different products, stating which they like better and why. Extracting information about product comparisons offers a number of challenges; recognizing and normalizing entities (products) in the informal language of blogs and discussion groups require different techniques than those used for entity extraction in the more formal text of newspapers and scientific articles. We present a case study in extracting information about comparisons between running shoes and between cars, describe an effective methodology, and show how it produces insight into how consumers view the running shoe and car markets. Ronen Feldman, Moshe Fresko, Jacob Goldenberg, Oded Netzer, Lyle H. Ungar |
ICDM | 5 |
| 2007 | Finding Cohesive Clusters for Analyzing Knowledge CommunitiesabstractDocuments and authors can be clustered into "knowledge communities" based on the overlap in the papers they cite. We introduce a new clustering algorithm, Streemer, which finds cohesive foreground clusters embedded in a diffuse background, and use it to identify knowledge communities as foreground clusters of papers which share common citations. To analyze the evolution of these communities over time, we build predictive models with features based on the citation structure, the vocabulary of the papers, and the affiliations and prestige of the authors. Findings include that scientific knowledge communities tend to grow more rapidly if their publications build on diverse information and if they use a narrow vocabulary. Vasileios Kandylas, S. Phineas Upham, Lyle H. Ungar |
ICDM | 3 |
| 2005 | Streaming feature selection using alpha-investingabstractIn Streaming Feature Selection (SFS), new features are sequentially considered for addition to a predictive model. When the space of potential features is large, SFS offers many advantages over traditional feature selection methods, which assume that all features are known in advance. Features can be generated dynamically, focusing the search for new features on promising subspaces, and overfitting can be controlled by dynamically adjusting the threshold for adding features to the model. We describe α-investing, an adaptive complexity penalty method for SFS which dynamically adjusts the threshold on the error reduction required for adding a new feature. α-investing gives false discovery rate-style guarantees against overfitting. It differs from standard penalty methods such as AIC, BIC or RIC, which always drastically over- or under-fit in the limit of infinite numbers of non-predictive features. Empirical results show that SFS is competitive with much more compute-intensive feature selection methods such as stepwise regression, and allows feature selection on problems with over a million potential features. Dean P. Foster, Robert A. Stine, Lyle H. Ungar |
KDD | 4 |
| 2004 | Cluster-based concept invention for statistical relational learningabstractWe use clustering to derive new relations which augment database schema used in automatic generation of predictive features in statistical relational learning. Entities derived from clusters increase the expressivity of feature spaces by creating new first-class concepts which contribute to the creation of new features. For example, in CiteSeer, papers can be clustered based on words or citations giving "topics", and authors can be clustered based on documents they co-author giving "communities". Such cluster-derived concepts become part of more complex feature expressions. Out of the large number of generated features, those which improve predictive accuracy are kept in the model, as decided by statistical feature selection criteria. We present results demonstrating improved accuracy on two tasks, venue prediction and link prediction, using CiteSeer data. Alexandrin Popescul, Lyle H. Ungar |
KDD | 2 |
| 2003 | Statistical Relational Learning for Document MiningabstractA major obstacle to fully integrated deployment of many data mining algorithms is the assumption that data sits in a single table, even though most real-world databases have complex relational structures. We propose an integrated approach to statistical modelling from relational databases. We structure the search space based on "refinement graphs", which are widely used in inductive logic programming for learning logic descriptions. The use of statistics allows us to extend the search space to include richer set of features, including many which are not Boolean. Search and model selection are integrated into a single process, allowing information criteria native to the statistical model, for example logistic regression, to make feature selection decisions in a step-wise manner. We present experimental results for the task of predicting where scientific papers will be published based on relational data taken from CiteSeer. Our approach results in classification accuracies superior to those achieved when using classical "flat" features. The resulting classifier can be used to recommend where to publish articles. Alexandrin Popescul, Lyle H. Ungar, Steve Lawrence, David M. Pennock |
ICDM | 2 |
| 2002 | Methods and metrics for cold-start recommendationsabstractWe have developed a method for recommending items that combines content and collaborative data under a single probabilistic framework. We benchmark our algorithm against a naïve Bayes classifier on the cold-start problem, where we wish to recommend items that no one in the community has yet rated. We systematically explore three testing methodologies using a publicly available data set, and explain how these methods apply to specific real-world applications. We advocate heuristic recommenders when benchmarking to give competent baseline performance. We introduce a new performance metric, the CROC curve, and demonstrate empirically that the various components of our testing strategy combine to obtain deeper understanding of the performance characteristics of recommender systems. Though the emphasis of our testing is on cold-start recommending, our methods for recommending and evaluation are general. Andrew I. Schein, Alexandrin Popescul, Lyle H. Ungar, David M. Pennock |
SIGIR | 3 |
| 2000 | Efficient clustering of high-dimensional data sets with application to reference matchingabstractMany important problems involve clustering large datasets. Although naive implementations of clustering are computationally expensive, there are established efficient techniques for clustering when the dataset has either (1) a limited number of clusters, (2) a low feature dimensionality, or (3) a small number of data points. However, there has been much less work on methods of efficiently clustering datasets that are large in all three ways at once---for example, having millions of data points that exist in many thousands of dimensions representing many thousands of clusters. We present a new technique for clustering these large, high-dimensional datasets. The key idea involves using a cheap, approximate distance measure to efficiently divide the data into overlapping subsets we call canopies. Then clustering is performed by measuring exact distances only between points that occur in a common canopy. Using canopies, large clustering problems that were formerly impossible become practical. U... Andrew McCallum, Kamal Nigam, Lyle H. Ungar |
KDD | 3 |