Daniel Preotiuc-Pietro

dblp:126/8668 · DBLP profile ↗
← Back
42ranked-venue papers
10as first author
14since 2021 · last 2025
0000-0002-4504-0212ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 10 first-author · 13 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2025 Improving Instruct Models for Free: A Study on Partial Adaptation
abstract
Instruct models, obtained from various instruction tuning or post-training steps, are commonly deemed superior and more usable than their base counterpart.While the model gains instruction following ability, instruction tuning may lead to forgetting the knowledge from pre-training or it may encourage the model to become overly conversational or verbose.This, in turn, can lead to degradation of in-context few-shot learning performance.In this work, we study the performance trajectory between base and instruct models by scaling down the strength of instruction-tuning via the partial adaption method.We show that, across several model families and model sizes, reducing the strength of instruction-tuning results in material improvement on a few-shot in-context learning benchmark covering a variety of classic natural language tasks.This comes at the cost of losing some degree of instruction following ability as measured by AlpacaEval.Our study shines light on the potential trade-off between in-context learning and instruction following abilities that is worth considering in practice.
Ozan Irsoy, Pengxiang Cheng 0001, Jennifer L. Chen, Daniel Preotiuc-Pietro, Shiyue Zhang 0001, Duccio Pappadopulo
EMNLP4
2025 Calibrating LLMs for Text-to-SQL Parsing by Leveraging Sub-clause Frequencies
abstract
While large language models (LLMs) achieve strong performance on text-to-SQL parsing, they sometimes exhibit unexpected failures in which they are confidently incorrect.Building trustworthy text-to-SQL systems thus requires eliciting reliable uncertainty measures from the LLM.In this paper, we study the problem of providing a calibrated confidence score that conveys the likelihood of an output query being correct.Our work is the first to establish a benchmark for post-hoc calibration of LLMbased text-to-SQL parsing.In particular, we show that Platt scaling, a canonical method for calibration, provides substantial improvements over directly using raw model output probabilities as confidence scores.Furthermore, we propose a method for text-to-SQL calibration that leverages the structured nature of SQL queries to provide more granular signals of correctness, named "sub-clause frequency" (SCF) scores.Using multivariate Platt scaling (MPS), our extension of the canonical Platt scaling technique, we combine individual SCF scores into an overall accurate and calibrated score.Empirical evaluation on two popular text-to-SQL datasets shows that our approach of combining MPS and SCF yields further improvements in calibration and the related task of error detection over traditional Platt scaling.
Terrance Liu, Daniel Preotiuc-Pietro, Yash Chandarana
EMNLP3
2025 STARQA: A Question Answering Dataset for Complex Analytical Reasoning over Structured Databases
abstract
Semantic parsing methods for converting text to SQL queries enable question answering over structured data and can greatly benefit analysts who routinely perform complex analytics on vast data stored in specialized relational databases.Although several benchmarks measure the abilities of text to SQL, the complexity of their questions is inherently limited by the level of expressiveness in query languages and none focus explicitly on questions involving complex analytical reasoning which require operations such as calculations over aggregate analytics, time series analysis or scenario understanding.In this paper, we introduce STARQA, the first public humancreated dataset of complex analytical reasoning questions and answers on three specializeddomain databases.In addition to generating SQL directly using LLMs, we evaluate a novel approach (TEXT2SQLCODE) that decomposes the task into a combination of SQL and Python: SQL is responsible for data fetching, and Python more naturally performs reasoning.Our results demonstrate that identifying and combining the abilities of SQL and Python is beneficial compared to using SQL alone, yet the dataset still remains quite challenging for the existing state-of-the-art LLMs.
Mounica Maddela, Lingjue Xie, Daniel Preotiuc-Pietro, Mausam
EMNLP3
2025 An Alternative to FLOPS Regularization to Effectively Productionize SPLADE-Doc
abstract
Learned Sparse Retrieval (LSR) models encode text as weighted term vectors, which need to be sparse to leverage inverted index structures during retrieval. SPLADE, the most popular LSR model, uses FLOPS regularization to encourage vector sparsity during training. However, FLOPS regularization does not ensure sparsity among terms-only within a given query or document. Terms with very high Document Frequencies (DFs) substantially increase latency in production retrieval engines, such as Apache Solr, due to their lengthy posting lists. To address the issue of high DFs, we present a new variant of FLOPS regularization: DF-FLOPS. This new regularization technique penalizes the usage of high-DF terms, thereby shortening posting lists and reducing retrieval latency. Unlike other inference-time sparsification methods, such as stopword removal, DF-FLOPS regularization allows for the selective inclusion of high-frequency terms in cases where the terms are truly salient. We find that DF-FLOPS successfully reduces the prevalence of high-DF terms and lowers retrieval latency (around 10x faster) in a production-grade engine while maintaining effectiveness both in-domain (only a 2.2-point drop in MRR@10) and cross-domain (improved performance in 12 out of 13 tasks on which we tested). With retrieval latencies on par with BM25, this work provides an important step towards making LSR practical for deployment in production-grade search engines.
Aldo Porco, Dhruv Mehra, Igor Malioutov, Karthik Radhakrishnan, Moniba Keymanesh, Daniel Preotiuc-Pietro, Sean MacAvaney, Pengxiang Cheng 0001
SIGIR6
2024 Who Is Bragging More Online? A Large Scale Analysis of Bragging in Social Media
abstract
Bragging is the act of uttering statements that are likely to be positively viewed by others and it is extensively employed in human communication with the aim to build a positive self-image of oneself. Social media is a natural platform for users to employ bragging in order to gain admiration, respect, attention and followers from their audiences. Yet, little is known about the scale of bragging online and its characteristics. This paper employs computational sociolinguistics methods to conduct the first large scale study of bragging behavior on Twitter (U.S.) by focusing on its overall prevalence, temporal dynamics and impact of demographic factors. Our study shows that the prevalence of bragging decreases over time within the same population of users. In addition, younger, more educated and popular users in the U.S. are more likely to brag. Finally, we conduct an extensive linguistics analysis to unveil specific bragging themes associated with different user traits.
Mali Jin, Daniel Preotiuc-Pietro, A. Seza Dogruöz, Nikolaos Aletras
LREC/COLING2
2024 Unsupervised Contrast-Consistent Ranking with Language Models
abstract
Niklas Stoehr, Pengxiang Cheng, Jing Wang, Daniel Preotiuc-Pietro, Rajarshi Bhowmik. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Niklas Stoehr, Pengxiang Cheng 0001, Jing Wang 0069, Daniel Preotiuc-Pietro, Rajarshi Bhowmik
EACL (1)4
2023 Towards a Unified Multi-Domain Multilingual Named Entity Recognition Model
abstract
Mayank Kulkarni, Daniel Preotiuc-Pietro, Karthik Radhakrishnan, Genta Indra Winata, Shijie Wu, Lingjue Xie, Shaohua Yang. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Mayank Kulkarni, Daniel Preotiuc-Pietro, Karthik Radhakrishnan, Genta Indra Winata, Lingjue Xie
EACL2
2023 EntSUMv2: Dataset, Models and Evaluation for More Abstractive Entity-Centric Summarization
abstract
Entity-centric summarization is a form of controllable summarization that aims to generate a summary for a specific entity given a document.Concise summaries are valuable in various reallife applications, as they enable users to quickly grasp the main points of the document focusing on an entity of interest.This paper presents ENTSUMV2, a more abstractive version of the original entity-centric ENTSUM summarization dataset.In ENTSUMV2 the annotated summaries are intentionally made shorter to benefit more specific and useful entity-centric summaries for downstream users.We conduct extensive experiments on this dataset using multiple abstractive summarization approaches that employ supervised fine-tuning or large-scale instruction tuning.Additionally, we perform comprehensive human evaluation that incorporates metrics for measuring crucial facets.These metrics provide a more fine-grained interpretation of the current state-of-the-art systems and highlight areas for future improvement.
Dhruv Mehra, Lingjue Xie, Ella Hofmann-Coyle, Mayank Kulkarni, Daniel Preotiuc-Pietro
EMNLP5
2023 Dataless Knowledge Fusion by Merging Weights of Language Models
Xisen Jin, Xiang Ren 0001, Daniel Preotiuc-Pietro, Pengxiang Cheng 0001
ICLR3
2023 Analyzing and Predicting Persistence of News Tweets
abstract
Maggie Liu, Jing Wang, Daniel Preotiuc-Pietro. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Maggie Liu, Jing Wang 0069, Daniel Preotiuc-Pietro
IJCNLP (1)3
2022 Automatic Identification and Classification of Bragging in Social Media
abstract
Bragging is a speech act employed with the goal of constructing a favorable self-image through positive statements about oneself.It is widespread in daily communication and especially popular in social media, where users aim to build a positive image of their persona directly or indirectly.In this paper, we present the first large scale study of bragging in computational linguistics, building on previous research in linguistics and pragmatics.To facilitate this, we introduce a new publicly available data set of tweets annotated for bragging and their types.We empirically evaluate different transformerbased models injected with linguistic information in (a) binary bragging classification, i.e., if tweets contain bragging statements or not; and (b) multi-class bragging type prediction including not bragging.Our results show that our models can predict bragging with macro F1 up to 72.42 and 35.95 in the binary and multi-class classification tasks respectively.Finally, we present an extensive linguistic and error analysis of bragging prediction to guide future research on this topic.1
Mali Jin, Daniel Preotiuc-Pietro, A. Seza Dogruöz, Nikolaos Aletras
ACL (1)2
2022 EntSUM: A Data Set for Entity-Centric Extractive Summarization
abstract
Controllable summarization aims to provide summaries that take into account userspecified aspects and preferences to better assist them with their information need, as opposed to the standard summarization setup which build a single generic summary of a document.We introduce a human-annotated data set (ENTSUM) for controllable summarization with a focus on named entities as the aspects to control.We conduct an extensive quantitative analysis to motivate the task of entity-centric summarization and show that existing methods for controllable summarization fail to generate entity-centric summaries.We propose extensions to state-of-the-art summarization approaches that achieve substantially better results on our data set.Our analysis and results show the challenging nature of this task and of the proposed data set.12
Mounica Maddela, Mayank Kulkarni, Daniel Preotiuc-Pietro
ACL (1)3
2022 Combining Humor and Sarcasm for Improving Political Parody Detection
abstract
Xiao Ao, Danae Sanchez Villegas, Daniel Preotiuc-Pietro, Nikolaos Aletras. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Xiao Ao, Danae Sanchez Villegas, Daniel Preotiuc-Pietro, Nikolaos Aletras
NAACL-HLT3
2021 Identifying Named Entities as they are Typed
abstract
Identifying named entities in written text is an essential component of the text processing pipeline used in applications such as text editors to gain a better understanding of the semantics of the text.However, the typical experimental setup for evaluating Named Entity Recognition (NER) systems is not directly applicable to systems that process text in real time as the text is being typed.Evaluation is performed on a sentence level assuming the end-user is willing to wait until the entire sentence is typed for entities to be identified and further linked to identifiers or coreferenced.We introduce a novel experimental setup for NER systems for applications where decisions about named entity boundaries need to be performed in an online fashion.We study how state-of-the-art methods perform under this setup in multiple languages and propose adaptations to these models to suit this new experimental setup.Experimental results show that the best systems that are evaluated on each token after its typed, reach performance within 1-5 F 1 points of systems that are evaluated at the end of the sentence.These show that entity recognition can be performed in this setup and open up the development of other NLP tools in a similar setup.
Ravneet Arora, Chen-Tse Tsai, Daniel Preotiuc-Pietro
EACL3
2020 Analyzing Political Parody in Social Media
abstract
Parody is a figurative device used to imitate an entity for comedic or critical purposes and represents a widespread phenomenon in social media through many popular parody accounts.In this paper, we present the first computational study of parody.We introduce a new publicly available data set of tweets from real politicians and their corresponding parody accounts.We run a battery of supervised machine learning models for automatically detecting parody tweets with an emphasis on robustness by testing on tweets from accounts unseen in training, across different genders and across countries.Our results show that political parody tweets can be predicted with an accuracy up to 90%.Finally, we identify the markers of parody through a linguistic analysis.Beyond research in linguistics and political communication, accurately and automatically detecting parody is important to improving fact checking for journalists and analytics such as sentiment analysis through filtering out parodical utterances. 1
Antonis Maronikolakis, Danae Sanchez Villegas, Daniel Preotiuc-Pietro, Nikolaos Aletras
ACL3
2020 Temporally-Informed Analysis of Named Entity Recognition
abstract
Natural language processing models often have to make predictions on text data that evolves over time as a result of changes in language use or the information described in the text. However, evaluation results on existing data sets are seldom reported by taking the timestamp of the document into account. We analyze and propose methods that make better use of temporally-diverse training data, with a focus on the task of named entity recognition. To support these experiments, we introduce a novel data set of English tweets annotated with named entities. We empirically demonstrate the effect of temporal drift on performance, and how the temporal information of documents can be used to obtain better models compared to those that disregard temporal information. Our analysis gives insights into why this information is useful, in the hope of informing potential avenues of improvement for named entity recognition as well as other NLP tasks under similar experimental setups.
Shruti Rijhwani, Daniel Preotiuc-Pietro
ACL2
2020 Multi-Domain Named Entity Recognition with Genre-Aware and Agnostic Inference
abstract
Named entity recognition is a key component of many text processing pipelines and it is thus essential for this component to be robust to different types of input.However, domain transfer of NER models with data from multiple genres has not been widely studied.To this end, we conduct NER experiments in three predictive setups on data from: a) multiple domains; b) multiple domains where the genre label is unknown at inference time; c) domains not encountered in training.We introduce a new architecture tailored to this task by using shared and private domain parameters and multi-task learning.This consistently outperforms all other baseline and competitive methods on all three experimental setups, with differences ranging between +1.95 to +3.11 average F1 across multiple genres when compared to standard approaches.These results illustrate the challenges that need to be taken into account when building real-world NLP applications that are robust to various types of text and the methods that can help, at least partially, alleviate these issues.
Jing Wang 0069, Mayank Kulkarni, Daniel Preotiuc-Pietro
ACL3
2020 Fact vs. Opinion: the Role of Argumentation Features in News Classification
abstract
A 2018 study led by the Media Insight Project showed that most journalists think that a clear marking of what is news reporting and what is commentary or opinion (e.g., editorial, op-ed) is essential for gaining public trust.We present an approach to classify news articles into news stories (i.e., reporting of factual information) and opinion pieces using models that aim to supplement the article content representation with argumentation features.Our hypothesis is that the nature of argumentative discourse is important in distinguishing between news stories and opinion articles.We show that argumentation features outperform linguistic features used previously and improve on fine-tuned transformer-based models when tested on data from publishers unseen in training.
Tariq Alhindi, Smaranda Muresan, Daniel Preotiuc-Pietro
COLING3
2019 Predicting and Analyzing Language Specificity in Social Media Posts
abstract
In computational linguistics, specificity quantifies how much detail is engaged in text. It is an important characteristic of speaker intention and language style, and is useful in NLP applications such as summarization and argumentation mining. Yet to date, expert-annotated data for sentence-level specificity are scarce and confined to the news genre. In addition, systems that predict sentence specificity are classifiers trained to produce binary labels (general or specific).We collect a dataset of over 7,000 tweets annotated with specificity on a fine-grained scale. Using this dataset, we train a supervised regression model that accurately estimates specificity in social media posts, reaching a mean absolute error of 0.3578 (for ratings on a scale of 1-5) and 0.73 Pearson correlation, significantly improving over baselines and previous sentence specificity prediction systems. We also present the first large-scale study revealing the social, temporal and mental health factors underlying language specificity on social media.
Daniel Preotiuc-Pietro, Junyi Jessy Li
AAAI3
2019 Multi-task Pairwise Neural Ranking for Hashtag Segmentation
abstract
Hashtags are often employed on social media and beyond to add metadata to a textual utterance with the goal of increasing discoverability, aiding search, or providing additional semantics.However, the semantic content of hashtags is not straightforward to infer as these represent ad-hoc conventions which frequently include multiple words joined together and can include abbreviations and unorthodox spellings.We build a dataset of 12,594 hashtags split into individual segments and propose a set of approaches for hashtag segmentation by framing it as a pairwise ranking problem between candidate segmentations. 1 Our novel neural approaches demonstrate 24.6% error reduction in hashtag segmentation accuracy compared to the current state-of-the-art method.Finally, we demonstrate that a deeper understanding of hashtag semantics obtained through segmentation is useful for downstream applications such as sentiment analysis, for which we achieved a 2.6% increase in average recall on the Se-mEval 2017 sentiment analysis dataset.
Mounica Maddela, Wei Xu 0004, Daniel Preotiuc-Pietro
ACL (1)3
2019 Analyzing Linguistic Differences between Owner and Staff Attributed Tweets
abstract
Research on social media has to date assumed that all posts from an account are authored by the same person.In this study, we challenge this assumption and study the linguistic differences between posts signed by the account owner or attributed to their staff.We introduce a novel data set of tweets posted by U.S. politicians who self-reported their tweets using a signature.We analyze the linguistic topics and style features that distinguish the two types of tweets.Predictive results show that we are able to distinguish between owner and staff attributed tweets with good accuracy, even when not using any training data from that account.
Daniel Preotiuc-Pietro, Rita Devlin Marier
ACL (1)1
2019 Automatically Identifying Complaints in Social Media
abstract
Complaining is a basic speech act regularly used in human and computer mediated communication to express a negative mismatch between reality and expectations in a particular situation.Automatically identifying complaints in social media is of utmost importance for organizations or brands to improve the customer experience or in developing dialogue systems for handling and responding to complaints.In this paper, we introduce the first systematic analysis of complaints in computational linguistics.We collect a new annotated data set of written complaints expressed in English on Twitter. 1 We present an extensive linguistic analysis of complaining as a speech act in social media and train strong feature-based and neural models of complaints across nine domains achieving a predictive performance of up to 79 F1 using distant supervision.
Daniel Preotiuc-Pietro, Mihaela Gaman, Nikolaos Aletras
ACL (1)1
2019 Categorizing and Inferring the Relationship between the Text and Image of Twitter Posts
abstract
Text in social media posts is frequently accompanied by images in order to provide content, supply context, or to express feelings.This paper studies how the meaning of the entire tweet is composed through the relationship between its textual content and its image.We build and release a data set of image tweets annotated with four classes which express whether the text or the image provides additional information to the other modality.We show that by combining the text and image information, we can build a machine learning approach that accurately distinguishes between the relationship types.Further, we derive insights into how these relationships are materialized through text and image content analysis and how they are impacted by user demographic traits.These methods can be used in several downstream applications including pre-training image tagging models, collecting distantly supervised data for image captioning, and can be directly used in end-user applications to optimize screen estate.
Alakananda Vempala, Daniel Preotiuc-Pietro
ACL (1)2
2019 What Twitter Profile and Posted Images Reveal about Depression and Anxiety
Sharath Chandra Guntuku, Daniel Preotiuc-Pietro, Johannes C. Eichstaedt, Lyle H. Ungar
ICWSM2
2018 Expressively vulgar: The socio-dynamics of vulgarity and its effects on sentiment analysis in social media
abstract
Vulgarity is a common linguistic expression and is used to perform several linguistic functions. Understanding their usage can aid both linguistic and psychological phenomena as well as benefit downstream natural language processing applications such as sentiment analysis. This study performs a large-scale, data-driven empirical analysis of vulgar words using social media data. We analyze the socio-cultural and pragmatic aspects of vulgarity using tweets from users with known demographics. Further, we collect sentiment ratings for vulgar tweets to study the relationship between the use of vulgar words and perceived sentiment and show that explicitly modeling vulgar words can boost sentiment analysis performance.
Isabel Cachola, Eric Holgate, Daniel Preotiuc-Pietro, Junyi Jessy Li
COLING3
2018 User-Level Race and Ethnicity Predictors from Twitter Text
abstract
User demographic inference from social media text has the potential to improve a range of downstream applications, including real-time passive polling or quantifying demographic bias. This study focuses on developing models for user-level race and ethnicity prediction. We introduce a data set of users who self-report their race/ethnicity through a survey, in contrast to previous approaches that use distantly supervised data or perceived labels. We develop predictive models from text which accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC and make these available to the research community.
Daniel Preotiuc-Pietro, Lyle H. Ungar
COLING1
2018 The Remarkable Benefit of User-Level Aggregation for Lexical-based Population-Level Predictions
abstract
Nowcasting based on social media text promises to provide unobtrusive and near real-time predictions of community-level outcomes.These outcomes are typically regarding people, but the data is often aggregated without regard to users in the Twitter populations of each community.This paper describes a simple yet effective method for building community-level models using Twitter language aggregated by user.Results on four different U.S. county-level tasks, spanning demographic, health, and psychological outcomes show large and consistent improvements in prediction accuracies (e.g. from Pearson r = .73to .82 for median income prediction or r = .37 to .47 for life satisfaction prediction) over the standard approach of aggregating all tweets.We make our aggregated and anonymized community-level data, derived from 37 billion tweets -over 1 billion of which were mapped to counties, available for research.
Salvatore Giorgi, Daniel Preotiuc-Pietro, Anneke Buffone, Daniel Rieman, Lyle H. Ungar, H. Andrew Schwartz
EMNLP2
2018 Why Swear? Analyzing and Inferring the Intentions of Vulgar Expressions
abstract
Vulgar words are employed in language use for several different functions, ranging from expressing aggression to signaling group identity or the informality of the communication.This versatility of usage of a restricted set of words is challenging for downstream applications and has yet to be studied quantitatively or using natural language processing techniques.We introduce a novel data set of 7,800 tweets from users with known demographic traits where all instances of vulgar words are annotated with one of the six categories of vulgar word use.Using this data set, we present the first analysis of the pragmatic aspects of vulgarity and how they relate to social factors.We build a model able to predict the category of a vulgar word based on the immediate context it appears in with 67.4 macro F1 across six classes.Finally, we demonstrate the utility of modeling the type of vulgar word use in context by using this information to achieve state-of-the-art performance in hate speech detection on a benchmark data set.
Eric Holgate, Isabel Cachola, Daniel Preotiuc-Pietro, Junyi Jessy Li
EMNLP3
2017 Beyond Binary Labels: Political Ideology Prediction of Twitter Users
abstract
Automatic political preference prediction from social media posts has to date proven successful only in distinguishing between publicly declared liberals and conservatives in the US.This study examines users' political ideology using a sevenpoint scale which enables us to identify politically moderate and neutral usersgroups which are of particular interest to political scientists and pollsters.Using a novel data set with political ideology labels self-reported through surveys, our goal is two-fold: a) to characterize the political groups of users through language use on Twitter; b) to build a fine-grained model that predicts political ideology of unseen users.Our results identify differences in both political leaning and engagement and the extent to which each group tweets using political keywords.Finally, we demonstrate how to improve ideology prediction accuracy by exploiting the relationships between the user groups.
Daniel Preotiuc-Pietro, Ye Liu 0002, Daniel Hopkins, Lyle H. Ungar
ACL (1)1
2017 Controlling Human Perception of Basic User Traits
abstract
Much of our online communication is textmediated and, lately, more common with automated agents.Unlike interacting with humans, these agents currently do not tailor their language to the type of person they are communicating to.In this pilot study, we measure the extent to which human perception of basic user trait information -gender and age -is controllable through text.Using automatic models of gender and age prediction, we estimate which tweets posted by a user are more likely to mis-characterize his traits.We perform multiple controlled crowdsourcing experiments in which we show that we can reduce the human prediction accuracy of gender to almost random -an over 20% drop in accuracy.Our experiments show that it is practically feasible for multiple applications such as text generation, text summarization or machine translation to be tailored to specific traits and perceived as such.
Daniel Preotiuc-Pietro, Sharath Chandra Guntuku, Lyle H. Ungar
EMNLP1
2017 Sub-story detection in Twitter with hierarchical Dirichlet processes
abstract
Social media has now become the de facto information source on real world events. The challenge, however, due to the high volume and velocity nature of social media streams, is in how to follow all posts pertaining to a given event over time – a task referred to as story detection. Moreover, there are often several different stories pertaining to a given event, which we refer to as sub-stories and the corresponding task of their automatic detection – as sub-story detection. This paper proposes hierarchical Dirichlet processes (HDP), a probabilistic topic model, as an effective method for automatic sub-story detection. HDP can learn sub-topics associated with sub-stories which enables it to handle subtle variations in sub-stories. It is compared with state-of-the-art story detection approaches based on locality sensitive hashing and spectral clustering. We demonstrate the superior performance of HDP for sub-story detection on real world Twitter data sets using various evaluation measures. The ability of HDP to learn sub-topics helps it to recall the sub-stories with high precision. This has resulted in an improvement of up to 60% in the F-score performance of HDP based sub-story detection approach compared to standard story detection approaches. A similar performance improvement is also seen using an information theoretic evaluation measure proposed for the sub-story detection task. Another contribution of this paper is in demonstrating that considering the conversational structures within the Twitter stream can bring up to 200% improvement in sub-story detection performance.
P. K. Srijith, Mark Hepple, Kalina Bontcheva, Daniel Preotiuc-Pietro
Inf. Process. Manag.4
2016 Discovering User Attribute Stylistic Differences via Paraphrasing
abstract
User attribute prediction from social media text has proven successful and useful for downstream tasks. In previous studies, differences in user trait language use have been limited primarily to the presence or absence of words that indicate topical preferences. In this study, we aim to find linguistic style distinctions across three different user attributes: gender, age and occupational class. By combining paraphrases with a simple yet effective method, we capture a wide set of stylistic differences that are exempt from topic bias. We show their predictive power in user profiling, conformity with human perception and psycholinguistic hypotheses, and potential use in generating natural language tailored to specific user traits.
Daniel Preotiuc-Pietro, Wei Xu 0004, Lyle H. Ungar
AAAI1
2016 Analyzing Biases in Human Perception of User Age and Gender from Text
abstract
Lucie Flekova, Jordan Carpenter, Salvatore Giorgi, Lyle Ungar, Daniel Preoţiuc-Pietro. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016.
Lucie Flek, Jordan Carpenter, Salvatore Giorgi, Lyle H. Ungar, Daniel Preotiuc-Pietro
ACL (1)5
2016 Studying the Dark Triad of Personality through Twitter Behavior
abstract
Research into the darker traits of human nature is growing in interest especially in the context of increased social media usage. This allows users to express themselves to a wider online audience. We study the extent to which the standard model of dark personality -- the dark triad -- consisting of narcissism, psychopathy and Machiavellianism, is related to observable Twitter behavior such as platform usage, posted text and profile image choice. Our results show that we can map various behaviors to psychological theory and study new aspects related to social media usage. Finally, we build a machine learning algorithm that predicts the dark triad of personality in out-of-sample users with reliable accuracy.
Daniel Preotiuc-Pietro, Jordan Carpenter, Salvatore Giorgi, Lyle H. Ungar
CIKM1
2016 Analyzing Personality through Social Media Profile Picture Choice
Liu Leqi, Daniel Preotiuc-Pietro, Zahra Riahi Samani, Mohsen Ebrahimi Moghaddam, Lyle H. Ungar
ICWSM2
2016 An Empirical Exploration of Moral Foundations Theory in Partisan News Sources
Dean Fulgoni, Jordan Carpenter, Lyle H. Ungar, Daniel Preotiuc-Pietro
LREC4
2016 Studying the Temporal Dynamics of Word Co-occurrences: An Application to Event Detection
Daniel Preotiuc-Pietro, P. K. Srijith, Mark Hepple, Trevor Cohn
LREC1
2015 An analysis of the user occupational class through Twitter content
abstract
Daniel Preoţiuc-Pietro, Vasileios Lampos, Nikolaos Aletras. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Daniel Preotiuc-Pietro, Vasileios Lampos, Nikolaos Aletras
ACL (1)1
2014 Predicting and Characterising User Impact on Twitter
abstract
The open structure of online social networks and their uncurated nature give rise to problems of user credibility and influence.In this paper, we address the task of predicting the impact of Twitter users based only on features under their direct control, such as usage statistics and the text posted in their tweets.We approach the problem as regression and apply linear as well as nonlinear learning methods to predict a user impact score, estimated by combining the numbers of the user's followers, followees and listings.The experimental results point out that a strong prediction performance is achieved, especially for models based on the Gaussian Processes framework.Hence, we can interpret various modelling components, transforming them into indirect 'suggestions' for impact boosting.
Vasileios Lampos, Nikolaos Aletras, Daniel Preotiuc-Pietro, Trevor Cohn
EACL3
2013 A user-centric model of voting intention from Social Media
Vasileios Lampos, Daniel Preotiuc-Pietro, Trevor Cohn
ACL (1)2
2013 A temporal model of text periodicities using Gaussian Processes
abstract
Temporal variations of text are usually ignored in NLP applications.However, text use changes with time, which can affect many applications.In this paper we model periodic distributions of words over time.Focusing on hashtag frequency in Twitter, we first automatically identify the periodic patterns.We use this for regression in order to forecast the volume of a hashtag based on past data.We use Gaussian Processes, a state-ofthe-art bayesian non-parametric model, with a novel periodic kernel.We demonstrate this in a text classification setting, assigning the tweet hashtag based on the rest of its text.This method shows significant improvements over competitive baselines.
Daniel Preotiuc-Pietro, Trevor Cohn
EMNLP1
2012 Unsupervised document zone identification using probabilistic graphical models
Andrea Varga, Daniel Preotiuc-Pietro, Fabio Ciravegna
LREC2