VLDB 2026 Research / reviewers in the wild / expert
Lingzi Hong
dblp:144/3339
· DBLP profile ↗
14ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0001-8412-8180ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Assessing the Human Likeness of AI-Generated CounterspeechabstractCounterspeech is a targeted response to counteract and challenge abusive or hateful content. It effectively curbs the spread of hatred and fosters constructive online communication. Previous studies have proposed different strategies for automatically generated counterspeech. Evaluations, however, focus on relevance, surface form, and other shallow linguistic characteristics. This paper investigates the human likeness of AI-generated counterspeech, a critical factor influencing effectiveness. We implement and evaluate several LLM-based generation strategies, and discover that AI-generated and human-written counterspeech can be easily distinguished by both simple classifiers and humans. Further, we reveal differences in linguistic characteristics, politeness, and specificity. The dataset used in this study is publicly available for further research. Sujana Mamidisetty, Eduardo Blanco 0002, Lingzi Hong |
COLING | 4 |
| 2024 | Outcome-Constrained Large Language Models for Countering Hate SpeechabstractAutomatic counterspeech generation methods have been developed to assist efforts in combating hate speech.Existing research focuses on generating counterspeech with linguistic attributes such as being polite, informative, and intent-driven.However, the real impact of counterspeech in online environments is seldom considered.This study aims to develop methods for generating counterspeech constrained by conversation outcomes and evaluate their effectiveness.We experiment with large language models (LLMs) to incorporate into the text generation process two desired conversation outcomes: low conversation incivility and nonhateful hater reentry.Specifically, we experiment with instruction prompts, LLM finetuning, and LLM reinforcement learning (RL).Evaluation results show that our methods effectively steer the generation of counterspeech towards the desired outcomes.Our analyses, however, show that there are differences in the quality and style depending on the model. Lingzi Hong, Pengcheng Luo, Eduardo Blanco 0002 |
EMNLP | 1 |
| 2024 | Hate Cannot Drive Out Hate: Forecasting Conversation Incivility following Replies to Hate SpeechabstractUser-generated counter hate speech is a promising means to combat hate speech, but questions about whether it can stop incivility in follow-up conversations linger. We argue that effective counter hate speech stops incivility from emerging in follow-up conversations—counter hate that elicits more incivility is counterproductive. This study introduces the task of predicting the incivility of conversations following replies to hate speech. We first propose a metric to measure conversation incivility based on the number of civil and uncivil comments as well as the unique authors involved in the discourse. Our metric approximates human judgments more accurately than previous metrics. We then use the metric to evaluate the outcomes of replies to hate speech. A linguistic analysis uncovers the differences in the language of replies that elicit follow-up conversations with high and low incivility. Experimental results show that forecasting incivility is challenging. We close with a qualitative analysis shedding light into the most common errors made by the best model. Xinchen Yu, Eduardo Blanco 0002, Lingzi Hong |
ICWSM | 3 |
| 2024 | An Emerging Adults' Patient Portal Behavioral Model: Integrating Perceived Risk Theory, Technology Acceptance Model, and Personal InnovativenessabstractHealth information technology provides patients the ability to manage their healthcare through patient portals. Such portals increase patient involvement, self-management, and satisfaction. Despite their benefits, patient portal adoption and usage remain low, especially among emerging adults who are newly self-managed. To investigate the behavioral intentions of emerging adults toward adopting and using patient portals, this study builds upon the Technology Acceptance Model, Perceived Risk Theory, and Personal Innovativeness. A survey was administered to emerging adults aged 18-29, and structural equation modeling was used to assess the posited model’s fit. Results show the importance of developing practical insights and strategies to overcome resistance behavior. Additionally, the research found that personal innovativeness plays a significant role in adoption and usage intention. These findings extend the literature by highlighting the specific needs of emerging adults regarding patient portal adoption and utilization. The study underscores the importance of providing guidance, training, and awareness programs. Navya Velverthi, Victor R. Prybutok, Lingzi Hong |
Int. J. Hum. Comput. Interact. | 3 |
| 2024 | Toward measuring data literacy for higher education: Developing and validating a data literacy self-efficacy scaleabstractAbstract Data literacy, a multifaceted competency in working with data, has emerged as an essential skill that holds significance in both personal and professional lives. Nonetheless, there is a lack of a precise definition of data literacy, and individuals' perceptions of their data literacy have not been thoroughly investigated. This study aims to develop and validate a scale designed for measuring self‐efficacy in data literacy within the context of higher education. Both exploratory and confirmatory factor analyses were conducted to determine construct validity and reliability. The resulting data literacy self‐efficacy scale comprises 31 items organized into three factors: data identification, data processing, and data management and sharing. These factors represent distinct yet interconnected dimensions, highlighting the multifaceted nature of data literacy. Jeonghyun Kim 0001, Lingzi Hong, Sarah A. Evans |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2023 | A Fine-Grained Taxonomy of Replies to Hate SpeechabstractCountering rather than censoring hate speech has emerged as a promising strategy to address hatred.There are many types of counterspeech in user-generated content: addressing the hateful content or its author, generic requests, well-reasoned counter arguments, insults, etc.The effectiveness of counterspeech, which we define as subsequent incivility, depends on these types.In this paper, we present a theoretically grounded taxonomy of replies to hate speech and a new corpus.We work with real, user-generated hate speech and all the replies it elicits rather than replies generated by a third party.Our analyses provide insights into the content real users reply with as well as which replies are empirically most effective.We also experiment with models to characterize the replies to hate speech, thereby opening the door to estimating whether a reply to hate speech will result in further incivility. Error Type % ExampleGround Truth Predicted Rhetorical 26 Hate: F**k worthless inbreds who've contributed nothing to society. Xinchen Yu, Ashley Zhao, Eduardo Blanco 0002, Lingzi Hong |
EMNLP | 4 |
| 2022 | Causal Impact Model to Evaluate the Diffusion Effect of Social Media Campaigns
Xinchen Yu, Afra J. Mashhadi, Jeremy Boy, René Clausen Nielsen, Lingzi Hong |
ECSCW | 5 |
| 2022 | Hate Speech and Counter Speech Detection: Conversational Context Does MatterabstractHate speech is plaguing the cyberspace along with user-generated content.This paper investigates the role of conversational context in the annotation and detection of online hate and counter speech, where context is defined as the preceding comment in a conversation thread.We created a context-aware dataset for a 3-way classification task on Reddit comments: hate speech, counter speech, or neutral.Our analyses indicate that context is critical to identify hate and counter speech: human judgments change for most comments depending on whether we show annotators the context.A linguistic analysis draws insights into the language people use to express hate and counter speech.Experimental results show that neural networks obtain significantly better results if context is taken into account.We also present qualitative error analyses shedding light into (a) when and why context is beneficial and (b) the remaining errors made by our best model when context is taken into account. Xinchen Yu, Eduardo Blanco 0002, Lingzi Hong |
NAACL-HLT | 3 |
| 2021 | Multi-faceted Classification for the Identification of Informative Communications during Crises: Case of COVID-19abstractSocial media data are used to enhance crisis management, as people widely adopt social media to share and acquire information to cope with uncertainties in crises. Identification and extraction of informative communications out of large volumes of data is critical for accurate situational awareness and timely response. Existing studies use conditions of geolocations, keywords, and topics separately or jointly to retrieve data that can be crisis related, but are not enough to filter subsets of data for different crisis management tasks. We propose that the crisis communication purposes of users can be detected to enhance data selection and prioritization for different crisis management tasks. A classification framework was built to identify three facets of a message: content type, audience type, and information source. The definitions of these categories are not dependent on a specific type of crises. So the classification framework can be potentially applied to different crisis scenarios. Machine learning models were created for the automatic classification of messages. Results showed the CNN-based model achieved the best accuracy (88.5%) for the classification of content type. The proposed Naive Bayes and logistic repression with predetermined features can best differentiate audience types and information source with an accuracy of 72.7% and 72.2%, respectively. Zhuoli Xie, Ajay Jayanth, Kapil Yadav, Guanghui Ye, Lingzi Hong |
COMPSAC | 5 |
| 2019 | Predicting Perceived Level of Cycling Safety for Cycling TripsabstractCycling provides various benefits to cyclists and cities. Nevertheless, the growth of cycling is still hindered by the lack of citywide information about perceived cycling safety. Providing cyclists with information about the safest routes could help increase cycling activity. In this paper, we aim to predict the perceived level of cycling safety for a trip (trip-PLOCS). We utilize LSTM-based architectures to incorporate the sequential information of segments in a trip, and predict its cycling safety. Our proposed method can achieve up to 76% F1 micro (65% F1 macro) score, 10% (19%) better than the state-of-the-art baseline. Finally, we use SHAP to extract insights about trip-PLOCS, showing that social features contribute to perceived danger while cycling facilities contributes to the perceived safety. Lingzi Hong, Vanessa Frías-Martínez |
SIGSPATIAL/GIS | 2 |
| 2019 | Characterization of internal migrant behavior in the immediate post-migration period using cell phone tracesabstractInternal migrations have been studied using two types of approaches: macro-level and micro-level analyses. Macro-level studies are typically carried out using a combination of various survey and census datasets to model large-scale behaviors, however these models fail to provide more nuanced information about the physical or social status of the migrants. Micro approaches, which successfully use interviews and diaries to provide a window into more individual behaviors, could benefit from methods to identify novel or under-studied behaviors that should be addressed in the migration research agenda. In this paper, we present a framework that uses information extracted from cell phone metadata to reveal internal migration behaviors that could guide or complement the research agenda of micro-level migration researchers working to understand the physical, social and psychological decision processes behind migration experiences. The proposed framework allows to carry out micro-level analyses of internal migration with a focus on immediate post-migration behaviors and the role of pre-migration activities from two perspectives: spatial behaviors and social ties. Ultimately, we expect our analyses to inform migration researchers of pre- and post-migration behaviors that would benefit from further qualitative analysis. Lingzi Hong, Enrique Frías-Martínez, Andrés Villarreal, Vanessa Frías-Martínez |
ICTD | 1 |
| 2018 | Predicting Perceived Cycling Safety Levels Using Open and Crowdsourced DataabstractCycling communities have been related to lower obesity rates and lower stress levels. Nevertheless, one of the main obstacles to increase ridership in cities is the lack of information regarding perceived cycling safety at the street level. City planners have typically used extensive road network and traffic information to approximate cycling safety levels. However, this approach requires the deployment of expensive sensors thus making it hard for many cities to get access to accurate cycling safety maps. In this paper, we evaluate several methods to predict urban cycling safety at the street level, exclusively using public information from open and crowdsourced datasets. We evaluate the proposed approach in the city of Washington D.C. and achieve F1 scores of 66%, 70% and 88% when five, four or three different cycling safety levels are considered. Lingzi Hong, Vanessa Frías-Martínez |
IEEE BigData | 2 |
| 2016 | Topic Models to Infer Socio-Economic MapsabstractSocio-economic maps contain important information regarding the population of a country. Computing these maps is critical given that policy makers often times make important decisions based upon such information. However, the compilation of socio-economic maps requires extensive resources and becomes highly expensive. On the other hand, the ubiquitous presence of cell phones, is generating large amounts of spatiotemporal data that can reveal human behavioral traits related to specific socio-economic characteristics. Traditional inference approaches have taken advantage of these datasets to infer regional socio-economic characteristics. In this paper, we propose a novel approach whereby topic models are used to infer socio-economic levels from large-scale spatio-temporal data. Instead of using a pre-determined set of features, we use latent Dirichlet Allocation (LDA) to extract latent recurring patterns of co-occurring behaviors across regions, which are then used in the prediction of socio-economic levels. We show that our approach improves state of the art prediction results by 9%. Lingzi Hong, Enrique Frías-Martínez, Vanessa Frías-Martínez |
AAAI | 1 |
| 2013 | Movie Recommendation Using Unrated DataabstractModel based movie recommender systems have been thoroughly investigated in the past few years, and they rely on rating data. In this paper, we take into account unrateddata of genre information to improve the performance of movie recommendation. We propose a novel method to measure users' preference on movie genres, and use Pearson Correlation Coefficient(PCC) to compute the user similarity. A matrix factorization framework is introduced for genre preference regularization. Experimental results on Movie Lens data set demonstrate that the approach performs well. Our method can also be used to increase the genre diversity of recommendations to some extent. Dong Nie, Lingzi Hong, Tingshao Zhu |
ICMLA (1) | 2 |