VLDB 2026 Research / reviewers in the wild / expert
Yuheng Hu
dblp:61/809
· DBLP profile ↗
17ranked-venue papers
11as first author
3since 2021 · last 2025
0000-0003-3665-1238ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 12 · 6 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 8 · 6 first-authorApplied, interdisciplinary, general and emerging computing · 8 · 5 first-authorArtificial intelligence and machine learning · 4 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorTheory of computation · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Engagement or entanglement? The dual impact of generative artificial intelligence in online knowledge exchange platforms
Aida Sanatizadeh, Yingda Lu, Keran Zhao, Yuheng Hu |
Inf. Manag. | 4 |
| 2022 | SHEDR: An End-to-End Deep Neural Event Detection and Recommendation Framework for Hyperlocal News Using Social MediaabstractResidents often rely on newspapers and television to gather hyperlocal news for community awareness and engagement. More recently, social media have emerged as an increasingly important source of hyperlocal news. Thus far, the literature on using social media to create desirable societal benefits, such as civic awareness and engagement, is still in its infancy. One key challenge in this research stream is to timely and accurately distill information from noisy social media data streams to community members. In this work, we develop SHEDR (social media–based hyperlocal event detection and recommendation), an end-to-end neural event detection and recommendation framework with a particular use case for Twitter to facilitate residents’ information seeking of hyperlocal events. The key model innovation in SHEDR lies in the design of the hyperlocal event detector and the event recommender. First, we harness the power of two popular deep neural network models, the convolutional neural network (CNN) and long short-term memory (LSTM), in a novel joint CNN-LSTM model to characterize spatiotemporal dependencies for capturing unusualness in a region of interest, which is classified as a hyperlocal event. Next, we develop a neural pairwise ranking algorithm for recommending detected hyperlocal events to residents based on their interests. To alleviate the sparsity issue and improve personalization, our algorithm incorporates several types of contextual information covering topic, social, and geographical proximities. We perform comprehensive evaluations based on two large-scale data sets comprising geotagged tweets covering Seattle and Chicago. We demonstrate the effectiveness of our framework in comparison with several state-of-the-art approaches. We show that our hyperlocal event detection and recommendation models consistently and significantly outperform other approaches in terms of precision, recall, and F-1 scores. Summary of Contribution: In this paper, we focus on a novel and important, yet largely underexplored application of computing—how to improve civic engagement in local neighborhoods via local news sharing and consumption based on social media feeds. To address this question, we propose two new computational and data-driven methods: (1) a deep learning–based hyperlocal event detection algorithm that scans spatially and temporally to detect hyperlocal events from geotagged Twitter feeds; and (2) A personalized deep learning–based hyperlocal event recommender system that systematically integrates several contextual cues such as topical, geographical, and social proximity to recommend the detected hyperlocal events to potential users. We conduct a series of experiments to examine our proposed models. The outcomes demonstrate that our algorithms are significantly better than the state-of-the-art models and can provide users with more relevant information about the local neighborhoods that they live in, which in turn may boost their community engagement. Yuheng Hu, Yili Hong 0002 |
INFORMS J. Comput. | 1 |
| 2021 | Characterizing Social TV Activity Around Televised Events: A Joint Topic Model ApproachabstractViewers often use social media platforms like Twitter to express their views about televised programs and events like the presidential debate, the Oscars, and the State of the Union speech. Although this promises tremendous opportunities to analyze the feedback on a program or an event using viewer-generated content on social media, there are significant technical challenges to doing so. Specifically, given a televised event and related tweets about this event, we need methods to effectively align these tweets and the corresponding event. In turn, this will raise many questions, such as how to segment the event and how to classify a tweet based on whether it is generally about the entire event or specifically about one particular event segment. In this paper, we propose and develop a novel joint Bayesian model that aligns an event and its related tweets based on the influence of the event’s topics. Our model allows the automated event segmentation and tweet classification concurrently. We present an efficient inference method for this model and a comprehensive evaluation of its effectiveness compared with the state-of-the-art methods. We find that the topics, segments, and alignment provided by our model are significantly more accurate and robust. Yuheng Hu |
INFORMS J. Comput. | 1 |
| 2017 | Mood Congruence or Mood Consistency? Examining Aggregated Twitter Sentiment Towards Ads in 2016 Super Bowl
Yuheng Hu, Tingting Nian |
ICWSM | 1 |
| 2016 | Predicting Perceived Brand Personality with Social Media
Anbang Xu, Liang Gou, Rama Akkiraju, Jalal Mahmud, Vibha Sinha, Yuheng Hu |
ICWSM | 7 |
| 2016 | Collective Sensemaking via Social Sensors: Extracting, Profiling, Analyzing, and Predicting Real-world EventsabstractSocial media platforms like Twitter and Facebook have emerged as some of the most important platforms for people to discover, report, share, and communicate with others about various public events, be they of global or local interest (some high profile examples include the U.S Presidential debates, the Boston bombings, the hurricane Sandy, etc). The burst of social media reaction can be seen as a valuable real-time reflection of events as they happen, and can be used for a variety of applications such as computational journalism. Until now, such analysis has been mostly done manually or through primitive tools. Scalable and automated approaches are needed given the massive amounts of both event and reaction information. These approaches must also be able to conduct in-depth analysis of complex interactions between an event and its audience. Supporting such automation and examination however poses several computational challenges. In recent years, research communities have witnessed a growing interest in tackling these challenges. Furthermore, much recent research has begun to focus on solving more complex event analytics tasks such as post-event effect quantification and event progress prediction. This tutorial aims to review and examine current state of the research progress on this emerging topic. Yuheng Hu, Yu-Ru Lin, Jiebo Luo 0001 |
KDD | 1 |
| 2015 | Predicting User Engagement on Twitter with Real-World Events
Yuheng Hu, Shelly Farnham, Kartik Talamadupula |
ICWSM | 1 |
| 2015 | Inferring Sentiment from Web Images with Joint Inference on Visual and Social Cues: A Regulated Matrix Factorization Approach
Yilin Wang 0002, Yuheng Hu, Subbarao Kambhampati, Baoxin Li |
ICWSM | 2 |
| 2014 | BayesWipe: A multimodal system for data cleaning and consistent query answering on structured bigdataabstractRecent efforts in data cleaning of structured data have focused exclusively on problems like data deduplication, record matching, and data standardization; none of these focus on fixing incorrect attribute values in tuples. Correcting values in tuples is typically performed by a minimum cost repair of tuples that violate static constraints like CFDs (which have to be provided by domain experts, or learned from a clean sample of the database). In this paper, we provide a method for correcting individual attribute values in a structured database using a Bayesian generative model and a statistical error model learned from the noisy database directly. We thus avoid the necessity for a domain expert or clean master data. We also show how to efficiently perform consistent query answering using this model over a dirty database, in case write permissions to the database are unavailable. We evaluate our methods over both synthetic and real data. Sushovan De, Yuheng Hu, Yi Chen 0001, Subbarao Kambhampati |
IEEE BigData | 2 |
| 2014 | What We Instagram: A First Analysis of Instagram Photo Content and User Types
Yuheng Hu, Lydia Manikonda, Subbarao Kambhampati |
ICWSM | 1 |
| 2013 | Whoo.ly: facilitating information seeking for hyperlocal communities using social mediaabstractSocial media systems promise powerful opportunities for people to connect to timely, relevant information at the hyper local level. Yet, finding the meaningful signal in noisy social media streams can be quite daunting to users. In this paper, we present and evaluate Whoo.ly, a web service that provides neighborhood-specific information based on Twitter posts that were automatically inferred to be hyperlocal. Whoo.ly automatically extracts and summarizes hyperlocal information about events, topics, people, and places from these Twitter posts. We provide an overview of our design goals with Whoo.ly and describe the system including the user interface and our unique event detection and summarization algorithms. We tested the usefulness of the system as a tool for finding neighborhood information through a comprehensive user study. The outcome demonstrated that most participants found Whoo.ly easier to use than Twitter and they would prefer it as a tool for exploring their neighborhoods. Yuheng Hu, Shelly Farnham, Andrés Monroy-Hernández |
CHI | 1 |
| 2013 | Dude, srsly?: The Surprisingly Formal Nature of Twitter's Language
Yuheng Hu, Kartik Talamadupula, Subbarao Kambhampati |
ICWSM | 1 |
| 2013 | Listening to the Crowd: Automated Analysis of Events via Aggregated Twitter Sentiment
Yuheng Hu, Subbarao Kambhampati |
IJCAI | 1 |
| 2012 | ET-LDA: Joint Topic Modeling for Aligning Events and their Twitter FeedbackabstractDuring broadcast events such as the Superbowl, the U.S. Presidential and Primary debates, etc., Twitter has become the de facto platform for crowds to share perspectives and commentaries about them. Given an event and an associated large-scale collection of tweets, there are two fundamental research problems that have been receiving increasing attention in recent years. One is to extract the topics covered by the event and the tweets; the other is to segment the event. So far these problems have been viewed separately and studied in isolation. In this work, we argue that these problems are in fact inter-dependent and should be addressed together. We develop a joint Bayesian model that performs topic modeling and event segmentation in one unified framework. We evaluate the proposed model both quantitatively and qualitatively on two large-scale tweet datasets associated with two events from different domains to show that it improves significantly over baseline models. Yuheng Hu, Ajita John, Subbarao Kambhampati |
AAAI | 1 |
| 2012 | What Were the Tweets About? Topical Associations between Public Events and Twitter Feeds
Yuheng Hu, Ajita John, Dorée D. Seligmann |
ICWSM | 1 |
| 2011 | Relevance-Based Retrieval on Hidden-Web Text Databases without Ranking SupportabstractMany online or local data sources provide powerful querying mechanisms but limited ranking capabilities. For instance, PubMed allows users to submit highly expressive Boolean keyword queries, but ranks the query results by date only. However, a user would typically prefer a ranking by relevance, measured by an information retrieval (IR) ranking function. A naive approach would be to submit a disjunctive query with all query keywords, retrieve all the returned matching documents, and then rerank them. Unfortunately, such an operation would be very expensive due to the large number of results returned by disjunctive queries. In this paper, we present algorithms that return the top results for a query, ranked according to an IR-style ranking function, while operating on top of a source with a Boolean query interface with no ranking capabilities (or a ranking capability of no interest to the end user). The algorithms generate a series of conjunctive queries that return only documents that are candidates for being highly ranked according to a relevance metric. Our approach can also be applied to other settings where the ranking is monotonic on a set of factors (query keywords in IR) and the source query interface is a Boolean expression of these factors. Our comprehensive experimental evaluation on the PubMed database and a TREC data set show that we achieve order of magnitude improvement compared to the current baseline approaches. Vagelis Hristidis, Yuheng Hu, Panagiotis G. Ipeirotis |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2010 | Ranked queries over sources with Boolean query interfaces without ranking supportabstractMany online or local data sources provide powerful querying mechanisms but limited ranking capabilities. For instance, PubMed allows users to submit highly expressive Boolean keyword queries, but ranks the query results by date only. However, a user would typically prefer a ranking by relevance, measured by an Information Retrieval (IR) ranking function. The naive approach would be to submit a disjunctive query with all query keywords, retrieve the returned documents, and then re-rank them. Unfortunately, such an operation would be very expensive due to the large number of results returned by disjunctive queries. In this paper we present algorithms that return the top results for a query, ranked according to an IR-style ranking function, while operating on top of a source with a Boolean query interface with no ranking capabilities (or a ranking capability of no interest to the end user). The algorithms generate a series of conjunctive queries that return only documents that are candidates for being highly ranked according to a relevance metric. Our approach can also be applied to other settings where the ranking is monotonic on a set of factors (query keywords in IR) and the source query interface is a Boolean expression of these factors. Our comprehensive experimental evaluation on the PubMed database and TREC dataset show that we achieve order of magnitude improvement compared to the current baseline approaches. Vagelis Hristidis, Yuheng Hu, Panagiotis G. Ipeirotis |
ICDE | 2 |