VLDB 2026 Research / reviewers in the wild / expert
Hemank Lamba
dblp:120/8503
· DBLP profile ↗
27ranked-venue papers
10as first author
7since 2021 · last 2025
0000-0002-9794-3587ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 16 · 7 first-author · 2 since 2021Artificial intelligence and machine learning · 15 · 4 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 8 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CEHA: A Dataset of Conflict Events in the Horn of AfricaabstractNatural Language Processing (NLP) of news articles can play an important role in understanding the dynamics and causes of violent conflict. Despite the availability of datasets categorizing various conflict events, the existing labels often do not cover all of the fine-grained violent conflict event types relevant to areas like the Horn of Africa. In this paper, we introduce a new benchmark dataset Conflict Events in the Horn of Africa region (CEHA) and propose a new task for identifying violent conflict events using online resources with this dataset. The dataset consists of 500 English event descriptions regarding conflict events in the Horn of Africa region with fine-grained event-type definitions that emphasize the cause of the conflict. This dataset categorizes the key types of conflict risk according to specific areas required by stakeholders in the Humanitarian-Peace-Development Nexus. Additionally, we conduct extensive experiments on two tasks supported by this dataset: Event-relevance Classification and Event-type Classification. Our baseline models demonstrate the challenging nature of these tasks and the usefulness of our dataset for model evaluations in low-resource settings. Di Lu 0003, Shihao Ran, Elizabeth M. Olson, Hemank Lamba, Aoife Cahill, Joel R. Tetreault, Alejandro Jaimes |
COLING | 5 |
| 2025 | Uchaguzi-2022: A Dataset of Citizen Reports on the 2022 Kenyan ElectionabstractOnline reporting platforms have enabled citizens around the world to collectively share their opinions and report in real time on events impacting their local communities. Systematically organizing (e.g., categorizing by attributes) and geotagging large amounts of crowdsourced information is crucial to ensuring that accurate and meaningful insights can be drawn from this data and used by policy makers to bring about positive change. These tasks, however, typically require extensive manual annotation efforts. In this paper we present Uchaguzi-2022, a dataset of 14k categorized and geotagged citizen reports related to the 2022 Kenyan General Election containing mentions of election-related issues such as official misconduct, vote count irregularities, and acts of violence. We use this dataset to investigate whether language models can assist in scalably categorizing and geotagging reports, thus highlighting its potential application in the AI for Social Good space. Roberto Mondini, Neema Kotonya, Robert L. Logan IV, Elizabeth M. Olson, Angela Oduor Lungati, Daniel Duke Odongo, Tim Ombasa, Hemank Lamba, Aoife Cahill, Joel R. Tetreault, Alejandro Jaimes |
COLING | 8 |
| 2025 | Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label DefinitionsabstractSeyedali Mohammadi, Bhaskara Hanuma Vedula, Hemank Lamba, Edward Raff, Ponnurangam Kumaraguru, Francis Ferraro, Manas Gaur. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Seyedali Mohammadi, Bhaskara Hanuma Vedula, Hemank Lamba, Edward Raff, Ponnurangam Kumaraguru, Francis Ferraro, Manas Gaur |
EMNLP | 3 |
| 2024 | Dissecting users' needs for search result explanationsabstractThere is a growing demand for transparency in search engines to understand how search results are curated and to enhance users’ trust. Prior research has introduced search result explanations with a focus on how to explain, assuming explanations are beneficial. Our study takes a step back to examine if search explanations are needed and when they are likely to provide benefits. Additionally, we summarize key characteristics of helpful explanations and share users’ perspectives on explanation features provided by Google and Bing. Interviews with non-technical individuals reveal that users do not always seek or understand search explanations and mostly desire them for complex and critical tasks. They find Google’s search explanations too obvious but appreciate the ability to contest search results. Based on our findings, we offer design recommendations for search engines and explanations to help users better evaluate search results and enhance their search experience. Prerna Juneja, Alison Smith-Renner, Hemank Lamba, Joel R. Tetreault, Alex Jaimes |
CHI | 4 |
| 2022 | "This Is Damn Slick!" Estimating the Impact of Tweets on Open Source Project Popularity and New ContributorsabstractTwitter is widely used by software developers. But how effective are tweets at promoting open source projects? How could one use Twitter to increase a project's popularity or attract new contributors? In this paper we report on a mixed-methods empirical study of 44,544 tweets containing links to 2,370 open-source GitHub repositories, looking for evidence of causal effects of these tweets on the projects attracting new GitHub stars and contributors, as well as characterizing the high-impact tweets, the people likely being attracted by them, and how they differ from contributors attracted otherwise. Among others, we find that tweets have a statistically significant and practically sizable effect on obtaining new stars and a small average effect on attracting new contributors. The popularity, content of the tweet, as well as the identity of tweet authors all affect the scale of the attraction effect. In addition, our qualitative analysis suggests that forming an active Twitter community for an open source project plays an important role in attracting new committers via tweets. We also report that developers who are new to GitHub or have a long history of Twitter usage but few tweets posted are most likely to be attracted as contributors to the repositories mentioned by tweets. Our work contributes to the literature on open source sustainability. Hongbo Fang, Hemank Lamba, James D. Herbsleb, Bogdan Vasilescu |
ICSE | 2 |
| 2022 | Effect of Popularity Shocks on User Behaviour
Omkar Gurjar, Tanmay Bansal, Hitkul Jangra, Hemank Lamba, Ponnurangam Kumaraguru |
ICWSM | 4 |
| 2022 | Glowing Experience or Bad Trip? A Quantitative Analysis of User Reported Drug Experiences on Erowid.org
Angelina Voggenreiter, Momin M. Malik, Hemank Lamba, Earth Erowid, Sylvia Thyssen, Jürgen Pfeffer |
ICWSM | 3 |
| 2020 | TellTail: Fast Scoring and Detection of Dense SubgraphsabstractSuppose you visit an e-commerce site, and see that 50 users each reviewed almost all of the same 500 products several times each: would you get suspicious? Similarly, given a Twitter follow graph, how can we design principled measures for identifying surprisingly dense subgraphs? Dense subgraphs often indicate interesting structure, such as network attacks in network traffic graphs. However, most existing dense subgraph measures either do not model normal variation, or model it using an Erdős-Renyi assumption - but this assumption has been discredited decades ago. What is the right assumption then? We propose a novel application of extreme value theory to the dense subgraph problem, which allows us to propose measures and algorithms which evaluate the surprisingness of a subgraph probabilistically, without requiring restrictive assumptions (e.g. Erdős-Renyi). We then improve the practicality of our approach by incorporating empirical observations about dense subgraph patterns in real graphs, and by proposing a fast pruning-based search algorithm. Our approach (a) provides theoretical guarantees of consistency, (b) scales quasi-linearly, and (c) outperforms baselines in synthetic and ground truth settings. Bryan Hooi, Kijung Shin, Hemank Lamba, Christos Faloutsos |
AAAI | 3 |
| 2020 | Provably Robust Node Classification via Low-Pass Message PassingabstractGraph Convolutional Networks (GCNs) have achieved state-of-the-art performance on node classification. However, recent works have shown that GCNs are vulnerable to adversarial attacks, such as additions or deletions of adversarially-chosen edges in the graph, in order to mislead the node classification algorithms. How can we design robust GCNs that are resistant to such adversarial attacks? More challengingly, how can we do this in a way that is provably robust? We propose a robust node classification approach based on a low-pass `message passing' mechanism, that (a) reduces the effectiveness of adversarial attacks in experiments, and (b) provides theoretical guarantees against adversarial attacks. Our approach can be embedded into the existing GCN architectures to enhance their robustness. Empirical results show that our loss-pass method effectively improves the performance of multiple GCNs under miscellaneous perturbations and helps them to achieve superior performance on various graphs. Yiwei Wang 0001, Shenghua Liu, Minji Yoon, Hemank Lamba, Wei Wang 0059, Christos Faloutsos, Bryan Hooi |
ICDM | 4 |
| 2020 | Driving the Last Mile: Characterizing and Understanding Distracted Driving Posts on Social Networks
Hemank Lamba, Shashank Srikanth, Dheeraj Reddy Pailla, Shwetanshu Singh, Karandeep Juneja, Ponnurangam Kumaraguru |
ICWSM | 1 |
| 2020 | Need for Tweet: How Open Source Developers Talk About Their GitHub Work on TwitterabstractSocial media, especially Twitter, has always been a part of the professional lives of software developers, with prior work reporting on a diversity of usage scenarios, including sharing information, staying current, and promoting one's work. However, previous studies of Twitter use by software developers typically lack information about activities of the study subjects (and their outcomes) on other platforms. To enable such future research, in this paper we propose a computational approach to cross-link users across Twitter and GitHub, revealing (at least) 70,427 users active on both. As a preliminary analysis of this dataset, we report on a case study of 786 tweets by open-source developers about GitHub work, combining automatic characterization of tweet authors in terms of their relationship to the GitHub items linked in their tweets with qualitative analysis of the tweet contents. We find that different developer roles tend to have different tweeting behaviors, with repository owners being perhaps the most distinctive group compared to other project contributors and followers. We also note a sizeable group of people who follow others on GitHub and tweet about these people's work, but do not otherwise contribute to those open-source projects. Our results and public dataset open up multiple future research directions. Hongbo Fang, Daniel Klug, Hemank Lamba, James D. Herbsleb, Bogdan Vasilescu |
MSR | 3 |
| 2020 | Heard it through the Gitvine: an empirical study of tool diffusion across the npm ecosystemabstractAutomation tools like continuous integration services, code coverage reporters, style checkers, dependency managers, etc. are all known to provide significant improvements in developer productivity and software quality. Some of these tools are widespread, others are not. How do these automation "best practices" spread? And how might we facilitate the diffusion process for those that have seen slower adoption? In this paper, we rely on a recent innovation in transparency on code hosting platforms like GitHub---the use of repository badges---to track how automation tools spread in open-source ecosystems through different social and technical mechanisms over time. Using a large longitudinal data set, multivariate network science techniques, and survival analysis, we study which socio-technical factors can best explain the observed diffusion process of a number of popular automation tools. Our results show that factors such as social exposure, competition, and observability affect the adoption of tools significantly, and they provide a roadmap for software engineers and researchers seeking to propagate best practices and tools. Hemank Lamba, Asher Trockman, Daniel Armanios, Christian Kästner, Heather Miller, Bogdan Vasilescu |
ESEC/SIGSOFT FSE | 1 |
| 2019 | Characterizing and detecting livestreaming chatbotsabstractLivestreaming platforms enable content producers, or streamers, to broadcast creative content to a potentially large viewer base. Chatrooms form an integral part of such platforms, enabling viewers to interact both with the streamer, and amongst themselves. Streams with high engagement (many viewers and active chatters) are typically considered engaging, and often promoted to end users by means of recommendation algorithms, and exposed to better monetization opportunities via revenue share from platform advertising, viewer donations, and third-party sponsorships. Given such incentives, some streamers make use of fraudulent means to increase perceived engagement by simulating chatter via fake "chatbots" which can be purchased from shady online marketplaces. This inauthentic engagement can negatively influence recommendation, hurt streamer and viewer trust in the platform, and harm monetization for honest streamers. In this paper, we tackle the novel problem of automating detection of chatbots on livestreaming platforms. To this end, we first formalize the livestreaming chatbot detection problem and characterize differences between botted and genuine chatter behavior observed from a real-world livestreaming chatter dataset collected from Twitch.tv. We then propose SHERLOCK, which posits a two-stage approach of detecting chatbotted streams, and subsequently detecting the constituent chatbots. Finally, we demonstrate effectiveness on both real and synthetic data: to this end, we propose a novel strategy for collecting labeled, synthetic chatter dataset (typically unavailable) from such platforms, enabling evaluation of proposed detection approaches against chatbot behaviors with varying signatures. Our approach achieves .97 precision/recall on the real-world dataset, and .80+ F1 scores across most simulated attack settings. Shreya Jain, Dipankar Niranjan, Hemank Lamba, Neil Shah, Ponnurangam Kumaraguru |
ASONAM | 3 |
| 2019 | Modeling Dwell Time Engagement on Visual MultimediaabstractVisual multimedia is one of the most prevalent sources of modern online content and engagement. However, despite its prevalence, little is known about user engagement with such content. For instance, how can we model engagement for a specific content or viewer sample, and across multiple samples? Can we model and discover patterns in these interactions, and detect outlying behaviors corresponding to abnormal engagement? In this paper, we study these questions in depth. Understanding these questions has implications in user modeling and understanding, ranking, trust and safety and more. For analysis, we consider content and viewer dwell time (engagement duration) behaviors with images and videos on Snapchat Stories, one of the largest multimedia-driven social sharing services. To our knowledge, we are the first to model and analyze dwell time behaviors on such media. Specifically, our contributions include (a) individual modeling: we propose and evaluate the ŁFmodel, ŁTmodel and \Vmodel parametric models to describe dwell times of unlooped/looped media and viewers which outperform alternatives, (b) aggregate modeling: we show how to flexibly summarize the respective joint distributions of multivariate parametrized fits across many samples using Vine Copulas in the analog \ALFmodel, \ALTmodel and \AVmodel models, which enable inferences regarding aggregate behavioral patterns, and offer the ability to simulate real-looking engagement data (c) anomaly detection: we demonstrate our aggregate models can robustly detect anomalies present during training ($0.9+$ AUROC across most attack models), and also enable discovery of real dwell time anomalies. Hemank Lamba, Neil Shah |
KDD | 1 |
| 2019 | Learning On-the-Job to Re-rank Anomalies from Top-1 FeedbackabstractIn many anomaly mining scenarios, a human expert verifies the anomaly at-the-top (as ranked by an anomaly detector) before they move on to the next. This verification produces a label—true positive (TP) or false positive (FP). In this work, we show how to leverage this label feedback for the top-1 instance to quickly re-rank the anomalies in an online fashion. In contrast to a detector that ranks once and goes offline, we propose a detector called OJRank that works alongside the human and continues to learn (how to rank) on-the-job, i.e., from every feedback. The benefits OJRank provides are two-fold; it reduces (i) the false positive rate by ‘muting’ the anomalies similar to FP instances; as well as (ii) the expert effort by elevating to the top the anomalies similar to a TP instance. We show that OJRank achieves statistically significant improvement on both detection precision and human effort over the offline detector as well as existing state-of-the-art ranking strategies, while keeping the per feedback response time (to re-rank) well below a second. Hemank Lamba, Leman Akoglu |
SDM | 1 |
| 2019 | WebHound: a data-driven intrusion detection from real-world web access logs
Te-En Wei, Hahn-Ming Lee, Albert B. Jeng, Hemank Lamba, Christos Faloutsos |
Soft Comput. | 4 |
| 2018 | xStream: Outlier Detection in Feature-Evolving Data StreamsabstractThis work addresses the outlier detection problem for feature-evolving streams, which has not been studied before. In this setting both (1) data points may evolve, with feature values changing, as well as (2) feature space may evolve, with newly-emerging features over time. This is notably different from row-streams, where points with fixed features arrive one at a time. We propose a density-based ensemble outlier detector, called xStream, for this more extreme streaming setting which has the following key properties: (1) it is a constant-space and constant-time (per incoming update) algorithm, (2) it measures outlierness at multiple scales or granularities, it can handle (3 i ) high-dimensionality through distance-preserving projections, and (3$ii$) non-stationarity via $O(1)$-time model updates as the stream progresses. In addition, xStream can address the outlier detection problem for the (less general) disk-resident static as well as row-streaming settings. We evaluate xStream rigorously on numerous real-life datasets in all three settings: static, row-stream, and feature-evolving stream. Experiments under static and row-streaming scenarios show that xStream is as competitive as state-of-the-art detectors and particularly effective in high-dimensions with noise. We also demonstrate that our solution is fast and accurate with modest space overhead for evolving streams, on which there exists no competition. Emaad Manzoor, Hemank Lamba, Leman Akoglu |
KDD | 2 |
| 2017 | The Many Faces of Link FraudabstractMost past work on social network link fraud detection tries to separate genuine users from fraudsters, implicitly assuming that there is only one type of fraudulent behavior. But is this assumption true? And, in either case, what are the characteristics of such fraudulent behaviors? In this work, we set up honeypots ("dummy" social network accounts), and buy fake followers (after careful IRB approval). We report the signs of such behaviors including oddities in local network connectivity, account attributes, and similarities and differences across fraud providers. Most valuably, we discover and characterize several types of fraud behaviors. We discuss how to leverage our insights in practice by engineering strongly performing entropy-based features and demonstrating high classification accuracy. Our contributions are (a) observations: we analyze our honeypot fraudster ecosystem and give surprising insights into the multifaceted behaviors of these fraudster types, and (b) features: we propose novel features that give strong (>0.95 precision/recall) discriminative power on ground-truth Twitter data. Neil Shah, Hemank Lamba, Alex Beutel, Christos Faloutsos |
ICDM | 2 |
| 2017 | From Camera to Deathbed: Understanding Dangerous Selfies on Social Media
Hemank Lamba, Varun Bharadhwaj, Mayank Vachher, Divyansh Agarwal, Megha Arora, Niharika Sachdeva, Ponnurangam Kumaraguru |
ICWSM | 1 |
| 2017 | zooRank: Ranking Suspicious Entities in Time-Evolving Tensors
Hemank Lamba, Bryan Hooi, Kijung Shin, Christos Faloutsos, Jürgen Pfeffer |
ECML/PKDD (1) | 1 |
| 2016 | Man-O-Meter: Modeling and Assessing the Evolution of Language Usage of Individuals on Microblogs
Kuntal Dey, Saroj Kaushik, Hemank Lamba, Seema Nagar |
APWeb (1) | 3 |
| 2015 | A Tempest in a Teacup? Analyzing Firestorms on Twitterabstract'Firestorms,' sudden bursts of negative attention in cases of controversy and outrage, are seemingly widespread on Twitter and are an increasing source of fascination and anxiety in the corporate, governmental, and public spheres. Using media mentions, we collect 80 candidate events from January 2011 to September 2014 that we would term 'firestorms.' Using data from the Twitter decahose (or gardenhose), a 10% random sample of all tweets, we describe the size and longevity of these firestorms. We take two firestorm exemplars, #myNYPD and #CancelColbert, as case studies to describe more fully. Then, taking the 20 firestorms with the most tweets, we look at the change in mention networks of participants over the course of the firestorm as one method of testing for possible impacts of firestorms. We find that the mention networks before and after the firestorms are more similar to each other than to those of the firestorms, suggesting that firestorms neither emerge from existing networks, nor do they result in lasting changes to social structure. To verify this, we randomly sample users and generate mention networks for baseline comparison, and find that the firestorms are not associated with a greater than random amount of change in mention networks. Hemank Lamba, Momin M. Malik, Jürgen Pfeffer |
ASONAM | 1 |
| 2015 | Experience-Aware Item Recommendation in Evolving Review CommunitiesabstractCurrent recommender systems exploit user and item similarities by collaborative filtering. Some advanced methods also consider the temporal evolution of item ratings as a global background process. However, all prior methods disregard the individual evolution of a user's experience level and how this is expressed in the user's writing in a review community. In this paper, we model the joint evolution of user experience, interest in specific item facets, writing style, and rating behavior. This way we can generate individual recommendations that take into account the user's maturity level (e.g., recommending art movies rather than blockbusters for a cinematography expert). As only item ratings and review texts are observables, we capture the user's experience and interests in a latent model learned from her reviews, vocabulary and writing style. We develop a generative HMM-LDA model to trace user evolution, where the Hidden Markov Model (HMM) traces her latent experience progressing over time -- with solely user reviews and ratings as observables over time. The facets of a user's interest are drawn from a Latent Dirichlet Allocation (LDA) model derived from her reviews, as a function of her (again latent) experience level. In experiments with four realworld datasets, we show that our model improves the rating prediction over state-of-the-art baselines, by a substantial margin. In addition, our model can also give some interpretations for the user experience level. Subhabrata Mukherjee, Hemank Lamba, Gerhard Weikum |
ICDM | 2 |
| 2014 | A Shapley Value-based Approach to Determine Gatekeepers in Social Networks with ApplicationsabstractInspired by emerging applications of social networks, we introduce in this paper a new centrality measure termed gate-keeper centrality. The new centrality is based on the well-known game-theoretic concept of Shapley value and, as we demonstrate, possesses unique qualities compared to the existing metrics. Furthermore, we present a dedicated approximate algorithm, based on the Monte Carlo sampling method, to compute the gatekeeper centrality. We also consider two well known applications in social network analysis, namely community detection and limiting the spread of mis-information; and show the merit of using the proposed framework to solve these two problems in comparison with the respective benchmark algorithms. Ramasuri Narayanam, Oskar Skibski, Hemank Lamba, Tomasz P. Michalak |
ECAI | 3 |
| 2013 | A Novel and Model Independent Approach for Efficient Influence Maximization in Social Networks
Hemank Lamba, Ramasuri Narayanam |
WISE (2) | 1 |
| 2012 | Incremental subclass discriminant analysis: A case study in face recognitionabstractSubclass discriminant analysis is found to be applicable under various scenarios. However, it is computationally expensive to update the between-class and within-class scatter matrices in batch mode. This research presents an incremental subclass discriminant analysis algorithm to update SDA in incremental manner with increasing number of samples per class. The effectiveness of the proposed algorithm is demonstrated using face recognition in terms of identification accuracy and training time. Experiments are performed on the AR face database and compared with other subspace based incremental and batch learning algorithms. The results illustrate that, compared to SDA, incremental SDA yields significant reduction in time along with comparable accuracy. Hemank Lamba, Tejas I. Dhamecha, Mayank Vatsa, Richa Singh 0001 |
ICIP | 1 |
| 2011 | Face recognition for look-alikes: A preliminary studyabstractOne of the major challenges efface recognition is to design a feature extractor and matcher that reduces the intra class variations and increases the inter-class variations. The feature extraction algorithm has to be robust enough to extract similar features for a particular subject despite variations in quality, pose, illumination, expression, aging, and disguise. The problem is exacerbated when there are two individuals with lower inter-class variations, i.e., look alikes. In such cases, the intra-class similarity is higher than the inter-class variation for these two individuals. This research explores the problem of look-alike faces and their effect on human performance and automatic face recognition algorithms. There is three fold contribution in this re search: firstly, we analyze the human recognition capabilities for look-alike appearances. Secondly, we compare human recognition performance with ten existing face recognition algorithms, and finally, proposed an algorithm to improve the face verification accuracy. The analysis shows that neither humans nor automatic face recognition algorithms are efficient in recognizing look-alikes. Hemank Lamba, Ankit Sarkar, Mayank Vatsa, Richa Singh 0001, Afzel Noore |
IJCB | 1 |