EDBT 2026 Demo / reviewers in the wild / expert
Ramya Tekumalla
dblp:261/9977
· DBLP profile ↗
4ranked-venue papers in the field
4as first author
3since 2021 · last 2023
0000-0002-1606-4856ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 3 (3 first)Information Retrieval & Web Search · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Towards automatic identification of self-reported COVID-19 tweets: Introducing a multilingual manually annotated dataset, baseline systems and exploratory evaluationsabstractIn recent times, social networks like Twitter have emerged as vital platforms for sharing personal thoughts, opinions, and most importantly, health-related information, especially pertaining to COVID-19. Users tend to share very detailed and personal narratives that could be utilized by researchers to capture true self-reported health data. While the data is easily accessible, the process to differentiate between health-related self-reports and informal discussion is quite tricky as it relies on either manual curation or the availability of large manually annotated datasets for machine learning models to be trained on. Manually annotating data is an immensely time-consuming task since, in general, the intervention of a subject matter expert is required, even more, in languages other than English, such as Spanish. In this work, we release two manually annotated datasets, one in English and one in Spanish, comprising of 36,548 tweets containing self-reported COVID-19 symptoms to aid machine learning models in extracting self-reported COVID-19 tweets. Using a very large set of experiments, we demonstrate how these datasets can be leveraged using classical and modern machine learning algorithms to identify unlabeled self-report tweets. Additionally, we perform a stratified analysis of how (and if) data augmentation and automatic translation could help train more generalizable models. Ramya Tekumalla, Luis Alberto Robles Hernandez, Juan M. Banda |
IEEE Big Data | 1 |
| 2022 | TweetDIS: A Large Twitter Dataset for Natural Disasters Built using Weak SupervisionabstractSocial media is often utilized as a lifeline for communication during natural disasters. Traditionally, natural disaster tweets are filtered f rom t he T witter s tream u sing t he n ame of the natural disaster and the filtered t weets a re s ent f or human annotation. The process of human annotation to create labeled sets for machine learning models is laborious, time consuming, at times inaccurate, and more importantly not scalable in terms of size and real-time use. In this work, we curated a silver standard dataset using weak supervision. In order to validate its utility, we train machine learning models on the weakly supervised data to identify three different types of natural disasters i.e earthquakes, hurricanes and floods. O ur r esults d emonstrate t hat models trained on the silver standard dataset achieved performance greater than 90% when classifying a manually curated, gold-standard dataset. To enable reproducible research and additional downstream utility, we release the silver standard dataset for the scientific community. Ramya Tekumalla, Juan M. Banda |
IEEE Big Data | 1 |
| 2022 | An Empirical Study on Characterizing Natural Disasters in Class Imbalanced Social Media Data using Weak SupervisionabstractSupervised learning has proven to be successful in classifying both class balanced and imbalanced data when a strong supervision signal is available. However, generating the supervision signal (eg: ground truth labels) is expensive and a major bottleneck of supervised learning. To curtail this, we rely on the theory of noisy learning and weak supervision to generate supervision signals. In this work, we utilize a noisy labeled dataset to train several class balanced and imbalanced machine learning models and compare the results to observe how efficient the models trained on silver standard dataset are in identifying ground truth labels. We demonstrate the approach on a natural disasters application which contains data from three different natural disasters. Our results demonstrate that theory of noisy learning can be utilized to build models via weak supervision for both class balanced and imbalanced data from social media sources for natural disasters application. Ramya Tekumalla, Juan M. Banda |
IEEE Big Data | 1 |
| 2020 | Mining Archive.org's Twitter Stream Grab for Pharmacovigilance Research Gold
Ramya Tekumalla, Javad Rafiei Asl, Juan M. Banda |
ICWSM | 1 |