EDBT 2026 Demo / reviewers in the wild / expert
Juan M. Banda
dblp:72/8217
· DBLP profile ↗
10ranked-venue papers in the field
4as first author
3since 2021 · last 2023
0000-0001-8499-824XORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 5 (1 first)Database Systems & Data Management · 3 (2 first)Information Retrieval & Web Search · 1Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Towards automatic identification of self-reported COVID-19 tweets: Introducing a multilingual manually annotated dataset, baseline systems and exploratory evaluationsabstractIn recent times, social networks like Twitter have emerged as vital platforms for sharing personal thoughts, opinions, and most importantly, health-related information, especially pertaining to COVID-19. Users tend to share very detailed and personal narratives that could be utilized by researchers to capture true self-reported health data. While the data is easily accessible, the process to differentiate between health-related self-reports and informal discussion is quite tricky as it relies on either manual curation or the availability of large manually annotated datasets for machine learning models to be trained on. Manually annotating data is an immensely time-consuming task since, in general, the intervention of a subject matter expert is required, even more, in languages other than English, such as Spanish. In this work, we release two manually annotated datasets, one in English and one in Spanish, comprising of 36,548 tweets containing self-reported COVID-19 symptoms to aid machine learning models in extracting self-reported COVID-19 tweets. Using a very large set of experiments, we demonstrate how these datasets can be leveraged using classical and modern machine learning algorithms to identify unlabeled self-report tweets. Additionally, we perform a stratified analysis of how (and if) data augmentation and automatic translation could help train more generalizable models. Ramya Tekumalla, Luis Alberto Robles Hernandez, Juan M. Banda |
IEEE Big Data | 3 |
| 2022 | TweetDIS: A Large Twitter Dataset for Natural Disasters Built using Weak SupervisionabstractSocial media is often utilized as a lifeline for communication during natural disasters. Traditionally, natural disaster tweets are filtered f rom t he T witter s tream u sing t he n ame of the natural disaster and the filtered t weets a re s ent f or human annotation. The process of human annotation to create labeled sets for machine learning models is laborious, time consuming, at times inaccurate, and more importantly not scalable in terms of size and real-time use. In this work, we curated a silver standard dataset using weak supervision. In order to validate its utility, we train machine learning models on the weakly supervised data to identify three different types of natural disasters i.e earthquakes, hurricanes and floods. O ur r esults d emonstrate t hat models trained on the silver standard dataset achieved performance greater than 90% when classifying a manually curated, gold-standard dataset. To enable reproducible research and additional downstream utility, we release the silver standard dataset for the scientific community. Ramya Tekumalla, Juan M. Banda |
IEEE Big Data | 2 |
| 2022 | An Empirical Study on Characterizing Natural Disasters in Class Imbalanced Social Media Data using Weak SupervisionabstractSupervised learning has proven to be successful in classifying both class balanced and imbalanced data when a strong supervision signal is available. However, generating the supervision signal (eg: ground truth labels) is expensive and a major bottleneck of supervised learning. To curtail this, we rely on the theory of noisy learning and weak supervision to generate supervision signals. In this work, we utilize a noisy labeled dataset to train several class balanced and imbalanced machine learning models and compare the results to observe how efficient the models trained on silver standard dataset are in identifying ground truth labels. We demonstrate the approach on a natural disasters application which contains data from three different natural disasters. Our results demonstrate that theory of noisy learning can be utilized to build models via weak supervision for both class balanced and imbalanced data from social media sources for natural disasters application. Ramya Tekumalla, Juan M. Banda |
IEEE Big Data | 2 |
| 2020 | Mining Archive.org's Twitter Stream Grab for Pharmacovigilance Research Gold
Ramya Tekumalla, Javad Rafiei Asl, Juan M. Banda |
ICWSM | 3 |
| 2019 | Solar Event Tracking with Deep Regression Networks: A Proof of Concept EvaluationabstractWith the advent of deep learning for computer vision tasks, the need for accurately labeled data in large volumes is vital for any application. The increasingly available large amounts of solar image data generated by the Solar Dynamic Observatory (SDO) mission make this domain particularly interesting for the development and testing of deep learning systems. The currently available labeled solar data is generated by the SDO mission's Feature Finding Team's (FFT) specialized detection modules. The major drawback of these modules is that detection and labeling is performed with a cadence of every 4 to 12 hours, depending on the module. Since SDO image data products are created every 10 seconds, there is a considerable gap between labeled observations and the continuous data stream. In order to address this shortcoming, we trained a deep regression network to track the movement of two solar phenomena: Active Region and Coronal Hole events. To the best of our knowledge, this is the first attempt of solar event tracking using a deep learning approach. Since it is impossible to fully evaluate the performance of the suggested event tracks with the original data (only partial ground truth is available), we demonstrate with several metrics the effectiveness of our approach. With the purpose of generating continuously labeled solar image data, we present this feasibility analysis showing the great promise of deep regression networks for this task. Toqi Tahamid Sarker, Juan M. Banda |
IEEE BigData | 2 |
| 2015 | Provenance-Centered Dataset of Drug-Drug Interactions
Juan M. Banda, Tobias Kuhn, Nigam H. Shah, Michel Dumontier |
ISWC (2) | 1 |
| 2014 | Scalable solar image Retrieval with LuceneabstractIn this work we present an alternative approach for large-scale retrieval of solar images using the highly-scalable retrieval engine Lucene. While Lucene is widely popular among text- based search engines, significant adjustments need to be made to take advantage of its fast indexing mechanism and highly-scalable architecture to enable search on image repositories. In this work we describe a novel way of representing image feature vectors in order to enable Lucene to perform search and retrieval of similar images. We compare our proposed method with other popular alternatives and provide commentary of the performance as well as the benefits and caveats of the proposed method. Juan M. Banda, Rafal A. Angryk |
IEEE BigData | 1 |
| 2013 | Big Data New Frontiers: Mining, Search and Management of Massive Repositories of Solar Image Data and Solar Events
Juan M. Banda, Michael A. Schuh, Rafal A. Angryk, Karthik Ganesan Pillai, Patrick McInerney |
ADBIS (2) | 1 |
| 2013 | When Too Similar Is Bad: A Practical Example of the Solar Dynamics Observatory Content-Based Image-Retrieval System
Juan M. Banda, Michael A. Schuh, Tim Wylie, Patrick McInerney, Rafal A. Angryk |
ADBIS (2) | 1 |
| 2013 | Spatiotemporal Co-occurrence Rules
Karthik Ganesan Pillai, Rafal A. Angryk, Juan M. Banda, Tim Wylie, Michael A. Schuh |
ADBIS (2) | 3 |