Juan M. Banda

dblp:72/8217 · DBLP profile ↗
← Back
10ranked-venue papers in the field
4as first author
3since 2021 · last 2023
0000-0001-8499-824XORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 5 (1 first)Database Systems & Data Management · 3 (2 first)Information Retrieval & Web Search · 1Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)
YearPublicationVenuePosition
2023 Towards automatic identification of self-reported COVID-19 tweets: Introducing a multilingual manually annotated dataset, baseline systems and exploratory evaluations
abstract
In recent times, social networks like Twitter have emerged as vital platforms for sharing personal thoughts, opinions, and most importantly, health-related information, especially pertaining to COVID-19. Users tend to share very detailed and personal narratives that could be utilized by researchers to capture true self-reported health data. While the data is easily accessible, the process to differentiate between health-related self-reports and informal discussion is quite tricky as it relies on either manual curation or the availability of large manually annotated datasets for machine learning models to be trained on. Manually annotating data is an immensely time-consuming task since, in general, the intervention of a subject matter expert is required, even more, in languages other than English, such as Spanish. In this work, we release two manually annotated datasets, one in English and one in Spanish, comprising of 36,548 tweets containing self-reported COVID-19 symptoms to aid machine learning models in extracting self-reported COVID-19 tweets. Using a very large set of experiments, we demonstrate how these datasets can be leveraged using classical and modern machine learning algorithms to identify unlabeled self-report tweets. Additionally, we perform a stratified analysis of how (and if) data augmentation and automatic translation could help train more generalizable models.
Ramya Tekumalla, Luis Alberto Robles Hernandez, Juan M. Banda
IEEE Big Data3
2022 TweetDIS: A Large Twitter Dataset for Natural Disasters Built using Weak Supervision
abstract
Social media is often utilized as a lifeline for communication during natural disasters. Traditionally, natural disaster tweets are filtered f rom t he T witter s tream u sing t he n ame of the natural disaster and the filtered t weets a re s ent f or human annotation. The process of human annotation to create labeled sets for machine learning models is laborious, time consuming, at times inaccurate, and more importantly not scalable in terms of size and real-time use. In this work, we curated a silver standard dataset using weak supervision. In order to validate its utility, we train machine learning models on the weakly supervised data to identify three different types of natural disasters i.e earthquakes, hurricanes and floods. O ur r esults d emonstrate t hat models trained on the silver standard dataset achieved performance greater than 90% when classifying a manually curated, gold-standard dataset. To enable reproducible research and additional downstream utility, we release the silver standard dataset for the scientific community.
Ramya Tekumalla, Juan M. Banda
IEEE Big Data2
2022 An Empirical Study on Characterizing Natural Disasters in Class Imbalanced Social Media Data using Weak Supervision
abstract
Supervised learning has proven to be successful in classifying both class balanced and imbalanced data when a strong supervision signal is available. However, generating the supervision signal (eg: ground truth labels) is expensive and a major bottleneck of supervised learning. To curtail this, we rely on the theory of noisy learning and weak supervision to generate supervision signals. In this work, we utilize a noisy labeled dataset to train several class balanced and imbalanced machine learning models and compare the results to observe how efficient the models trained on silver standard dataset are in identifying ground truth labels. We demonstrate the approach on a natural disasters application which contains data from three different natural disasters. Our results demonstrate that theory of noisy learning can be utilized to build models via weak supervision for both class balanced and imbalanced data from social media sources for natural disasters application.
Ramya Tekumalla, Juan M. Banda
IEEE Big Data2
2020 Mining Archive.org's Twitter Stream Grab for Pharmacovigilance Research Gold
Ramya Tekumalla, Javad Rafiei Asl, Juan M. Banda
ICWSM3
2019 Solar Event Tracking with Deep Regression Networks: A Proof of Concept Evaluation
abstract
With the advent of deep learning for computer vision tasks, the need for accurately labeled data in large volumes is vital for any application. The increasingly available large amounts of solar image data generated by the Solar Dynamic Observatory (SDO) mission make this domain particularly interesting for the development and testing of deep learning systems. The currently available labeled solar data is generated by the SDO mission's Feature Finding Team's (FFT) specialized detection modules. The major drawback of these modules is that detection and labeling is performed with a cadence of every 4 to 12 hours, depending on the module. Since SDO image data products are created every 10 seconds, there is a considerable gap between labeled observations and the continuous data stream. In order to address this shortcoming, we trained a deep regression network to track the movement of two solar phenomena: Active Region and Coronal Hole events. To the best of our knowledge, this is the first attempt of solar event tracking using a deep learning approach. Since it is impossible to fully evaluate the performance of the suggested event tracks with the original data (only partial ground truth is available), we demonstrate with several metrics the effectiveness of our approach. With the purpose of generating continuously labeled solar image data, we present this feasibility analysis showing the great promise of deep regression networks for this task.
Toqi Tahamid Sarker, Juan M. Banda
IEEE BigData2
2015 Provenance-Centered Dataset of Drug-Drug Interactions
Juan M. Banda, Tobias Kuhn, Nigam H. Shah, Michel Dumontier
ISWC (2)1
2014 Scalable solar image Retrieval with Lucene
abstract
In this work we present an alternative approach for large-scale retrieval of solar images using the highly-scalable retrieval engine Lucene. While Lucene is widely popular among text- based search engines, significant adjustments need to be made to take advantage of its fast indexing mechanism and highly-scalable architecture to enable search on image repositories. In this work we describe a novel way of representing image feature vectors in order to enable Lucene to perform search and retrieval of similar images. We compare our proposed method with other popular alternatives and provide commentary of the performance as well as the benefits and caveats of the proposed method.
Juan M. Banda, Rafal A. Angryk
IEEE BigData1
2013 Big Data New Frontiers: Mining, Search and Management of Massive Repositories of Solar Image Data and Solar Events
Juan M. Banda, Michael A. Schuh, Rafal A. Angryk, Karthik Ganesan Pillai, Patrick McInerney
ADBIS (2)1
2013 When Too Similar Is Bad: A Practical Example of the Solar Dynamics Observatory Content-Based Image-Retrieval System
Juan M. Banda, Michael A. Schuh, Tim Wylie, Patrick McInerney, Rafal A. Angryk
ADBIS (2)1
2013 Spatiotemporal Co-occurrence Rules
Karthik Ganesan Pillai, Rafal A. Angryk, Juan M. Banda, Tim Wylie, Michael A. Schuh
ADBIS (2)3