VLDB 2026 Research / reviewers in the wild / expert
Doina Caragea
dblp:40/2098
· DBLP profile ↗
61ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0002-6440-0914ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 6 first-author · 14 since 2021Databases, data management, data science and information retrieval · 33 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 since 2021Theory of computation · 2 · 2 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated behavior analysis in the novel object recognition testabstractAccurate identification of animal behaviors is fundamental to behavioral neuroscience, yet manual annotation remains a major bottleneck due to its subjectivity, limited scalability, and high labor cost. This work presents a unified deep learning and heuristic-based framework for automated behavior classification in the Novel Object Recognition Test (NORT), a widely used paradigm for assessing memory and cognitive function in rodents. We fine-tuned a YOLOv11 Pose model to detect nose and tail-base keypoints specific to Long-Evans rats, and used these keypoints to derive spatial heuristics for identifying NORT behaviors. In parallel, we trained YOLOv11 classification models to predict Standing, Object Interaction, and Other behaviors from individual frames, using datasets containing either 2-object or 5-object videos, as well as a combined dataset containing both 2-object and 5-object videos. Our model trained on the combined dataset, enhanced with heuristic-based post-processing to identify the specific objects that the rat is interacting with, achieved the best overall accuracy and generalization across arenas, outperforming heuristic-only and subset-specific 2-object or 5-object models. Detailed confusion matrices and error analyses reveal that most errors occur near behavioral transitions, reflecting the inherent ambiguity of manual labels at frame-level. The proposed framework enables scalable, reproducible, and objective annotation of NORT videos and provides a foundation for future extensions toward temporally-aware behavioral analysis. All the code and model weights, together with sample NORT videos, will be made publicly available to support future rodent behavior recognition research. Emily Alfs-Votipka, Bhavana Sivayokan, Aliva Bakshi, Sanaz Gheibuni, Doina Caragea, Bethany Plakke, Dave Turner, Daniel Andresen |
Neurocomputing | 5 |
| 2025 | Multimodal Disaster-Related Tweet Classification with Parameter-Efficient Fine-Tuning of Large Language Models
Dongping Guo, Xinli Xiao, Hongmin Li 0001, Doina Caragea |
ASONAM (3) | 5 |
| 2025 | Predictive Modeling and Explainable AI for Veterinary Safety Profiles, Residue Assessment, and Health Outcomes Using Real-World Data and Physicochemical Properties
Hossein Sholehrasa, Doina Caragea, Jim Edmond S. Riviere, Majid Jaberi Douraki |
IEEE Big Data | 3 |
| 2025 | Semi-Supervised Relation Extraction Informed by Area Under the Margin Ranking and Large Language ModelsabstractRelation extraction is an important task for understanding relationships between entities, building knowledge graphs, and facilitating knowledge discovery. Pre-trained models can be fine-tuned for relation extraction if a substantial amount of labeled data is available. However, acquiring extensive labeled data is generally challenging. Semi-supervised techniques for low-resource relation extraction, such as self-training, offer a promising solution by leveraging both limited labeled data and vast unlabeled data to mitigate this challenge. Traditional self-training methods use a teacher-student framework, where a student is iteratively trained with pseudo-labels generated by the teacher. This may lead to noisy pseudo-labels and impact performance. To address this limitation, we introduce a new model called RE-AUM-LLM that generates high-quality pseudo-labels using self-training combined with Area Under the Margin (AUM) and Large Language Models (LLMs), such as Llama 3.1. Experimental results on two benchmark datasets show that the proposed approach achieves state-of-the-art results for low-resource relation extraction by comparison with several strong baselines. We will make the code publicly available to enable reproducibility and further research in this area. Nikita Gautam, Bipin Paudel, Doina Caragea, Cornelia Caragea |
DSAA | 3 |
| 2025 | AutoPK: Leveraging LLMs and a Hybrid Similarity Metric for Advanced Retrieval of Pharmacokinetic Data From Complex Tables and DocumentsabstractPharmacokinetics (PK) plays a critical role in drug development and regulatory decision-making for human and veterinary medicine, directly affecting public health through drug safety and efficacy assessments. However, PK data are often embedded in complex, heterogeneous tables with variable structures and inconsistent terminologies, posing significant challenges for automated PK data retrieval and standardization. AutoPK, a novel two-stage framework for accurate and scalable extraction of PK data from complex scientific tables. In the first stage, AutoPK identifies and extracts PK parameter variants using large language models (LLMs), a hybrid similarity metric, and LLM-based validation. The second stage filters relevant rows, converts the table into a key-value text format, and uses an LLM to reconstruct a standardized, machine-readable table. Evaluated on a real-world dataset of 605 annotated PK tables, including captions and footnotes, AutoPK demonstrates significant improvements in precision and recall over direct LLM baselines. For instance, AutoPK with LLaMA 3.1-70B achieved an F1-score of 0.92 on half-life and 0.91 on clearance parameters, outperforming direct use of LLaMA 3.1-70B by margins of 0.10 and 0.21, respectively. Smaller models such as Gemma 3-27B and Phi 312B with AutoPK achieved 2-7 fold F1 gains over their direct use, with Gemma's hallucination rates reduced from 60-95% down to 8-14%. Notably, AutoPK enabled open-source models like Gemma 3-27B to outperform commercial systems such as GPT-4o Mini on several PK parameters. AutoPK enables scalable and high-confidence PK data extraction, making it well-suited for critical applications in veterinary pharmacology, drug safety monitoring, and public health decision-making, while addressing heterogeneous table structures and terminology and demonstrating generalizability across key PK parameters. Code and data are available at: https://github.com/hosseinsholehrasa/AutoPK. Hossein Sholehrasa, Amirhossein Ghanaatian, Doina Caragea, Lisa Ann Tell, Jim Edmond S. Riviere, Majid Jaberi Douraki |
ICTAI | 3 |
| 2024 | GunStance: Stance Detection for Gun Control and Gun RegulationabstractNikesh Gyawali, Iustin Sirbu, Tiberiu Sosea, Sarthak Khanal, Doina Caragea, Traian Rebedea, Cornelia Caragea. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Nikesh Gyawali, Iustin Sirbu, Tiberiu Sosea, Sarthak Khanal, Doina Caragea, Traian Rebedea, Cornelia Caragea |
ACL (1) | 5 |
| 2024 | Predicting Surface Water Bacteria Levels Using Transfer Learning and Domain AdaptationabstractSurface water contaminated by fecal bacteria can cause diarrheal illness, threatening human's health (especially among children). In recent years, supervised machine learning (ML) has been used to predict fecal indicator bacteria (FIB) levels. However, training ML models is challenging and, in some cases, even impractical due to sparsity of labeled data in all locations (e.g., in rural areas or low-income countries). In this paper, we introduce the largest water quality dataset available collected from beaches in Chicago and San Diego, USA. We utilized various models to predict historical FIB levels on this dataset establishing strong baseline models for supervised learning and transfer learning. Our models include Random Forest (RF), extreme gradient boosting (XGBoost), and attentionbased tabular deep learning (TabNet) models. Additionally, given the widespread use of large language models (LLMs), we have fine-tuned the LLaMA3-8B model for regression in a tabular-to-text setting. Our results show that supervised and unsupervised domain adaptation methods can enhance transfer learning performance. Specifically, the supervised methods, especially RF, represent a promising solution for FIB level prediction, while domain adaptation could be successfully employed to predict FIB levels in locations where they are rarely measured. Our code and dataset are available on: https://github.com/aliielahi/ONR-WQ. Ali Elahi, David Shumway, Megan Kowalcyk, Abhilasha Shrestha, Nikita Gautam, Doina Caragea, Cornelia Caragea, Samuel Dorevitch |
IEEE Big Data | 6 |
| 2024 | Contrastive Learning for Multimodal Classification of Crisis related TweetsabstractMultimodal tasks require learning a joint representation of the constituent modalities of data. Contrastive learning learns a joint representation by using a contrastive loss. For example, CLIP takes as input image-caption pairs and is trained to maximize the similarity between an image and its corresponding caption in actual image-caption pairs, while minimizing the similarity for arbitrary image-caption pairs. This approach operates on the premise that the caption depicts the image's content. However, this assumption does not always hold true for tweets that contain both text and images. Previous studies have indicated that the connection between the image and the text in a tweet is more intricate and complex. We study the effectiveness of pre-trained multimodal contrastive learning models, specifically, CLIP, and ALIGN, on the task of classifying multimodal crisis related tweets. Our experiments using two publicly available datasets, CrisisMMD and DMD, show that despite the intricate relationships in tweets, pre-trained contrastive learning models fine-tuned with task-specific data produce better results than prior approaches used for the multimodal classification of crisis related tweets. Additionally, the experiments show that the contrastive learning models are effective in low-data few-shot and cross-domain settings. Bishwas Mandal, Sarthak Khanal, Doina Caragea |
WWW | 3 |
| 2023 | A Comparison Study for Disaster Tweet Classification Using Deep Learning Models
Soudabeh Taghian Dinani, Doina Caragea |
DATA | 2 |
| 2023 | Disaster Image Classification Using Pre-trained Transformer and Contrastive Learning ModelsabstractNatural disasters can have devastating consequences for communities, causing loss of life and significant economic damage. To mitigate these impacts, it is crucial to quickly and accurately identify situational awareness and actionable information useful for disaster relief and response organizations. In this paper, we study the use of advanced transformer and contrastive learning models for disaster image classification in a humanitarian context, with focus on state-of-the-art pre-trained vision transformers such as ViT, CSWin and a state-of-the-art pre-trained contrastive learning model, CLIP. We evaluate the performance of these models across various disaster scenarios, including in-domain and cross-domain settings, as well as few-shot learning and zero-shot learning settings. Our results show that the CLIP model outperforms the two transformer models (ViT and CSWin) and also ConvNeXts, a competitive CNN-based model resembling transformers, in all the settings. By improving the performance of disaster image classification, our work can contribute to the goal of reducing the number of deaths and economic losses caused by disasters, as well as helping to decrease the number of people affected by these events. Soudabeh Taghian Dinani, Doina Caragea |
DSAA | 2 |
| 2023 | Leveraging Existing Literature on the Web and Deep Neural Models to Build a Knowledge Graph Focused on Water Quality and Health RisksabstractA knowledge graph focusing on water quality in relation to health risks posed by water activities (such as diving or swimming) is not currently available. To address this limitation, we first use existing resources to construct a knowledge graph relevant to water quality and health risks using KNowledge Acquisition and Representation Methodology (KNARM). Subsequently, we explore knowledge graph completion approaches for maintaining and updating the graph. Specifically, we manually identify a set of domain-specific UMLS concepts and use them to extract a graph of approximately 75,000 semantic triples from the Semantic MEDLINE database (which contains head-relation-tail triples extracted from PubMed). Using the resulting knowledge graph, we experiment with the KG-BERT approach for graph completion by employing pre-trained BERT/RoBERTa models and also models fine-tuned on a collection of water quality and health risks abstracts retrieved from the Web of Science. Experimental results show that KG-BERT with BERT/RoBERTa models fine-tuned on a domain-specific corpus improves the performance of KG-BERT with pre-trained models. Furthermore, KG-BERT gives better results than several translational distance or semantic matching baseline models. Nikita Gautam, David Shumway, Megan Kowalcyk, Sarthak Khanal, Doina Caragea, Cornelia Caragea, Hande McGinty, Samuel Dorevitch |
WWW | 5 |
| 2022 | Multimodal Semi-supervised Learning for Disaster Tweet ClassificationabstractDuring natural disasters, people often use social media platforms, such as Twitter, to post information about casualties and damage produced by disasters. This information can help relief authorities gain situational awareness in nearly real time, and enable them to quickly distribute resources where most needed. However, annotating data for this purpose can be burdensome, subjective and expensive. In this paper, we investigate how to leverage the copious amounts of unlabeled data generated on social media by disaster eyewitnesses and affected individuals during disaster events. To this end, we propose a semi-supervised learning approach to improve the performance of neural models on several multimodal disaster tweet classification tasks. Our approach shows significant improvements, obtaining up to 7.7% improvements in F-1 in low-data regimes and 1.9% when using the entire training data. We make our code and data publicly available at https://github.com/iustinsirbu13/multimodal-ssl-for-disaster-tweet-classification. Iustin Sirbu, Tiberiu Sosea, Cornelia Caragea, Doina Caragea, Traian Rebedea |
COLING | 4 |
| 2022 | Using Deep Learning to Improve Detection and Decoding Of BarcodesabstractWe propose an end-to-end pipeline for transforming raw images containing barcodes into sharp barcode images that can be accurately decoded. Our pipeline leverages recent deep learning approaches and consists of a rotation-decoupled detector (RDD) for oriented barcode detection and a deblurring model. The deblurring model uses a generative adversarial network (specifically, DeblurGAN-v2) trained on pairs of noisy and sharp images generated using ground truth numeric codes. Evaluation of the proposed pipeline using real barcode images enhanced by the DeblurGAN-v2 model shows a 14% improvement of the decoding rate as compared to the decoding rate obtained on the original images. Chaoxin Wang, Nicolais Guevara, Doina Caragea |
ICIP | 3 |
| 2022 | Identification of Fine-Grained Location Mentions in Crisis TweetsabstractIdentification of fine-grained location mentions in crisis tweets is central in transforming situational awareness information extracted from social media into actionable information. Most prior works have focused on identifying generic locations, without considering their specific types. To facilitate progress on the fine-grained location identification task, we assemble two tweet crisis datasets and manually annotate them with specific location types. The first dataset contains tweets from a mixed set of crisis events, while the second dataset contains tweets from the global COVID-19 pandemic. We investigate the performance of state-of-the-art deep learning models for sequence tagging on these datasets, in both in-domain and cross-domain settings. Sarthak Khanal, Maria Traskowsky, Doina Caragea |
LREC | 3 |
| 2021 | Author Homepage Discovery in CiteSeerXabstractScholarly digital libraries provide access to scientific publications and comprise useful resources for researchers. CiteSeerX is one such digital library search engine that provides access to more than 10 million academic documents. We propose a novel search-driven approach to build and maintain a large collection of homepages that can be used as seed URLs in any digital library including CiteSeerX to crawl scientific documents. Precisely, we integrate Web search and classification in a unified approach to discover new homepages: first, we use publicly-available author names and research paper titles as queries to a Web search engine to find relevant content, and then we identify the correct homepages from the search results using a powerful deep learning classifier based on Convolutional Neural Networks. Moreover, we use Self-Training in order to reduce the labeling effort and to utilize the unlabeled data to train the efficient researcher homepage classifier. Our experiments on a large scale dataset highlight the effectiveness of our approach, and position Web search as an effective method for acquiring authors' homepages. We show the development and deployment of the proposed approach in CiteSeerX and the maintenance requirements. Krutarth Patel, Cornelia Caragea, Doina Caragea, C. Lee Giles |
AAAI | 3 |
| 2021 | Stance Detection in COVID-19 TweetsabstractKyle Glandt, Sarthak Khanal, Yingjie Li, Doina Caragea, Cornelia Caragea. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Kyle Glandt, Sarthak Khanal, Yingjie Li 0008, Doina Caragea, Cornelia Caragea |
ACL/IJCNLP (1) | 4 |
| 2021 | Disaster Image Classification Using Capsule NetworksabstractWhen a disaster happens, affected individuals may use social media platforms, such as Twitter or Facebook, to ask for help or post information about the disaster. From a disaster response point of view, it is important to filter posts, in particular, text and images that provide situational awareness information, in a timely manner. For image classification, capsule networks have shown superiority over convolutional neural networks (CNN). Given their success in other application domains, in this study, we used capsule networks to classify disaster images as Informative or Non-informative. Using publicly available images collected from several disasters, we compared capsule network models with ResNet-18 models, for both in-domain and cross-domain settings. The results showed that the capsule network models had better performance for all the disaster datasets considered in the in-domain experiments, and also for most of the cross-domain pairs of disasters used in the study. Soudabeh Taghian Dinani, Doina Caragea |
IJCNN | 2 |
| 2020 | On Identifying Hashtags in Disaster Twitter DataabstractTweet hashtags have the potential to improve the search for information during disaster events. However, there is a large number of disaster-related tweets that do not have any user-provided hashtags. Moreover, only a small number of tweets that contain actionable hashtags are useful for disaster response. To facilitate progress on automatic identification (or extraction) of disaster hashtags for Twitter data, we construct a unique dataset of disaster-related tweets annotated with hashtags useful for filtering actionable information. Using this dataset, we further investigate Long Short-Term Memory-based models within a Multi-Task Learning framework. The best performing model achieves an F1-score as high as $92.22%$. The dataset, code, and other resources are available on Github.1 Jishnu Ray Chowdhury, Cornelia Caragea, Doina Caragea |
AAAI | 3 |
| 2020 | Identifying FinTech Innovations Using BERTabstractAdvancements in technology have resulted in the emergence of numerous FinTech innovations. However, a global understanding of such innovations is limited, due to a lack of an underlying taxonomy and benchmark datasets in the FinTech domain. To address this limitation, we develop a FinTech taxonomy and manually annotate a set of FinTech patent abstracts according to the taxonomy. We use the annotated dataset to train deep learning models, specifically recurrent neural networks and convolutional neural networks combined with state-of-the-art BERT transformers. Experimental results show that the deep learning models can accurately identify FinTech innovations. We use our best performing BERT-based model on a large dataset of financial patent abstracts, and shortlist a set of 25,580 FinTech patent applications submitted to the European and US Patent Offices between 2000 and 2017. We illustrate how an analysis of the shortlisted set can be used to gain understanding of what FinTech innovations are, where and when they emerge, and provide the basis for further work on what their impact is on the companies investing in them, and ultimately on society. Doina Caragea, Theodor Cojoianu, Mihai Dobri, Kyle Glandt, George Mihaila |
IEEE BigData | 1 |
| 2020 | Domain Adaptation with Reconstruction for Disaster Tweet ClassificationabstractIdentifying critical information in real time in the beginning of a disaster is a challenging but important task. This task has been recently addressed using domain adaptation approaches, which eliminate the need for target labeled data, and can thus accelerate the process of identifying useful information. We propose to investigate the effectiveness of the Domain Reconstruction Classification Network (DRCN) approach on disaster tweets. DRCN adapts information from target data by reconstructing it with an autoencoder. Experimental results using a sequence-to-sequence autoencodershow that the DRCN approach can improve the performance of both supervised and domain adaptation baseline models. Xukun Li, Doina Caragea |
SIGIR | 2 |
| 2020 | Using AI and Social Media Multimodal Content for Disaster Response and Management: Opportunities, Challenges, and Future Directions
Muhammad Imran 0002, Ferda Ofli, Doina Caragea, Antonio Torralba 0001 |
Inf. Process. Manag. | 3 |
| 2019 | Identifying Android Malware Using Network-Based ApproachesabstractThe proliferation of Android apps has resulted in many malicious apps entering the market and causing significant damage. Robust techniques that determine if an app is malicious are greatly needed. We propose the use of a network-based approach to effectively separate malicious from benign apps, based on a small labeled dataset. The apps in our dataset come from the Google Play Store and have been scanned for malicious behavior using Virus Total to produce a ground truth dataset with labels malicous or benign. The apps in the resulting dataset have been represented using binary feature vectors (where the features represent permissions, intent actions, discriminative APIs, obfuscation signatures, and native code signatures). We have used the feature vectors corresponding to apps to build a weighted network that captures the “closeness” between apps. We propagate labels from the labeled apps to unlabeled apps, and evaluate the effectiveness of the proposed approach using the F1-measure. We have conducted experiments to compare three variants of the label propagation approaches on datasets that include increasingly larger amounts of labeled data. The results have shown that a variant proposed in this study gives the best results overall. Emily Alfs, Doina Caragea, Nathan Albin, Pietro Poggi-Corradini |
AAAI | 2 |
| 2019 | Identifying Android malware using network-based approachesabstractThe proliferation of Android applications has resulted in many malicious apps entering the market and causing significant damage. Robust techniques that determine if an app is malicious are greatly needed. We propose the use of network-based approaches to effectively separate malicious from benign apps, based on a small labeled dataset. The apps in our dataset come from the Google Play Store and have been scanned for malicious behavior using VirusTotal to produce a ground truth dataset with labels malicious or benign. The apps in the resulting dataset have been represented in the form of binary feature vectors (where the features represent permissions, intent actions, discriminative APIs, obfuscation signatures, and native code signatures). We have used these vectors to build a weighted network that captures the "closeness" between apps. We propagate labels from the labeled apps to unlabeled apps, and evaluate the effectiveness of the approaches studied using the F1-measure. We have conducted experiments to compare three variants of the label propagation approaches on datasets that consist of increasingly larger amounts of labeled data. Emily Alfs, Doina Caragea, Dewan Chaulagain, Sankardas Roy, Nathan Albin, Pietro Poggi-Corradini |
ASONAM | 2 |
| 2019 | Keyphrase Extraction from Disaster-related TweetsabstractWhile keyphrase extraction has received considerable attention in recent years, relatively few studies exist on extracting keyphrases from social media platforms such as Twitter, and even fewer for extracting disaster-related keyphrases from such sources. During a disaster, keyphrases can be extremely useful for filtering relevant tweets that can enhance situational awareness. Previously, joint training of two different layers of a stacked Recurrent Neural Network for keyword discovery and keyphrase extraction had been shown to be effective in extracting keyphrases from general Twitter data. We improve the model's performance on both general Twitter data and disaster-related Twitter data by incorporating contextual word embeddings, POS-tags, phonetics, and phonological features. Moreover, we discuss the shortcomings of the often used F1-measure for evaluating the quality of predicted keyphrases with respect to the ground truth annotations. Instead of the F1-measure, we propose the use of embedding-based metrics to better capture the correctness of the predicted keyphrases. In addition, we also present a novel extension of an embedding-based metric. The extension allows one to better control the penalty for the difference in the number of ground-truth and predicted keyphrases. Jishnu Ray Chowdhury, Cornelia Caragea, Doina Caragea |
WWW | 3 |
| 2018 | Localizing and Quantifying Damage in Social Media ImagesabstractTraditional post-disaster assessment of damage heavily relies on expensive GIS data, especially remote sensing image data. In recent years, social media has become a rich source of disaster information that may be useful in assessing damage at a lower cost. Such information includes text (e.g., tweets) or images posted by eyewitnesses of a disaster. Most of the existing research explores the use of text in identifying situational awareness information useful for disaster response teams. The use of social media images to assess disaster damage is limited. In this paper, we propose a novel approach, based on convolutional neural networks and class activation maps, to locate damage in a disaster image and to quantify the degree of the damage. Our proposed approach enables the use of social network images for post-disaster damage assessment, and provides an inexpensive and feasible alternative to the more expensive GIS approach. Xukun Li, Doina Caragea, Huaiyu Zhang, Muhammad Imran 0002 |
ASONAM | 2 |
| 2017 | Android Malware Detection with Weak Ground Truth DataabstractFor Android malware detection, precise ground truth is a rare commodity. As security knowledge evolves, what may be considered ground truth at one moment in time may change, and apps once considered benign may turn out to be malicious. The inevitable noise in data labels poses a challenge to inferring effective machine learning classifiers. Our work is focused on approaches for learning classifiers for Android malware detection in a manner that is methodologically sound with regard to the uncertain and ever-changing ground truth in the problem space. We leverage the fact that although data labels are unavoidably noisy, a malware label is much more precise than a benign label. While you can be confident that an app is malicious, you can never be certain that a benign app is really benign, or just undetected malware. Based on this insight, we leverage a modified Logistic Regression classifier that allows us to learn from only positive and unlabeled data, without making any assumptions about benign labels. We find Label Regularized Logistic Regression to perform well for noisy app datasets, as well as datasets where there is a limited amount of positive labeled data, both of which are representative of real-world situations. Jordan DeLoach, Doina Caragea, Xinming Ou |
AAAI | 2 |
| 2017 | Twitter-enhanced Android malware detectionabstractIn data-driven Android malware detection, large numbers of both malicious and benign apps are used to train machine learning classifiers to detect malware. Existing approaches have nearly exclusively focused on app contents to extract features for classification. We seek to understand if auxiliary data, specifically Twitter data, can be used to improve the performance of existing approaches for Android malware detection. Throughout the course of our research, we collected over 50 million tweets potentially related to Android apps. We propose to link tweets with apps using approaches inspired from the standard vector space model, and subsequently study the usefulness of the linked tweets in malware detection. We find that Twitter data accurately linked to apps through HTTP links can be used to improve the machine learning classifier performance across a variety of common malware detection classifiers. However, classification experiments with Twitter data automatically linked to apps reveal the need for future work on more robust linking approaches. Jordan DeLoach, Doina Caragea |
IEEE BigData | 2 |
| 2016 | Study of transductive learning and unsupervised feature construction methods for biological sequence classificationabstractNext Generation Sequencing (NGS) technologies have led to fast and inexpensive production of large amounts of biological sequence data, including nucleotide sequences and derived protein sequences. These fast-increasing volumes of data pose challenges to computational methods for annotation. Machine learning approaches, primarily supervised algorithms, have been widely used to assist with classification tasks in bioinformatics. However, supervised algorithms rely on large amounts of labeled data in order to produce quality predictors. Oftentimes, labeled data is difficult and expensive to acquire in sufficiently large quantities. When only limited amounts of labeled data but considerably larger amounts of unlabeled data are available for a specific annotation problem, semi-supervised learning approaches represent a cost-effective alternative. In this work, we focus on a special case of semi-supervised learning, namely transductive learning, in which the algorithm has access during the training phase to the instances that need to be labeled. Transduction is particularly suitable for biological sequence classification, where the goal is generally to label a given set of unlabeled instances. However, a challenge that needs to be addressed in this context consists of identification of compact sets of informative features. Given the lack of labeled data, standard supervised feature selection methods may result in unreliable features. Therefore, we study recently proposed unsupervised feature construction approaches together with transductive learning. Experimental results on two classification problems, namely cassette exon identification and protein localization, show that the unsupervised features result in better performance than the supervised features. Ana Stanescu 0001, Karthik Tangirala, Doina Caragea |
ASONAM | 3 |
| 2016 | Android malware detection with weak ground truth dataabstractFor Android malware detection, precise ground truth is a rare commodity. As security knowledge evolves, what may be considered ground truth at one moment in time may change, and apps once considered benign turn out to be malicious. The inevitable noise in data labels poses a challenge to creating effective machine learning models. Our work is focused on approaches for learning classifiers for Android malware detection in a manner that is methodologically sound with regard to the uncertain and ever-changing ground truth in the problem space. We leverage the fact that although data labels are unavoidably noisy, a malware label is much more precise than a benign label. While you can be confident that an app is malicious, you can never be certain that a benign app is really benign or just an undetected malware. Based on this insight, we leverage a modified Logistic Regression classifier that allows us to learn from only positive and unlabeled data, without making any assumptions about benign labels. We find Label Regularized Logistic Regression to perform well for noisy app datasets, as well as datasets where there is a limited amount of positive labeled data, both of which are representative of real-world situations. Jordan DeLoach, Doina Caragea, Xinming Ou |
IEEE BigData | 2 |
| 2015 | Experimental Study with Real-world Data for Android App Security Analysis using Machine LearningabstractAlthough Machine Learning (ML) based approaches have shown promise for Android malware detection, a set of critical challenges remain unaddressed. Some of those challenges arise in relation to proper evaluation of the detection approach while others are related to the design decisions of the same. In this paper, we systematically study the impact of these challenges as a set of research questions (i.e., hypotheses). We design an experimentation framework where we can reliably vary several parameters while evaluating ML-based Android malware detection approaches. The results from the experiments are then used to answer the research questions. Meanwhile, we also demonstrate the impact of some challenges on some existing ML-based approaches. The large (market-scale) dataset (benign and malicious apps) we use in the above experiments represents the real-world Android app security analysis scale. We envision this study to encourage the practice of employing a better evaluation strategy and better designs of future ML-based approaches for Android malware detection. Sankardas Roy, Jordan DeLoach, Nic Herndon, Doina Caragea, Xinming Ou, Venkatesh Prasad Ranganath, Hongmin Li 0001, Nicolais Guevara |
ACSAC | 5 |
| 2015 | An Evaluation of Self-training Styles for Domain Adaptation on the Task of Splice Site PredictionabstractWe consider the problem of adding a large unlabeled sample from the target domain to boost the performance of a domain adaptation algorithm when only a small set of labeled examples are available from the target domain. In particular, we consider the problem setting motivated by the task of splice site prediction. For this task, annotating a genome using machine learning requires a lot of labeled data, whereas for non-model organisms, there is only some labeled data and lots of unlabeled data. With domain adaptation one can leverage the large amount of data from a related model organism, along with the labeled and unlabeled data from the organism of interest to train a classifier for the latter. Our goal is to analyze the three ways of incorporating the unlabeled data -- with soft labels only (i.e., Expectation-Maximization), with hard labels only (i.e., self-training), or with both soft and hard labels -- for the splice site prediction in particular, and more broadly for a general iterative domain adaptation setting. We provide empirical results on splice site prediction indicating that using soft labels only can lead to better classifier compared to the other two ways. Nic Herndon, Doina Caragea |
ASONAM | 2 |
| 2015 | Predicting cassette exons using transductive learning approachesabstractRecent advances in biotechnology have resulted in large volumes of genomic and proteomic data leading to the emergence of numerous in silico methods for annotation, such as supervised machine learning approaches. Such algorithms, however, require large amounts of labeled data for training. In practice, labeled data is oftentimes limited because it is difficult to obtain. Therefore, semi-supervised machine learning is preferable, in which classifiers trained on limited amounts of labeled data can be improved by exploiting the large amounts of unlabeled data. In this work, we focus on transductive learning, a special case of semi-supervised learning. A semi-supervised algorithm builds an inductive model that generalizes well to new, unseen (test) instances. In contrast, during the training phase, a transductive algorithm has access to the (test) instances that need to be classified, allowing advantageous utilization of these points in order to reach the best separation function. Compared to learning a classifier for use with future data, cassette exon identification is a suitable application for transductive learning, since the goal is to annotate a sequenced genome for which a limited amount of labeled data is available. We study the applicability of three popular transductive techniques and their compatibility with various kernels to the binary DNA classification problem of cassette exon identification. The results of our experiments suggest that transductive learning is a useful approach for assisting genome annotation. Ana Stanescu 0001, Doina Caragea |
CIBCB | 2 |
| 2015 | Domain Adaptation with Logistic Regression for the Task of Splice Site Prediction
Nic Herndon, Doina Caragea |
ISBRA | 2 |
| 2015 | Community Detection-Based Feature Construction for Protein Sequence Classification
Karthik Tangirala, Nic Herndon, Doina Caragea |
ISBRA | 3 |
| 2014 | Ensemble-based semi-supervised learning approaches for imbalanced splice site datasetsabstractProducing accurate classifiers depends on the quality and quantity of labeled data. The lack of labeled data, due to its expensive generation, critically affects the application of machine learning algorithms to biological problems. However, unlabeled data may be acquired relatively faster and in larger quantities thanks to current biochemical technologies, called Next Generation Sequencing. In such cases, when the number of labeled instances is overwhelmed by the number of unlabeled instances, semi-supervised learning represents a cost-effective alternative that can improve supervised classifiers by utilizing unlabeled data. In practice, data oftentimes exhibits imbalanced class distributions, which represents an obstacle for both supervised and semi-supervised learning. The problem of supervised learning from imbalanced datasets has been extensively studied, and various solutions have been proposed to produce classifiers with optimal performance on highly skewed class distributions. In the case of semi-supervised learning, there are not as many efforts aimed at the imbalance data problem. In this paper, we study several ensemble-based semi-supervised learning approaches for predicting splice sites, a problem for which the imbalance ratio is very high. We run experiments on five imbalanced datasets with the goal of identifying which variants are the most effective. Ana Stanescu 0001, Doina Caragea |
BIBM | 2 |
| 2014 | Predicting protein localization using a domain adaptation na¨ıve Bayes classifier with burrows wheeler transform featuresabstractThe reduced cost of the next generation sequencing technologies provides opportunities to study non-model organisms. However, one challenge is the large volume of data generated and, thus, the need to use automated approaches to annotate these data. Machine learning algorithms could provide a cost-effective solution but they need lots of labeled data and informative features to represent these data. Our proposed approach addresses both these problems by using a domain adaptation classifier in conjunction with features generated with unsupervised techniques to annotate biological sequence data. Nic Herndon, Karthik Tangirala, Doina Caragea |
BIBM | 3 |
| 2013 | Aiding Intrusion Analysis Using Machine LearningabstractIntrusion analysis, i.e., the process of combing through IDS alerts and audit logs to identify real successful and attempted attacks, remains a difficult problem in practical network security defense. The major contributing cause to this problem is the high false-positive rate in the sensors used by IDS systems to detect malicious activities. The goal of our work is to examine whether a machine-learned classifier can help a human analyst filter out non-interesting scenarios reported by an IDS alert correlator, so that analysts' time can be saved. This research is conducted in the open-source SnIPS intrusion analysis framework. Throughout observing the output of SnIPS running on our departmental network, we found that an analyst would need to perform repetitive tasks in pruning out the false positives in the correlation graphs produced by it. We hypothesized that such repetitive tasks can yield (limited) labeled data that can enable the use of a machine learning-based approach to prune SnIPS' output based on the human analysts' feedback, much similar to spam filters that can learn from users' past judgment to prune emails. Our goal is to classify the correlation graphs produced from SnIPS into "interesting" and "non-interesting", where "interesting" means that a human analyst would want to conduct further analysis on the events. We spent significant amount of time manually labeling SnIPS' output correlations based on this criterion, and built prediction models using both supervised and semi-supervised learning approaches. Our experiments revealed a number of interesting observations that give insights into the pitfalls and challenges of applying machine learning in intrusion analysis. The experimentation results also indicate that semi-supervised learning is a promising approach towards practical machine learning-based tools that can aid human analysts, when a limited amount of labeled data is available. Loai Zomlot, Sathya Chandran Sundaramurthy, Doina Caragea, Xinming Ou |
ICMLA (2) | 3 |
| 2013 | Economic Development through Business Profiling: A Text Analysis Based ApproachabstractThe tremendous improvements in the field of web technologies have contributed to the accumulation of large amounts of text data, particularly in the form of websites. Among others, most businesses, smaller or bigger, present themselves to the world through their websites. Economic development analysts could make use of the information available on business websites to identify ways in which businesses in a region can be grouped together into clusters, and possibly partnership for mutual benefit, and thus for economics gains of that particular region. Automated clustering of businesses is especially useful, as the existing NAICS codes-based clustering is not very accurate, according to domain experts, and does not scale up well. Text analysis is a blooming field whose goal is to automatically extract useful information from natural language text. In this work, we perform a preliminary text analysis of business websites to build business profiles and to organize businesses into clusters. Our approach is based on representing businesses as a mixture of ``topics" using a technique called Latent Dirichlet Allocation (LDA). Given the business profiles represented as the topic distributions obtained with LDA, we construct preliminary clusters of businesses. Manual analysis of the results shows that the proposed approach has the potential for giving interesting and useful clusters, which have the potential to replace the existing NAICS codes-based clusters. We identify further challenges associated with the existing business data, and several ideas for future work. Rohit Parimi, Doina Caragea, Dale Wunderlich |
Web Intelligence | 2 |
| 2013 | A Hybrid Recommender System: User Profiling from Keywords and RatingsabstractOver the last decade, user-generated content has grown continuously. Recommender systems that exploit user feedback are widely used in e-commerce and quite necessary for business enhancement. To make use of such user feedback, we propose a new content/collaborative hybrid approach, which is built on top of the recently released hetrec2011-movielens-2k dataset and is an extension of a previously proposed neighborhood based approach, called Weighted Tag Recommender (WTR). Our approach has two versions. Both versions make use of ratings to enable collaborative filtering and use either user tags, available in the hetrec2011-movielens-2k dataset, or movie keywords retrieved from IMDB, to capture movie content information. Experimental results show that the information from keywords can help build a movie recommender system competitive with other neighborhood based approaches and even with more sophisticated state-of-the-art approaches. Ana Stanescu 0001, Swapnil Nagar, Doina Caragea |
Web Intelligence | 3 |
| 2011 | Semi-supervised Learning of Alternatively Spliced Exons Using Co-trainingabstractAlternative splicing is a phenomenon that gives rise to multiple mRNA transcripts from a single gene. It is believed that a large number of genes undergoes alternative splicing. Predicting alternative splicing events is a problem of great interest, as it can help the understanding of transcript diversity. Supervised machine learning approaches can be used to predict alternative splicing events at genome level. However, supervised approaches require large amounts of labeled data to learn accurate classifiers. While large amounts of genomic data are produced by the new sequencing technologies, labeling these data can be costly and time consuming. Therefore, semi- supervised learning approaches that can make use of large amounts of unlabeled data, in addition to small amounts of labeled data are highly desirable. In this work, we study the usefulness of a semi-supervised learning approach, co-training, for classifying exons as alternatively spliced or constitutive. The co-training algorithm makes use of two views of the data to iteratively learn two classifiers that can inform each other, at each step, with their best predictions on the unlabeled data. We consider two sets of features for constructing views for the problem of predicting alternatively spliced exons: exonic splicing enhancers and intronic regulatory sequences. We use the Naive Bayes Multinomial algorithm as a base classifier in our study. Experimental results show that the usage of the unlabeled data can result in better classifiers as compared to those obtained from the small amount of labeled data alone. Karthik Tangirala, Doina Caragea |
BIBM | 2 |
| 2011 | An Empirical Study on Using the National Vulnerability Database to Predict Software Vulnerabilities
Doina Caragea, Xinming Ou |
DEXA (1) | 2 |
| 2011 | Predicting Friendship Links in Social Networks Using a Topic Modeling Approach
Rohit Parimi, Doina Caragea |
PAKDD (2) | 2 |
| 2010 | Abstraction Augmented Markov ModelsabstractHigh accuracy sequence classification often requires the use of higher order Markov models (MMs). However, the number of MM parameters increases exponentially with the range of direct dependencies between sequence elements, thereby increasing the risk of overfitting when the data set is limited in size. We present abstraction augmented Markov models (AAMMs) that effectively reduce the number of numeric parameters of k(th) order MMs by successively grouping strings of length k (i.e., k-grams) into abstraction hierarchies. We evaluate AAMMs on three protein subcellular localization prediction tasks. The results of our experiments show that abstraction makes it possible to construct predictive models that use significantly smaller number of features (by one to three orders of magnitude) as compared to MMs. AAMMs are competitive with and, in some cases, significantly outperform MMs. Moreover, the results show that AAMMs often perform significantly better than variable order Markov models, such as decomposed context tree weighting, prediction by partial match, and probabilistic suffix trees. Cornelia Caragea, Adrian Silvescu, Doina Caragea, Vasant G. Honavar |
ICDM | 3 |
| 2010 | Boosting Biomedical Entity Extraction by Using Syntactic Patterns for Semantic Relation DiscoveryabstractBiomedical entity extraction from unstructured web documents is an important task that needs to be performed in order to discover knowledge in the veterinary medicine domain. In general, this task can be approached by applying domain specific ontologies, but a review of the literature shows that there is no universal dictionary, or ontology for this domain. To address this issue, we manually construct an ontology for extracting entities such as: animal disease names, viruses and serotypes. We then use an automated ontology expansion approach to extract semantic relationships between concepts. Such relationships include asserted synonymy, hyponymy and causality. Specifically, these relationships are extracted by using a set of syntactic patterns and part-of-speech tagging. The resulting ontology contains richer semantics compared to the manually constructed ontology. We compare our approach for extracting synonyms, hyponyms and other disease related concepts, with an approach where the ontology is expanded using GoogleSets, on the veterinary medicine entity extraction task. Experimental results show that our semantic relationship extraction approach produces a significant increase in precision and recall as compared to the GoogleSets approach. Svitlana Volkova, Doina Caragea, William H. Hsu, John Drouhard, Landon Fowles |
Web Intelligence | 2 |
| 2010 | Semi-supervised prediction of protein subcellular localization using abstraction augmented Markov modelsabstractBACKGROUND: Determination of protein subcellular localization plays an important role in understanding protein function. Knowledge of the subcellular localization is also essential for genome annotation and drug discovery. Supervised machine learning methods for predicting the localization of a protein in a cell rely on the availability of large amounts of labeled data. However, because of the high cost and effort involved in labeling the data, the amount of labeled data is quite small compared to the amount of unlabeled data. Hence, there is a growing interest in developing semi-supervised methods for predicting protein subcellular localization from large amounts of unlabeled data together with small amounts of labeled data. RESULTS: In this paper, we present an Abstraction Augmented Markov Model (AAMM) based approach to semi-supervised protein subcellular localization prediction problem. We investigate the effectiveness of AAMMs in exploiting unlabeled data. We compare semi-supervised AAMMs with: (i) Markov models (MMs) (which do not take advantage of unlabeled data); (ii) an expectation maximization (EM); and (iii) a co-training based approaches to semi-supervised training of MMs (that make use of unlabeled data). CONCLUSIONS: The results of our experiments on three protein subcellular localization data sets show that semi-supervised AAMMs: (i) can effectively exploit unlabeled data; (ii) are more accurate than both the MMs and the EM based semi-supervised MMs; and (iii) are comparable in performance, and in some cases outperform, the co-training based semi-supervised MMs. Cornelia Caragea, Doina Caragea, Adrian Silvescu, Vasant G. Honavar |
BMC Bioinform. | 2 |
| 2009 | Bi-relational Network Analysis Using a Fast Random Walk with RestartabstractIdentification of nodes relevant to a given node in a relational network is a basic problem in network analysis with great practical importance. Most existing network analysis algorithms utilize one single relation to define relevancy among nodes. However, in real world applications multiple relationships exist between nodes in a network. Therefore, network analysis algorithms that can make use of more than one relation to identify the relevance set for a node are needed. In this paper, we show how the Random Walk with Restart (RWR) approach can be used to study relevancy in a bi-relational network from the bibliographic domain, and show that making use of two relations results in better results as compared to approaches that use a single relation. As relational networks can be very large, we also propose a fast implementation for RWR by adapting an existing Iterative Aggregation and Disaggregation (IAD) approach. The IAD-based RWR exploits the block-wise structure of real world networks. Experimental results show significant increase in running time for the IAD-based RWR compared to the traditional power method based RWR. Doina Caragea, William H. Hsu |
ICDM | 2 |
| 2009 | Learning Link-Based Classifiers from Ontology-Extended Textual DataabstractReal-world data mining applications call for effective strategies for learning predictive models from richly structured relational data. In this paper, we address the problem of learning classifiers from structured relational data that are annotated with relevant meta data. Specifically, we show how to learn classifiers at different levels of abstraction in a relational setting, where the structured relational data are organized in an abstraction hierarchy that describes the semantics of the content of the data. We show how to cope with some of the challenges presented by partial specification in the case of structured data, that unavoidably results from choosing a particular level of abstraction. Our solution to partial specification is based on a statistical method, called shrinkage. We present results of experiments in the case of learning link-based Naive Bayes classifiers on a text classification task that (i) demonstrate that the choice of the level of abstraction can impact the performance of the resulting link-based classifiers and (ii) examine the effect of partially specified data. Cornelia Caragea, Doina Caragea, Vasant G. Honavar |
ICTAI | 2 |
| 2009 | Towards Bridging the Web and the Semantic WebabstractThe World Wide Web (WWW) has provided us with a plethora of information. However, given its unstructured format, this information is useful mainly to humans and cannot be effectively interpreted by machines. The Semantic Web provides information in computer understandable structures (e.g., RDF), but the amount of information on the Semantic Web is limited compared to the amount of information available on the Web. The problem of generating a bridge between the Web and Semantic Web has recently gained a lot of attention. In this paper, we propose a Concept Extractor and Relationship Identifier (CE-RI) system, which acts as a bridge between Web and Semantic Web by providing a “semantic” way of presenting the search results to the user. The Concept Extractor (CE) component of our system makes use of the power of existing search engines coupled with the elegance of PageRank to extract high quality concepts related to the given query. The Relationship Identifier (RI) component finds relationships between the extracted concepts and the given query and presents them to the user in the form of a graph. It also stores the generated results formally, in the form of RDF triples, to facilitate better inferences as compared to traditional search engines. We evaluate our system by comparing its components CE and RI with other similar ”state of the art” concept detection and relationship identification systems, respectively. The results produced by our system are either similar or better than those generated by other systems. Swarnim Kulkarni, Doina Caragea |
Web Intelligence | 2 |
| 2008 | Exploring Alternative Splicing Features Using Support Vector MachinesabstractAlternative splicing is a mechanism for generating different gene transcripts (called isoforms) from the same genomic sequence. Finding alternative splicing events experimentally is both expensive and time consuming. Computational methods, in general, and machine learning algorithms,in particular, can be used to complement experimental methods in the process of identifying alternative splicing events. In this paper, we explore the predictive power of a rich set of features that have been experimentally shown to affect alternative splicing. We use these features to build support vector machine (SVM) classifiers for distinguishing between alternatively spliced exons and constitutive exons.Our results show that simple linear SVM classifiers built from a rich set of features give results comparable to those of more sophisticated SVM classifiers that use more basic sequence features. Furthermore, we use feature selection methods to identify computationally the most informative features for the prediction problem considered. Doina Caragea, Susan J. Brown |
BIBM | 2 |
| 2008 | Learning Classifiers from Large Databases Using Statistical QueriesabstractWe describe an approach to learning predictive models from large databases in settings where direct access to data is not available because of massive size of data, access restrictions, or bandwidth requirements. We outline some techniques for minimizing the number of statistical queries needed; and for efficiently coping with missing values in the data. We provide open source implementation of the decision tree and naive Bayes algorithms to demonstrate the feasibility of the proposed approach. Neeraj Koul, Cornelia Caragea, Vasant G. Honavar, Vikas Bahirwani, Doina Caragea |
Web Intelligence | 5 |
| 2007 | Structural Prediction of Protein-Protein Interactions in Saccharomyces cerevisiaeabstractProtein-protein interactions (PPI) refer to the associations between proteins and the study of these associations. Several approaches have been used to address the problem of predicting PPI. Some of them are based on biological features extracted from a protein sequence (such as, amino acid composition, GO terms, etc.); others use relational and structural features extracted from the PPI network, which can be represented as a graph. Our approach falls in the second category. We adapt a general approach to graph feature extraction that has previously been applied to collaborative recommendation of friends in social networks. Several structural features are identified based on the PPI graph and used to learn classifiers for predicting new interactions. Two datasets containing Saccharomyces cerevisiae PPI are used to test the proposed approach. Both these datasets were assembled from the Database of Interacting Proteins (DIP). We assembled the first data set directly from DIP in April 2006, while the second data set has been used in previous studies, thus making it easy to compare our approach with previous approaches. Several classifiers are trained using the structural features extracted from the interactions graph. The results show good performance (accuracy, sensitivity and specificity), proving that the structural features are highly predictive with respect to PPI. Martin S. R. Paradesi, Doina Caragea, William H. Hsu |
BIBE | 2 |
| 2006 | Learning Classifiers from Distributed, Ontology-Extended Data Sources
Doina Caragea, Jun Zhang 0002, Jyotishman Pathak, Vasant G. Honavar |
DaWaK | 1 |
| 2006 | On the Semantics of Linking and Importing in Modular Ontologies
Jie Bao 0001, Doina Caragea, Vasant G. Honavar |
ISWC | 2 |
| 2006 | Package-Based Description Logics - Preliminary Results
Jie Bao 0001, Doina Caragea, Vasant G. Honavar |
ISWC | 2 |
| 2006 | A Tableau-Based Federated Reasoning Algorithm for Modular OntologiesabstractMany real world applications of ontologies call for reasoning with modular ontologies. We describe a tableau-based reasoning algorithm based on package-based description logics (P-DL), an modular ontology language that extends description logics. Unlike classical approaches that assume a single centralized, consistent ontology, the proposed algorithm adopts a federated approach to reasoning with modular ontologies wherein each ontology module has associated with it, a local reasoner. The local reasoners communicate with each other as needed in an asynchronous fashion. Hence, the proposed approach offers an attractive approach to reasoning with multiple, autonomously developed ontology modules, in settings where it is neither possible nor desirable to integrate all involved modules into a single centralized ontology Jie Bao 0001, Doina Caragea, Vasant G. Honavar |
Web Intelligence | 2 |
| 2005 | Learning Support Vector Machines from Distributed Data Sources
Cornelia Caragea, Doina Caragea, Vasant G. Honavar |
AAAI | 2 |
| 2005 | Algorithms and Software for Collaborative Discovery from Autonomous, Semantically Heterogeneous, Distributed Information Sources
Doina Caragea, Jun Zhang 0002, Jie Bao 0001, Jyotishman Pathak, Vasant G. Honavar |
ALT | 1 |
| 2005 | Algorithms and Software for Collaborative Discovery from Autonomous, Semantically Heterogeneous, Distributed Information Sources
Doina Caragea, Jun Zhang 0002, Jie Bao 0001, Jyotishman Pathak, Vasant G. Honavar |
Discovery Science | 1 |
| 2005 | Learning Ontology-Aware Classifiers
Jun Zhang 0002, Doina Caragea, Vasant G. Honavar |
Discovery Science | 2 |
| 2003 | Towards Simple, Easy-to-Understand, yet Accurate ClassifiersabstractWe design a method for weighting linear support vector machine classifiers or random hyperplanes, to obtain classifiers whose accuracy is comparable to the accuracy of a nonlinear support vector machine classifier, and whose results can be readily visualized. We conduct a simulation study to examine how our weighted linear classifiers behave in the presence of known structure. The results show that the weighted linear classifiers might perform well compared to the nonlinear support vector machine classifiers, while they are more readily interpretable than the nonlinear classifiers. Doina Caragea, Dianne Cook, Vasant G. Honavar |
ICDM | 1 |
| 2001 | Gaining insights into support vector machine pattern classifiers using projection-based tour methodsabstractThis paper discusses visual methods that can be used to understand and interpret the results of classification using support vector machines (SVM) on data with continuous real-valued variables. SVM induction algorithms build pattern classifiers by identifying a maximal margin separating hyperplane from training examples in high dimensional pattern spaces or spaces induced by suitable nonlinear kernel transformations over pattern spaces. SVM have been demonstrated to be quite effective in a number of practical pattern classification tasks. Since the separating hyperplane is defined in terms of more than two variables it is necessary to use visual techniques that can navigate the viewer through high-dimensional spaces. We demonstrate the use of projection-based tour methods to gain useful insights into SVM classifiers with linear kernels on 8-dimensional data. Doina Caragea, Dianne Cook, Vasant G. Honavar |
KDD | 1 |