Ramakanth Kavuluru

dblp:17/265 · DBLP profile ↗
← Back
44ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0003-1238-9378ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 27 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Security and privacy · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 LitTx: A New Treatment Relation Extraction Dataset
abstract
for identifying treatment relationships discussed in literature given the lack of such datasets in the recent past. Besides confirmed or implied positive relations, we also introduce a new "conditional treatment" relation type where hedging or a potential relationship is indicated. Our baseline RE models with this new dataset demonstrate promising results, while also revealing clear areas for improvement. To foster innovation and ensure replicability in the biomedical RE community, we release our dataset, code, and annotation guidelines publicly: https://github.com/bionlproc/LitTx_dataset.
Md Sultan Al Nahian, Li Hao Richie Xu, Rani Chikkanna, Ramakanth Kavuluru
LREC5
2025 How Important is Domain-Specific Language Model Pretraining and Instruction Finetuning for Biomedical Relation Extraction?
Aviv Brokman, Ramakanth Kavuluru
NLDB (1)2
2025 Comparison of Pipelines, Seq2seq Models, and LLMs for Rare Disease Information Extraction
Xuguang Ai, Ramakanth Kavuluru
NLDB (1)4
2024 Revisiting Document-Level Relation Extraction with Context-Guided Link Prediction
abstract
Document-level relation extraction (DocRE) poses the challenge of identifying relationships between entities within a document. Existing approaches rely on logical reasoning or contextual cues from entities. This paper reframes document-level RE as link prediction over a Knowledge Graph (KG) with distinct benefits: 1) Our approach amalgamates entity context and document-derived logical reasoning, enhancing link prediction quality. 2) Predicted links between entities offer interpretability, elucidating employed reasoning. We evaluate our approach on benchmark datasets - DocRED, ReDocRED, and DWIE. The results indicate that our proposed method outperforms the state-of-the-art models and suggests that incorporating context-based Knowledge Graph link prediction techniques can enhance the performance of document-level relation extraction models.
Raghava Mutharaju, Ramakanth Kavuluru
AAAI3
2023 An open natural language processing (NLP) framework for EHR-based clinical research: a case demonstration using the National COVID Cohort Collaborative (N3C)
abstract
Despite recent methodology advancements in clinical natural language processing (NLP), the adoption of clinical NLP models within the translational research community remains hindered by process heterogeneity and human factor variations. Concurrently, these factors also dramatically increase the difficulty in developing NLP models in multi-site settings, which is necessary for algorithm robustness and generalizability. Here, we reported on our experience developing an NLP solution for Coronavirus Disease 2019 (COVID-19) signs and symptom extraction in an open NLP framework from a subset of sites participating in the National COVID Cohort (N3C). We then empirically highlight the benefits of multi-site data for both symbolic and statistical methods, as well as highlight the need for federated annotation and evaluation to resolve several pitfalls encountered in the course of these efforts.
Sijia Liu 0002, Andrew Wen, Liwei Wang 0010, Sunyang Fu, Robert T. Miller, Andrew E. Williams, Daniel R. Harris, Ramakanth Kavuluru, Noor Abu-El-Rub, Dalton Schutte, Rui Zhang 0028, Masoud Rouhizadeh, John D. Osborne, Yongqun He, Umit Topaloglu, Stephanie S. Hong, Joel H. Saltz, Thomas Schaffter, Emily R. Pfaff, Christopher G. Chute, Tim Duong, Melissa A. Haendel, Rafael Fuentes, Peter Szolovits, Hua Xu 0001
J. Am. Medical Informatics Assoc.9
2023 A Comparative Effectiveness Study on Opioid Use Disorder Prediction Using Artificial Intelligence and Existing Risk Models
abstract
Opioid use disorder (OUD) is a leading cause of death in the United States placing a tremendous burden on patients, their families, and health care systems. Artificial intelligence (AI) can be harnessed with available healthcare data to produce automated OUD prediction tools. In this retrospective study, we developed AI based models for OUD prediction and showed that AI can predict OUD more effectively than existing clinical tools including the unweighted opioid risk tool (ORT). Data include 474,208 patients' data over 10 years; 269,748 were females with an average age of 56.78 years. Cases are prescription opioid users with at least one diagnosis of OUD or at least one prescription for buprenorphine or methadone. Controls are prescription opioid users with no OUD diagnoses or buprenorphine or methadone prescriptions. On 100 randomly selected test sets including 47,396 patients, our proposed transformer-based AI model can predict OUD more efficiently (AUC = 0.742 ± 0.021) compared to logistic regression (AUC = 0.651 ± 0.025), random forest (AUC = 0.679 ± 0.026), xgboost (AUC = 0.690 ± 0.027), long short-term memory model (AUC = 0.706 ± 0.026), transformer (AUC = 0.725 ± 0.024), and unweighted ORT model (AUC = 0.559 ± 0.025). Our results show that embedding AI algorithms into clinical care may assist clinicians in risk stratification and management of patients receiving opioid therapy.
Sajjad Fouladvand, Jeffery C. Talbert, Linda P. Dwoskin, Heather Bush, Amy Lynn Meadows, Lars E. Peterson, Yash R. Mishra, Steven K. Roggenkamp, Ramakanth Kavuluru, Jin Chen 0004
IEEE J. Biomed. Health Informatics10
2022 DeepPhe: Natural Language Processing Tools for Cancer Research and Surveillance
Harry Hochheiser, Sean Finan, Zhou Yuan, John D. Levander, Eric B. Durbin, Isaac Hands, Ramakanth Kavuluru, Jeremy L. Warner, Guergana K. Savova
AMIA7
2022 Toward informatics-enabled preparedness for natural hazards to minimize health impacts of climate change
abstract
Natural hazards (NHs) associated with climate change have been increasing in frequency and intensity. These acute events impact humans both directly and through their effects on social and environmental determinants of health. Rather than relying on a fully reactive incident response disposition, it is crucial to ramp up preparedness initiatives for worsening case scenarios. In this perspective, we review the landscape of NH effects for human health and explore the potential of health informatics to address associated challenges, specifically from a preparedness angle. We outline important components in a health informatics agenda for hazard preparedness involving hazard-disease associations, social determinants of health, and hazard forecasting models, and call for novel methods to integrate them toward projecting healthcare needs in the wake of a hazard. We describe potential gaps and barriers in implementing these components and propose some high-level ideas to address them.
Jimmy Phuong, Naomi O. Riches, Luca Calzoni, Gora Datta, Deborah Duran, Asiyah Yu Lin, Ramesh P. Singh, Tony Solomonides, Noreen Whysel, Ramakanth Kavuluru
J. Am. Medical Informatics Assoc.10
2022 Contrastive Cross-Modal Pre-Training: A General Strategy for Small Sample Medical Imaging
abstract
A key challenge in training neural networks for a given medical imaging task is the difficulty of obtaining a sufficient number of manually labeled examples. In contrast, textual imaging reports are often readily available in medical records and contain rich but unstructured interpretations written by experts as part of standard clinical practice. We propose using these textual reports as a form of weak supervision to improve the image interpretation performance of a neural network without requiring additional manually labeled examples. We use an image-text matching task to train a feature extractor and then fine-tune it in a transfer learning setting for a supervised task using a small labeled dataset. The end result is a neural network that automatically interprets imagery without requiring textual reports during inference. We evaluate our method on three classification tasks and find consistent performance improvements, reducing the need for labeled data by 67%-98%.
Gongbo Liang, Connor Greenwell, Yu Zhang 0094, Xin Xing 0002, Ramakanth Kavuluru, Nathan Jacobs
IEEE J. Biomed. Health Informatics6
2022 ReCDroid+: Automated End-to-End Crash Reproduction from Bug Reports for Android Apps
abstract
The large demand of mobile devices creates significant concerns about the quality of mobile applications (apps). Developers heavily rely on bug reports in issue tracking systems to reproduce failures (e.g., crashes). However, the process of crash reproduction is often manually done by developers, making the resolution of bugs inefficient, especially given that bug reports are often written in natural language. To improve the productivity of developers in resolving bug reports, in this paper, we introduce a novel approach, called ReCDroid+, that can automatically reproduce crashes from bug reports for Android apps. ReCDroid+ uses a combination of natural language processing (NLP) , deep learning, and dynamic GUI exploration to synthesize event sequences with the goal of reproducing the reported crash. We have evaluated ReCDroid+ on 66 original bug reports from 37 Android apps. The results show that ReCDroid+ successfully reproduced 42 crashes (63.6% success rate) directly from the textual description of the manually reproduced bug reports. A user study involving 12 participants demonstrates that ReCDroid+ can improve the productivity of developers when resolving crash bug reports.
Yu Zhao 0010, Ting Su 0001, Yang Liu 0003, Wei Zheng 0006, Xiaoxue Wu 0001, Ramakanth Kavuluru, William G. J. Halfond, Tingting Yu 0001
ACM Trans. Softw. Eng. Methodol.6
2021 Identifying Opioid Use Disorder from Longitudinal Healthcare Data using a Multi-stream Transformer
Sajjad Fouladvand, Jeffery C. Talbert, Linda P. Dwoskin, Heather Bush, Amy Lynn Meadows, Lars E. Peterson, Steven K. Roggenkamp, Ramakanth Kavuluru, Jin Chen 0004
AMIA8
2021 Therapeutic Claims in Cannabidiol (CBD) Marketing Messages on Twitter
abstract
Although the U.S. FDA has only approved exactly one cannabidiol (CBD) drug product (specifically to treat seizures), CBD products are proliferating rapidly through different modes of usage including food products, cosmetics, vaping pods, and supplements (typically, oils). Despite the FDA clearly warning consumers about unproven health claims made by manufacturers selling CBD products over the counter, the CBD market share was nearly 3 billion USD in 2020 and is expected to top 55 billion USD in 2028. In this context, it is important to assess the presence of health claims being made on social media, especially claims that are part of marketing messages. To this end, we collected over two million English tweets discussing CBD themes. We created a hand-labeled dataset and built machine learned classifiers to identify marketing tweets from regular tweets that may be generated by consumers. The best classifier achieved 85% precision, 83% recall, and 84% F-score. Our analyses showed that pain, anxiety disorders, sleep disorders, and stress are the four main therapeutic claims made constituting 31.67%, 27.11%, 13.77%, and 10.37% of all medical claims made on Twitter, respectively. Also, more than 93% of advertised CBD products are edibles or oil/tinctures. Our effort is the first to demonstrate the feasibility of surveillance of marketing claims for CBD products. We believe this could pave way for more explorations into this indispensable task in the current landscape of social media driven health (mis)information and communication.
Mohammad Soleymanpour, Sofia Saderholm, Ramakanth Kavuluru
BIBM3
2021 Attention-Gated Graph Convolutions for Extracting Drug Interaction Information from Drug Labels
abstract
Preventable adverse events as a result of medical errors present a growing concern in the healthcare system. As drug-drug interactions (DDIs) may lead to preventable adverse events, being able to extract DDIs from drug labels into a machine-processable form is an important step toward effective dissemination of drug safety information. Herein, we tackle the problem of jointly extracting mentions of drugs and their interactions, including interactionoutcome, from drug labels. Our deep learning approach entails composing various intermediate representations, including graph-based context derived using graph convolutions (GCs) with a novel attention-based gating mechanism (holistically called GCA), which are combined in meaningful ways to predict on all subtasks jointly. Our model is trained and evaluated on the 2018 TAC DDI corpus. Our GCA model in conjunction with transfer learning performs at 39.20% F1 and 26.09% F1 on entity recognition (ER) and relation extraction (RE), respectively, on the first official test set and at 45.30% F1 and 27.87% F1 on ER and RE, respectively, on the second official test set. These updated results lead to improvements over our prior best by up to 6 absolute F1 points. After controlling for available training data, the proposed model exhibits state-of-the-art performance for this task.
Tung Tran 0001, Ramakanth Kavuluru, Halil Kilicoglu
ACM Trans. Comput. Heal.2
2021 Improved biomedical word embeddings in the transformer era
Jiho Noh, Ramakanth Kavuluru
J. Biomed. Informatics2
2019 Non-Negative Matrix Factorization for Drug Repositioning: Experiments with the repoDB Dataset
Mehmet G. Bakal, Halil Kilicoglu, Ramakanth Kavuluru
AMIA3
2019 Clinical Text Mining in Mental Health
Jessica D. Tenenbaum, Ramakanth Kavuluru, Thomas H. McCoy, Özlem Uzuner, Sumithra Velupillai
AMIA2
2019 Knowledge-aware Assessment of Severity of Suicide Risk for Early Intervention
abstract
Mental health illness such as depression is a significant risk factor for suicide ideation, behaviors, and attempts. A report by Substance Abuse and Mental Health Services Administration (SAMHSA) shows that 80% of the patients suffering from Borderline Personality Disorder (BPD) have suicidal behavior, 5-10% of whom commit suicide. While multiple initiatives have been developed and implemented for suicide prevention, a key challenge has been the social stigma associated with mental disorders, which deters patients from seeking help or sharing their experiences directly with others including clinicians. This is particularly true for teenagers and younger adults where suicide is the second highest cause of death in the US. Prior research involving surveys and questionnaires (e.g. PHQ-9) for suicide risk prediction failed to provide a quantitative assessment of risk that informed timely clinical decision-making for intervention. Our interdisciplinary study concerns the use of Reddit as an unobtrusive data source for gleaning information about suicidal tendencies and other related mental health conditions afflicting depressed users. We provide details of our learning framework that incorporates domain-specific knowledge to predict the severity of suicide risk for an individual. Our approach involves developing a suicide risk severity lexicon using medical knowledge bases and suicide ontology to detect cues relevant to suicidal thoughts and actions. We also use language modeling, medical entity recognition and normalization and negation detection to create a dataset of 2181 redditors that have discussed or implied suicidal ideation, behavior, or attempt. Given the importance of clinical knowledge, our gold standard dataset of 500 redditors (out of 2181) was developed by four practicing psychiatrists following the guidelines outlined in Columbia Suicide Severity Rating Scale (C-SSRS), with the pairwise annotator agreement of 0.79 and group-wise agreement of 0.73. Compared to the existing four-label classification scheme (no risk, low risk, moderate risk, and high risk), our proposed C-SSRS-based 5-label classification scheme distinguishes people who are supportive, from those who show different severity of suicidal tendency. Our 5-label classification scheme outperforms the state-of-the-art schemes by improving the graded recall by 4.2% and reducing the perceived risk measure by 12.5%. Convolutional neural network (CNN) provided the best performance in our scheme due to the discriminative features and use of domain-specific knowledge resources, in comparison to SVM-L that has been used in the state-of-the-art tools over similar dataset.
Manas Gaur, Amanuel Alambo, Joy Prakash Sain, Ugur Kursuncu, Krishnaprasad Thirunarayan, Ramakanth Kavuluru, Amit P. Sheth, Randy S. Welton, Jyotishman Pathak
WWW6
2019 Neural transfer learning for assigning diagnosis codes to EMRs
Anthony Rios, Ramakanth Kavuluru
Artif. Intell. Medicine2
2019 Distant supervision for treatment relation extraction by leveraging MeSH subheadings
Tung Tran 0001, Ramakanth Kavuluru
Artif. Intell. Medicine2
2019 Cross-registry neural domain adaptation to extract mutational test results from pathology reports
Anthony Rios, Eric B. Durbin, Isaac Hands, Susanne M. Arnold, Darshil Shah, Stephen M. Schwartz, Bernardo H. L. Goulart, Ramakanth Kavuluru
J. Biomed. Informatics8
2018 Few-Shot and Zero-Shot Multi-Label Learning for Structured Label Spaces
abstract
Large multi-label datasets contain labels that occur thousands of times (frequent group), those that occur only a few times (few-shot group), and labels that never appear in the training dataset (zero-shot group). Multi-label few- and zero-shot label prediction is mostly unexplored on datasets with large label spaces, especially for text classification. In this paper, we perform a fine-grained evaluation to understand how state-of-the-art methods perform on infrequent labels. Furthermore, we develop few- and zero-shot methods for multi-label text classification when there is a known structure over the label space, and evaluate them on two publicly available medical text datasets: MIMIC II and MIMIC III. For few-shot labels we achieve improvements of 6.2% and 4.8% in R@10 for MIMIC II and MIMIC III, respectively, over prior efforts; the corresponding R@10 improvements for zero-shot labels are 17.3% and 19%.
Anthony Rios, Ramakanth Kavuluru
EMNLP2
2018 Document Retrieval for Biomedical Question Answering with Neural Sentence Matching
abstract
Document retrieval (DR) forms an important component in end-to-end question-answering (QA) systems where particular answers are sought for well-formed questions. DR in the QA scenario is also useful by itself even without a more involved natural language processing component to extract exact answers from the retrieved documents. This latter step may simply be done by humans like in traditional search engines granted the retrieved documents contain the answer. In this paper, we take advantage of datasets made available through the BioASQ end-to-end QA shared task series and build an effective biomedical DR system that relies on relevant answer snippets in the BioASQ training datasets. At the core of our approach is a question-answer sentence matching neural network that learns a measure of relevance of a sentence to an input question in the form of a matching score. In addition to this matching score feature, we also exploit two auxiliary features for scoring document relevance: the name of the journal in which a document is published and the presence/absence of semantic relations (subject-predicate-object triples) in a candidate answer sentence connecting entities mentioned in the question. We rerank our baseline sequential dependence model scores using these three additional features weighted via adaptive random research and other learning-to-rank methods. Our full system placed 2nd in the final batch of Phase A (DR) of task B (QA) in BioASQ 2018. Our ablation experiments highlight the significance of the neural matching network component in the full system.
Jiho Noh, Ramakanth Kavuluru
ICMLA2
2018 EMR Coding with Semi-Parametric Multi-Head Matching Networks
abstract
Coding EMRs with diagnosis and procedure codes is an indispensable task for billing, secondary data analyses, and monitoring health trends. Both speed and accuracy of coding are critical. While coding errors could lead to more patient-side financial burden and mis-interpretation of a patient's well-being, timely coding is also needed to avoid backlogs and additional costs for the healthcare facility. In this paper, we present a new neural network architecture that combines ideas from few-shot learning matching networks, multi-label loss functions, and convolutional neural networks for text classification to significantly outperform other state-of-the-art models. Our evaluations are conducted using a well known deidentified EMR dataset (MIMIC) with a variety of multi-label performance measures.
Anthony Rios, Ramakanth Kavuluru
NAACL-HLT2
2018 Generalizing biomedical relation classification with neural adversarial domain adaptation
abstract
Motivation: Creating large datasets for biomedical relation classification can be prohibitively expensive. While some datasets have been curated to extract protein-protein and drug-drug interactions (PPIs and DDIs) from text, we are also interested in other interactions including gene-disease and chemical-protein connections. Also, many biomedical researchers have begun to explore ternary relationships. Even when annotated data are available, many datasets used for relation classification are inherently biased. For example, issues such as sample selection bias typically prevent models from generalizing in the wild. To address the problem of cross-corpora generalization, we present a novel adversarial learning algorithm for unsupervised domain adaptation tasks where no labeled data are available in the target domain. Instead, our method takes advantage of unlabeled data to improve biased classifiers through learning domain-invariant features via an adversarial process. Finally, our method is built upon recent advances in neural network (NN) methods. Results: We experiment by extracting PPIs and DDIs from text. In our experiments, we show domain invariant features can be learned in NNs such that classifiers trained for one interaction type (protein-protein) can be re-purposed to others (drug-drug). We also show that our method can adapt to different source and target pairs of PPI datasets. Compared to prior convolutional and recurrent NN-based relation classification methods without domain adaptation, we achieve improvements as high as 30% in F1-score. Likewise, we show improvements over state-of-the-art adversarial methods. Availability and implementation: Experimental code is available at https://github.com/bionlproc/adversarial-relation-classification. Supplementary information: Supplementary data are available at Bioinformatics online.
Anthony Rios, Ramakanth Kavuluru, Zhiyong Lu
Bioinform.2
2018 Data and systems for medication-related text classification and concept normalization from Twitter: insights from the Social Media Mining for Health (SMM4H)-2017 shared task
abstract
Objective: We executed the Social Media Mining for Health (SMM4H) 2017 shared tasks to enable the community-driven development and large-scale evaluation of automatic text processing methods for the classification and normalization of health-related text from social media. An additional objective was to publicly release manually annotated data. Materials and Methods: We organized 3 independent subtasks: automatic classification of self-reports of 1) adverse drug reactions (ADRs) and 2) medication consumption, from medication-mentioning tweets, and 3) normalization of ADR expressions. Training data consisted of 15 717 annotated tweets for (1), 10 260 for (2), and 6650 ADR phrases and identifiers for (3); and exhibited typical properties of social-media-based health-related texts. Systems were evaluated using 9961, 7513, and 2500 instances for the 3 subtasks, respectively. We evaluated performances of classes of methods and ensembles of system combinations following the shared tasks. Results: Among 55 system runs, the best system scores for the 3 subtasks were 0.435 (ADR class F1-score) for subtask-1, 0.693 (micro-averaged F1-score over two classes) for subtask-2, and 88.5% (accuracy) for subtask-3. Ensembles of system combinations obtained best scores of 0.476, 0.702, and 88.7%, outperforming individual systems. Discussion: Among individual systems, support vector machines and convolutional neural networks showed high performance. Performance gains achieved by ensembles of system combinations suggest that such strategies may be suitable for operational systems relying on difficult text classification tasks (eg, subtask-1). Conclusions: Data imbalance and lack of context remain challenges for natural language processing of social media text. Annotated data from the shared task have been made available as reference standards for future studies (http://dx.doi.org/10.17632/rxwfb3tysd.1).
Abeed Sarker, Maksim Belousov, Jasper Friedrichs, Kai Hakala, Svetlana Kiritchenko, Farrokh Mehryary, Sifei Han, Tung Tran 0001, Anthony Rios, Ramakanth Kavuluru, Berry de Bruijn, Filip Ginter, Debanjan Mahata, Saif M. Mohammad, Goran Nenadic, Graciela Gonzalez-Hernandez
J. Am. Medical Informatics Assoc.10
2018 Exploiting semantic patterns over biomedical knowledge graphs for predicting treatment and causative relations
Gokhan Bakal, Preetham Talari, Elijah V. Kakani, Ramakanth Kavuluru
J. Biomed. Informatics4
2017 Knowledge-Based Biomedical Word Sense Disambiguation with Neural Concept Embeddings
abstract
Biomedical word sense disambiguation (WSD) is an important intermediate task in many natural language processing applications such as named entity recognition, syntactic parsing, and relation extraction. In this paper, we employ knowledge-based approaches that also exploit recent advances in neural word/concept embeddings to improve over the state-of-the-art in biomedical WSD using the public MSH WSD dataset [1] as the test set. Our methods involve weak supervision - we do not use any hand-labeled examples for WSD to build our prediction models; however, we employ an existing concept mapping program, MetaMap, to obtain our concept vectors. Over the MSH WSD dataset, our linear time (in terms of numbers of senses and words in the test instance) method achieves an accuracy of 92.24% which is a 3% improvement over the best known results [2] obtained via unsupervised means. A more expensive approach that we developed relies on a nearest neighbor framework and achieves accuracy of 94.34%, essentially cutting the error rate in half. Employing dense vector representations learned from unlabeled free text has been shown to benefit many language processing tasks recently and our efforts show that biomedical WSD is no exception to this trend. For a complex and rapidly evolving domain such as biomedicine, building labeled datasets for larger sets of ambiguous terms may be impractical. Here, we show that weak supervision that leverages recent advances in representation learning can rival supervised approaches in biomedical WSD. However, external knowledge bases (here sense inventories) play a key role in the improvements achieved.
Akm Sabbir, Antonio Jimeno-Yepes, Ramakanth Kavuluru
BIBE3
2016 Mining EMR Data to Hypothesize Causal Associations for Depressive Disorders
Orhan Abar, Ramakanth Kavuluru
AMIA2
2016 On the Predictive Potential of Graph Patterns for Biomedical Relation Extraction
Gokhan Bakal, Sergei Wallace, Ramakanth Kavuluru
AMIA3
2016 Exploratory Analysis of Marketing Vs. Non-Marketing Tweets on E-Cigarettes
Sifei Han, Ramakanth Kavuluru
AMIA2
2016 Toward automated e-cigarette surveillance: Spotting e-cigarette proponents on Twitter
Ramakanth Kavuluru, A. K. M. Sabbir
J. Biomed. Informatics1
2015 Automatic Assignment of Non-Leaf MeSH Terms to Biomedical Articles
Ramakanth Kavuluru, Anthony Rios
AMIA1
2015 An empirical evaluation of supervised learning approaches in assigning diagnosis codes to electronic medical records
Ramakanth Kavuluru, Anthony Rios
Artif. Intell. Medicine1
2015 Context-driven automatic subgraph creation for literature-based discovery
Delroy Cameron, Ramakanth Kavuluru, Thomas C. Rindflesch, Amit P. Sheth, Krishnaprasad Thirunarayan, Olivier Bodenreider
J. Biomed. Informatics2
2014 A Knowledge-Based Collaborative Clinical Case Mining Framework
Ramakanth Kavuluru, Anthony Rios, Brandon Kulengowski, Patrick McNamara
AMIA1
2014 Leveraging output term co-occurrence frequencies and latent associations in predicting medical subject headings
Ramakanth Kavuluru
Data Knowl. Eng.1
2014 Using Common Table Expressions to Build a Scalable Boolean Query Generator for Clinical Data Warehouses
abstract
We present a custom, Boolean query generator utilizing common-table expressions (CTEs) that is capable of scaling with big datasets. The generator maps user-defined Boolean queries, such as those interactively created in clinical-research and general-purpose healthcare tools, into SQL. We demonstrate the effectiveness of this generator by integrating our study into the Informatics for Integrating Biology and the Bedside (i2b2) query tool and show that it is capable of scaling. Our custom generator replaces and outperforms the default query generator found within the Clinical Research Chart cell of i2b2. In our experiments, 16 different types of i2b2 queries were identified by varying four constraints: date, frequency, exclusion criteria, and whether selected concepts occurred in the same encounter. We generated nontrivial, random Boolean queries based on these 16 types; the corresponding SQL queries produced by both generators were compared by execution times. The CTE-based solution significantly outperformed the default query generator and provided a much more consistent response time across all query types (M = 2.03, SD = 6.64 versus M = 75.82, SD = 238.88 s). Without costly hardware upgrades, we provide a scalable solution based on CTEs with very promising empirical results centered on performance gains. The evaluation methodology used for this provides a means of profiling clinical data warehouse performance.
Daniel R. Harris, Darren W. Henderson, Ramakanth Kavuluru, Arnold J. Stromberg, Todd R. Johnson
IEEE J. Biomed. Health Informatics3
2013 Phrase Based Topic Modeling for Semantic Information Processing in Biomedicine
abstract
Given that unstructured data is increasing exponentially everyday, extracting and understanding the information, themes, and relationships from large collections of documents is increasingly important to researchers in many disciplines including biomedicine. Latent Dirichlet Allocation (LDA) is an unsupervised topic modeling technique based on the "bag-of-words" assumption that has been applied extensively to unveil hidden semantic themes within large sets of textual documents. Recently, it was extended using the "bag-of-n-grams" paradigm to account for word order. In this paper, we present an alternative phrase based LDA model to move from a bag of words or n-grams paradigm to a "bag-of-key-phrases" setting by applying a key phrase extraction technique, the C-value method, to further explore latent themes. We evaluate our approach by using a phrase intrusion user study and demonstrate that our model can help LDA generate better and more interpretable topics than those generated using the bag-of-n-grams approach. Given topic models essentially are statistical tools, an important problem in topic modeling is that of visualizing and interacting with the models to understand and extract new information from a collection. To evaluate our phrase based modeling approach in this context, we incorporate it in an open source interactive topic browser. Qualitative evaluations of this browser with biomedical experts demonstrate that our approach can aid biomedical researchers gain better and faster understanding of their document collections.
Todd R. Johnson, Ramakanth Kavuluru
ICMLA (1)3
2013 Unsupervised Medical Subject Heading Assignment Using Output Label Co-occurrence Statistics and Semantic Predications
Ramakanth Kavuluru, Zhenghao He
NLDB1
2012 Improving Scalability and Performance of i2b2 Query Processing Using Common Table Expressions
Darren W. Henderson, Daniel R. Harris, Ramakanth Kavuluru, Todd R. Johnson
AMIA3
2011 Semantic Predications for Complex Information Needs in Biomedical Literature
abstract
Many complex information needs that arise in biomedical disciplines require exploring multiple documents in order to obtain information. While traditional information retrieval techniques that return a single ranked list of documents are quite common for such tasks, they may not always be adequate. The main issue is that ranked lists typically impose a significant burden on users to filter out irrelevant documents. Additionally, users must intuitively reformulate their search query when relevant documents have not been not highly ranked. Furthermore, even after interesting documents have been selected, very few mechanisms exist that enable document-to-document transitions. In this paper, we demonstrate the utility of assertions extracted from biomedical text (called semantic predications) to facilitate retrieving relevant documents for complex information needs. Our approach offers an alternative to query reformulation by establishing a framework for transitioning from one document to another. We evaluate this novel knowledge-driven approach using precision and recall metrics on the 2006 TREC Genomics Track.
Delroy Cameron, Ramakanth Kavuluru, Olivier Bodenreider, Pablo N. Mendes, Amit P. Sheth, Krishnaprasad Thirunarayan
BIBM2
2011 RASP: efficient multidimensional range query on attack-resilient encrypted databases
abstract
Range query is one of the most frequently used queries for online data analytics. Providing such a query service could be expensive for the data owner. With the development of services computing and cloud computing, it has become possible to outsource large databases to database service providers and let the providers maintain the range-query service. With outsourced services, the data owner can greatly reduce the cost in maintaining computing infrastructure and data-rich applications. However, the service provider, although honestly processing queries, may be curious about the hosted data and received queries. Most existing encryption based approaches require linear scan over the entire database, which is inappropriate for online data analytics on large databases. While a few encryption solutions are more focused on efficiency side, they are vulnerable to attackers equipped with certain prior knowledge. We propose the Random Space Encryption (RASP) approach that allows efficient range search with stronger attack resilience than existing efficiency-focused approaches. We use RASP to generate indexable auxiliary data that is resilient to prior knowledge enhanced attacks. Range queries are securely transformed to the encrypted data space and then efficiently processed with a two-stage processing algorithm. We thoroughly studied the potential attacks on the encrypted data and queries at three different levels of prior knowledge available to an attacker. Experimental results on synthetic and real datasets show that this encryption approach allows efficient processing of range queries with high resilience to attacks.
Keke Chen, Ramakanth Kavuluru, Shumin Guo
CODASPY2
2009 Characterization of 2n-periodic binary sequences with fixed 2-error or 3-error linear complexity
Ramakanth Kavuluru
Des. Codes Cryptogr.1
2008 2n-Periodic Binary Sequences with Fixed k-Error Linear Complexity for k=2 or 3
Ramakanth Kavuluru
SETA1