Satya Sanket Sahoo

dblp:62/2107 · also Satya S. Sahoo · DBLP profile ↗
← Back
28ranked-venue papers
10as first author
6since 2021 · last 2024
0000-0001-9190-4256ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 25 · 9 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-authorSecurity and privacy · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2024 A 360° View for Large Language Models: Early Detection of Amblyopia in Children Using Multi-view Eye Movement Recordings
Dipak P. Upadhyaya, Aasef G. Shaikh, Gokce Busra Cakir, Katrina Prantzalos, Pedram Golnari, Fatema F. Ghasia, Satya Sanket Sahoo
AIME (2)7
2024 Large language models for biomedicine: foundations, opportunities, challenges, and best practices
abstract
OBJECTIVES: Generative large language models (LLMs) are a subset of transformers-based neural network architecture models. LLMs have successfully leveraged a combination of an increased number of parameters, improvements in computational efficiency, and large pre-training datasets to perform a wide spectrum of natural language processing (NLP) tasks. Using a few examples (few-shot) or no examples (zero-shot) for prompt-tuning has enabled LLMs to achieve state-of-the-art performance in a broad range of NLP applications. This article by the American Medical Informatics Association (AMIA) NLP Working Group characterizes the opportunities, challenges, and best practices for our community to leverage and advance the integration of LLMs in downstream NLP applications effectively. This can be accomplished through a variety of approaches, including augmented prompting, instruction prompt tuning, and reinforcement learning from human feedback (RLHF). TARGET AUDIENCE: Our focus is on making LLMs accessible to the broader biomedical informatics community, including clinicians and researchers who may be unfamiliar with NLP. Additionally, NLP practitioners may gain insight from the described best practices. SCOPE: We focus on 3 broad categories of NLP tasks, namely natural language understanding, natural language inferencing, and natural language generation. We review the emerging trends in prompt tuning, instruction fine-tuning, and evaluation metrics used for LLMs while drawing attention to several issues that impact biomedical NLP applications, including falsehoods in generated text (confabulation/hallucinations), toxicity, and dataset contamination leading to overfitting. We also review potential approaches to address some of these current challenges in LLMs, such as chain of thought prompting, and the phenomena of emergent capabilities observed in LLMs that can be leveraged to address complex NLP challenge in biomedical applications.
Satya Sanket Sahoo, Joseph M. Plasek, Hua Xu 0001, Özlem Uzuner, Trevor Cohen, Meliha Yetisgen, Stéphane M. Meystre, Yanshan Wang
J. Am. Medical Informatics Assoc.1
2022 Enabling Scientific Reproducibility through FAIR Data Management: An ontology-driven deep learning approach in the NeuroBridge Project
Yue Wang 0035, José Luis Ambite, Abhishek M. Appaji, Howard Lander, Arcot Rajasekar, Jessica A. Turner, Matthew D. Turner, Lei Wang 0032, Satya Sanket Sahoo
AMIA11
2021 Parkinson and movement disorders ontology for clinically-oriented and clinicians-driven data mining of multi-center cohorts in Parkinson's disease
Massimo Marano, Raj Aurora, James Boyd, Satya Sanket Sahoo
AMIA5
2021 Epilepsy-Connect: An Integrated Knowledgebase for Characterizing Alterations in Consciousness State of Pharmacoresistant Epilepsy Patients
Katrina Prantzalos, Jianzhe Zhang, Nassim Shafiabadi, Guadalupe Fernandez-BacaVaca, Satya Sanket Sahoo
AMIA5
2021 Characterizing Brain Network Dynamics using Persistent Homology in Patients with Refractory Epilepsy
Jianzhe Zhang, Roland Bauman, Nassim Shafiabadi, Nick Gurski, Guadalupe Fernandez-BacaVaca, Satya Sanket Sahoo
AMIA6
2020 NeuroIntegrative Connectivity (NIC) Informatics Tool for Brain Functional Connectivity Network Analysis in Cohort Studies
Satya Sanket Sahoo, Arthur L. Gershon, Nassim Shafiabadi, Curtis Tatsuoka, Samden D. Lhatoo, Guadalupe Fernandez-BacaVaca
AMIA1
2019 Enhancing Multi-Center Patient Cohort Studies in the Managing Epilepsy Well (MEW) Network: Integrated Data Integration and Statistical Analysis
Xinting Hong, Hasina Momotaz, Kristin Cassidy, Martha Sajatovic, Satya Sanket Sahoo
AMIA6
2017 A Flexible Computational Neuroinformatics Workflow for Computing Functional Networks in Epilepsy Neurological Disorder
Arthur L. Gershon, Bilal Zonjy, Curtis Tatsuoka, Satya Sanket Sahoo
AMIA5
2017 ProvCaRe Semantic Provenance Knowledgebase: Evaluating Scientific Reproducibility of Research Studies
Joshua Valdez, Matthew Kim, Michael Rueschman, Vimig Socrates, Satya Sanket Sahoo
AMIA5
2016 Scientific Reproducibility in Biomedical Research: Provenance Metadata Ontology for Semantic Annotation of Study Description
Satya Sanket Sahoo, Joshua Valdez, Michael Rueschman
AMIA1
2015 Insight: Semantic provenance and analysis platform for multi-center neurology healthcare research
abstract
Insight is a Semantic Web technology-based platform to support large-scale secondary analysis of healthcare data for neurology clinical research. Insight features the novel use of: (1) provenance metadata, which describes the history or origin of patient data, in clinical research analysis, and (2) support for patient cohort queries across multiple institutions conducting research in epilepsy, which is the one of the most common neurological disorders affecting 50 million persons worldwide. Insight is being developed as a healthcare informatics infrastructure to support a national network of eight epilepsy research centers across the U.S. funded by the U.S. Centers for Disease Control and Prevention (CDC). This paper describes the use of the World Wide Web Consortium (W3C) PROV recommendation for provenance metadata that allows researchers to create patient cohorts based on the provenance of the research studies. In addition, the paper describes the use of descriptive logic-based OWL2 epilepsy ontology for cohort queries with “expansion of query expression” using ontology reasoning. Finally, the evaluation results for the data integration and query performance are described using data from three research studies with 180 epilepsy patients. The experiment results demonstrate that Insight is a scalable approach to use Semantic provenance metadata for context-based data analysis in healthcare informatics.
Priya Ramesh, Annan Wei, Elisabeth Welter, Yvan Bamps, Shelley Stoll, Ashley Bukach, Martha Sajatovic, Satya Sanket Sahoo
BIBM8
2014 MEDCIS: Multi-Modality Epilepsy Data Capture and Integration System
Guo-Qiang Zhang 0001, Licong Cui, Samden D. Lhatoo, Satya Sanket Sahoo
AMIA4
2014 Domain Ontology As Conceptual Model for Big Data Management: Application in Biomedical Informatics
Catherine P. Jayapandian, Aman Dabir, Samden D. Lhatoo, Guo-Qiang Zhang 0001, Satya Sanket Sahoo
ER6
2014 Heart beats in the cloud: distributed analysis of electrophysiological 'Big Data' using cloud computing for epilepsy clinical research
abstract
OBJECTIVE: The rapidly growing volume of multimodal electrophysiological signal data is playing a critical role in patient care and clinical research across multiple disease domains, such as epilepsy and sleep medicine. To facilitate secondary use of these data, there is an urgent need to develop novel algorithms and informatics approaches using new cloud computing technologies as well as ontologies for collaborative multicenter studies. MATERIALS AND METHODS: We present the Cloudwave platform, which (a) defines parallelized algorithms for computing cardiac measures using the MapReduce parallel programming framework, (b) supports real-time interaction with large volumes of electrophysiological signals, and (c) features signal visualization and querying functionalities using an ontology-driven web-based interface. Cloudwave is currently used in the multicenter National Institute of Neurological Diseases and Stroke (NINDS)-funded Prevention and Risk Identification of SUDEP (sudden unexplained death in epilepsy) Mortality (PRISM) project to identify risk factors for sudden death in epilepsy. RESULTS: Comparative evaluations of Cloudwave with traditional desktop approaches to compute cardiac measures (eg, QRS complexes, RR intervals, and instantaneous heart rate) on epilepsy patient data show one order of magnitude improvement for single-channel ECG data and 20 times improvement for four-channel ECG data. This enables Cloudwave to support real-time user interaction with signal data, which is semantically annotated with a novel epilepsy and seizure ontology. DISCUSSION: Data privacy is a critical issue in using cloud infrastructure, and cloud platforms, such as Amazon Web Services, offer features to support Health Insurance Portability and Accountability Act standards. CONCLUSION: The Cloudwave platform is a new approach to leverage of large-scale electrophysiological data for advancing multicenter clinical research.
Satya Sanket Sahoo, Catherine P. Jayapandian, Farhad Kaffashi, Stephanie Chung, Alireza Bozorgi, Kenneth A. Loparo, Samden D. Lhatoo, Guo-Qiang Zhang 0001
J. Am. Medical Informatics Assoc.1
2014 Epilepsy and seizure ontology: towards an epilepsy informatics infrastructure for clinical research and patient care
abstract
OBJECTIVE: Epilepsy encompasses an extensive array of clinical and research subdomains, many of which emphasize multi-modal physiological measurements such as electroencephalography and neuroimaging. The integration of structured, unstructured, and signal data into a coherent structure for patient care as well as clinical research requires an effective informatics infrastructure that is underpinned by a formal domain ontology. METHODS: We have developed an epilepsy and seizure ontology (EpSO) using a four-dimensional epilepsy classification system that integrates the latest International League Against Epilepsy terminology recommendations and National Institute of Neurological Disorders and Stroke (NINDS) common data elements. It imports concepts from existing ontologies, including the Neural ElectroMagnetic Ontologies, and uses formal concept analysis to create a taxonomy of epilepsy syndromes based on their seizure semiology and anatomical location. RESULTS: EpSO is used in a suite of informatics tools for (a) patient data entry, (b) epilepsy focused clinical free text processing, and (c) patient cohort identification as part of the multi-center NINDS-funded study on sudden unexpected death in epilepsy. EpSO is available for download at http://prism.case.edu/prism/index.php/EpilepsyOntology. DISCUSSION: An epilepsy ontology consortium is being created for community-driven extension, review, and adoption of EpSO. We are in the process of submitting EpSO to the BioPortal repository. CONCLUSIONS: EpSO plays a critical role in informatics tools for epilepsy patient care and multi-center clinical research.
Satya Sanket Sahoo, Samden D. Lhatoo, Licong Cui, Catherine P. Jayapandian, Alireza Bozorgi, Guo-Qiang Zhang 0001
J. Am. Medical Informatics Assoc.1
2014 Complex epilepsy phenotype extraction from narrative clinical discharge summaries
Licong Cui, Satya Sanket Sahoo, Samden D. Lhatoo, Prashant Rai, Alireza Bozorgi, Guo-Qiang Zhang 0001
J. Biomed. Informatics2
2013 Cloudwave: Distributed Processing of "Big Data" from Electrophysiological Recordings for Epilepsy Clinical Research Using Hadoop
Catherine P. Jayapandian, Alireza Bozorgi, Samden D. Lhatoo, Guo-Qiang Zhang 0001, Satya Sanket Sahoo
AMIA6
2012 EpiDEA: Extracting Structured Epilepsy and Seizure Information from Patient Discharge Summaries for Cohort Identification
Licong Cui, Samden D. Lhatoo, Guo-Qiang Zhang 0001, Satya Sanket Sahoo, Alireza Bozorgi
AMIA4
2012 OPIC: Ontology-driven Patient Information Capturing System for Epilepsy
Satya Sanket Sahoo, Lingyun Luo, Alireza Bozorgi, Samden D. Lhatoo, Guo-Qiang Zhang 0001
AMIA1
2012 A Distributed Semantic Web Approach for Cohort Identification
Joseph Teagno, Richard C. Kiefer, Jyotishman Pathak, G. Q. Zhang, Satya Sanket Sahoo
AMIA5
2012 An Analysis of Multi-type Relational Interactions in FMA Using Graph Motifs with Disjointness Constraints
Guo-Qiang Zhang 0001, Lingyun Luo, Chimezie Ogbuji, Cliff A. Joslyn, José L. V. Mejino Jr., Satya Sanket Sahoo
AMIA6
2011 A unified framework for managing provenance information in translational research
abstract
BACKGROUND: A critical aspect of the NIH Translational Research roadmap, which seeks to accelerate the delivery of "bench-side" discoveries to patient's "bedside," is the management of the provenance metadata that keeps track of the origin and history of data resources as they traverse the path from the bench to the bedside and back. A comprehensive provenance framework is essential for researchers to verify the quality of data, reproduce scientific results published in peer-reviewed literature, validate scientific process, and associate trust value with data and results. Traditional approaches to provenance management have focused on only partial sections of the translational research life cycle and they do not incorporate "domain semantics", which is essential to support domain-specific querying and analysis by scientists. RESULTS: We identify a common set of challenges in managing provenance information across the pre-publication and post-publication phases of data in the translational research lifecycle. We define the semantic provenance framework (SPF), underpinned by the Provenir upper-level provenance ontology, to address these challenges in the four stages of provenance metadata:(a) Provenance collection - during data generation(b) Provenance representation - to support interoperability, reasoning, and incorporate domain semantics(c) Provenance storage and propagation - to allow efficient storage and seamless propagation of provenance as the data is transferred across applications(d) Provenance query - to support queries with increasing complexity over large data size and also support knowledge discovery applicationsWe apply the SPF to two exemplar translational research projects, namely the Semantic Problem Solving Environment for Trypanosoma cruzi (T.cruzi SPSE) and the Biomedical Knowledge Repository (BKR) project, to demonstrate its effectiveness. CONCLUSIONS: The SPF provides a unified framework to effectively manage provenance of translational research data during pre and post-publication phases. This framework is underpinned by an upper-level provenance ontology called Provenir that is extended to create domain-specific provenance ontologies to facilitate provenance interoperability, seamless propagation of provenance, automated querying, and analysis.
Satya Sanket Sahoo, Vinh Nguyen 0002, Olivier Bodenreider, Priti Parikh, Todd Minning, Amit P. Sheth
BMC Bioinform.1
2010 Provenance Context Entity (PaCE): Scalable Provenance Tracking for Scientific RDF Data
Satya Sanket Sahoo, Olivier Bodenreider, Pascal Hitzler, Amit P. Sheth, Krishnaprasad Thirunarayan
SSDBM1
2008 Capturing Workflow Event Data for Monitoring, Performance Analysis, and Management of Scientific Workflows
abstract
To effectively support real-time monitoring and performance analysis of scientific workflow execution, varying levels of event data must be captured and made available to interested parties. This paper discusses the creation of an ontology-aware workflow monitoring system for use in the Trident system which utilizes a distributed publish/subscribe event model. The implementation of the publish/subscribe system is discussed and performance results are presented.
Matthew D. Valerio, Satya Sanket Sahoo, Roger S. Barga, Jared Jackson
eScience2
2008 An ontology-driven semantic mashup of gene and biological pathway information: Application to the domain of nicotine dependence
Satya Sanket Sahoo, Olivier Bodenreider, Joni L. Rutter, Karen J. Skinner, Amit P. Sheth
J. Biomed. Informatics1
2006 Knowledge modeling and its application in life sciences: a tale of two ontologies
abstract
High throughput glycoproteomics, similar to genomics and proteomics, involves extremely large volumes of distributed, heterogeneous data as a basis for identification and quantification of a structurally diverse collection of biomolecules. The ability to share, compare, query for and most critically correlate datasets using the native biological relationships are some of the challenges being faced by glycobiology researchers. As a solution for these challenges, we are building a semantic structure, using a suite of ontologies, which supports management of data and information at each step of the experimental lifecycle. This framework will enable researchers to leverage the large scale of glycoproteomics data to their benefit.In this paper, we focus on the design of these biological ontology schemas with an emphasis on relationships between biological concepts, on the use of novel approaches to populate these complex ontologies including integrating extremely large datasets ( 500MB) as part of the instance base and on the evaluation of ontologies using OntoQA [38] metrics. The application of these ontologies in providing informatics solutions, for high throughput glycoproteomics experimental domain, is also discussed. We present our experience as a use case of developing two ontologies in one domain, to be part of a set of use cases, which are used in the development of an emergent framework for building and deploying biological ontologies.
Satya Sanket Sahoo, Christopher Thomas 0001, Amit P. Sheth, William S. York, Samir Tartir
WWW1
2005 Template Based Semantic Similarity for Security Applications
Boanerges Aleman-Meza, Christian Halaschek-Wiener, Satya Sanket Sahoo, Amit P. Sheth, Ismailcem Budak Arpinar
ISI3