EDBT 2026 Demo / reviewers in the wild / expert
Karthik Soman
dblp:213/0089
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2024
0000-0002-3490-9306ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 72% Medical and health informatics · 28% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › knowledge representation in biology
biomedical knowledge graph |
1.4 | 2 | 2024 | Biomedical knowledge graph-optimized prompt generation for large language models · Bioinform. 2024 The scalable precision medicine open knowledge engine (SPOKE): a massive knowledge graph of biomedical information · Bioinform. 2023 |
Bioinformatics and computational biology
knowledge graph |
0.9 | 2 | 2024 | The scalable precision medicine open knowledge engine (SPOKE): a massive knowledge graph of biomedical information · Bioinform. 2023 Biomedical knowledge graph-optimized prompt generation for large language models · Bioinform. 2024 |
Bioinformatics and computational biology › biomedical text mining
biomedical question answering |
0.8 | 1 | 2024 | Biomedical knowledge graph-optimized prompt generation for large language models · Bioinform. 2024 |
Medical and health informatics
retrieval-augmented generation |
0.8 | 1 | 2024 | Biomedical knowledge graph-optimized prompt generation for large language models · Bioinform. 2024 |
Bioinformatics and computational biology › data integration
biomedical data integration |
0.7 | 1 | 2023 | The scalable precision medicine open knowledge engine (SPOKE): a massive knowledge graph of biomedical information · Bioinform. 2023 |
Medical and health informatics
precision medicine |
0.7 | 1 | 2023 | The scalable precision medicine open knowledge engine (SPOKE): a massive knowledge graph of biomedical information · Bioinform. 2023 |
Methods — techniques the papers use, named apart from their topics
retrieval-augmented generation · 0.8large language model · 0.8embedding-based context pruning · 0.8ontology-based integration · 0.7REST API · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Biomedical knowledge graph-optimized prompt generation for large language modelsabstractMOTIVATION: Large language models (LLMs) are being adopted at an unprecedented rate, yet still face challenges in knowledge-intensive domains such as biomedicine. Solutions such as pretraining and domain-specific fine-tuning add substantial computational overhead, requiring further domain-expertise. Here, we introduce a token-optimized and robust Knowledge Graph-based Retrieval Augmented Generation (KG-RAG) framework by leveraging a massive biomedical KG (SPOKE) with LLMs such as Llama-2-13b, GPT-3.5-Turbo, and GPT-4, to generate meaningful biomedical text rooted in established knowledge. RESULTS: Compared to the existing RAG technique for Knowledge Graphs, the proposed method utilizes minimal graph schema for context extraction and uses embedding methods for context pruning. This optimization in context extraction results in more than 50% reduction in token consumption without compromising the accuracy, making a cost-effective and robust RAG implementation on proprietary LLMs. KG-RAG consistently enhanced the performance of LLMs across diverse biomedical prompts by generating responses rooted in established knowledge, accompanied by accurate provenance and statistical evidence (if available) to substantiate the claims. Further benchmarking on human curated datasets, such as biomedical true/false and multiple-choice questions (MCQ), showed a remarkable 71% boost in the performance of the Llama-2 model on the challenging MCQ dataset, demonstrating the framework's capacity to empower open-source models with fewer parameters for domain-specific questions. Furthermore, KG-RAG enhanced the performance of proprietary GPT models, such as GPT-3.5 and GPT-4. In summary, the proposed framework combines explicit and implicit knowledge of KG and LLM in a token optimized fashion, thus enhancing the adaptability of general-purpose LLMs to tackle domain-specific questions in a cost-effective fashion. AVAILABILITY AND IMPLEMENTATION: SPOKE KG can be accessed at https://spoke.rbvi.ucsf.edu/neighborhood.html. It can also be accessed using REST-API (https://spoke.rbvi.ucsf.edu/swagger/). KG-RAG code is made available at https://github.com/BaranziniLab/KG_RAG. Biomedical benchmark datasets used in this study are made available to the research community in the same GitHub repository. Karthik Soman, Peter W. Rose, John Scotter Morris, Rabia E. Akbas, Brett Smith, Braian Peetoom, Catalina Villouta-Reyes, Gabriel Cerono, Yongmei Shi, Angela Rizk-Jackson, Sharat Israni, Charlotte A. Nelson, Sui Huang, Sergio Baranzini |
Bioinform. | 1 |
| 2023 | The scalable precision medicine open knowledge engine (SPOKE): a massive knowledge graph of biomedical informationabstractMOTIVATION: Knowledge graphs (KGs) are being adopted in industry, commerce and academia. Biomedical KG presents a challenge due to the complexity, size and heterogeneity of the underlying information. RESULTS: In this work, we present the Scalable Precision Medicine Open Knowledge Engine (SPOKE), a biomedical KG connecting millions of concepts via semantically meaningful relationships. SPOKE contains 27 million nodes of 21 different types and 53 million edges of 55 types downloaded from 41 databases. The graph is built on the framework of 11 ontologies that maintain its structure, enable mappings and facilitate navigation. SPOKE is built weekly by python scripts which download each resource, check for integrity and completeness, and then create a 'parent table' of nodes and edges. Graph queries are translated by a REST API and users can submit searches directly via an API or a graphical user interface. Conclusions/Significance: SPOKE enables the integration of seemingly disparate information to support precision medicine efforts. AVAILABILITY AND IMPLEMENTATION: The SPOKE neighborhood explorer is available at https://spoke.rbvi.ucsf.edu. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. John Scotter Morris, Karthik Soman, Rabia E. Akbas, Xiaoyuan Zhou, Brett Smith, Elaine C. Meng, Conrad C. Huang, Gabriel Cerono, Gundolf Schenk, Angela Rizk-Jackson, Adil Harroud, Lauren M. Sanders, Sylvain V. Costes, Krish Bharat, Arjun Chakraborty, Alexander R. Pico, Taline Mardirossian, Michael J. Keiser, Alice Tang, Josef Hardi, Yongmei Shi, Mark A. Musen, Sharat Israni, Sui Huang, Peter W. Rose, Charlotte A. Nelson, Sergio Baranzini |
Bioinform. | 2 |