VLDB 2026 Research / reviewers in the wild / expert
Yongmei Shi
dblp:72/1103
· DBLP profile ↗
9ranked-venue papers
2as first author
4since 2021 · last 2024
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 72% Medical and health informatics · 28% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Knowledge graphs · 50% Data mining · 50% | |
| Artificial intelligence
2 papers |
Information extraction and text analysis · 38% Language models and text generation · 38% Speech recognition and synthesis · 25% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › knowledge representation in biology
biomedical knowledge graph |
1.4 | 2 | 2024 | Biomedical knowledge graph-optimized prompt generation for large language models · Bioinform. 2024 The scalable precision medicine open knowledge engine (SPOKE): a massive knowledge graph of biomedical information · Bioinform. 2023 |
Bioinformatics and computational biology
knowledge graph |
0.9 | 2 | 2024 | The scalable precision medicine open knowledge engine (SPOKE): a massive knowledge graph of biomedical information · Bioinform. 2023 Biomedical knowledge graph-optimized prompt generation for large language models · Bioinform. 2024 |
Bioinformatics and computational biology › biomedical text mining
biomedical question answering |
0.8 | 1 | 2024 | Biomedical knowledge graph-optimized prompt generation for large language models · Bioinform. 2024 |
Medical and health informatics
retrieval-augmented generation |
0.8 | 1 | 2024 | Biomedical knowledge graph-optimized prompt generation for large language models · Bioinform. 2024 |
Bioinformatics and computational biology › data integration
biomedical data integration |
0.7 | 1 | 2023 | The scalable precision medicine open knowledge engine (SPOKE): a massive knowledge graph of biomedical information · Bioinform. 2023 |
Medical and health informatics
precision medicine |
0.7 | 1 | 2023 | The scalable precision medicine open knowledge engine (SPOKE): a massive knowledge graph of biomedical information · Bioinform. 2023 |
Parallel and multicore computing › parallel algorithms › graph algorithms
all-pairs shortest paths |
0.6 | 1 | 2022 | Exaflops Biomedical Knowledge Graph Analytics · SC 2022 |
Parallel and multicore computing
graph processing |
0.6 | 1 | 2022 | Exaflops Biomedical Knowledge Graph Analytics · SC 2022 |
Knowledge graphs › domain-specific knowledge graph
biomedical knowledge graph |
0.2 | 1 | 2022 | Exaflops Biomedical Knowledge Graph Analytics · SC 2022 |
Data mining › structured data mining
relational data mining |
0.2 | 1 | 2022 | Exaflops Biomedical Knowledge Graph Analytics · SC 2022 |
Natural language and speech › Information extraction and text analysis › text classification
deception detection |
0.1 | 1 | 2008 | A Statistical Language Modeling Approach to Online Deception Detection · IEEE Trans. Knowl. Data Eng. 2008 |
Natural language and speech › Language models and text generation › language modeling
statistical language modeling |
0.1 | 1 | 2008 | A Statistical Language Modeling Approach to Online Deception Detection · IEEE Trans. Knowl. Data Eng. 2008 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.1 | 1 | 2005 | Data Mining for Detecting Errors in Dictation Speech Recognition · IEEE Trans. Speech Audio Process. 2005 |
Methods — techniques the papers use, named apart from their topics
tropical semiring · 1.1hyperbolic performance modeling · 1.1retrieval-augmented generation · 0.8large language model · 0.8embedding-based context pruning · 0.8ontology-based integration · 0.7REST API · 0.7vocabulary pruning · 0.1smoothing · 0.1support vector machine · 0.1neural network · 0.1naive bayes · 0.1link grammar · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Biomedical knowledge graph-optimized prompt generation for large language modelsabstractMOTIVATION: Large language models (LLMs) are being adopted at an unprecedented rate, yet still face challenges in knowledge-intensive domains such as biomedicine. Solutions such as pretraining and domain-specific fine-tuning add substantial computational overhead, requiring further domain-expertise. Here, we introduce a token-optimized and robust Knowledge Graph-based Retrieval Augmented Generation (KG-RAG) framework by leveraging a massive biomedical KG (SPOKE) with LLMs such as Llama-2-13b, GPT-3.5-Turbo, and GPT-4, to generate meaningful biomedical text rooted in established knowledge. RESULTS: Compared to the existing RAG technique for Knowledge Graphs, the proposed method utilizes minimal graph schema for context extraction and uses embedding methods for context pruning. This optimization in context extraction results in more than 50% reduction in token consumption without compromising the accuracy, making a cost-effective and robust RAG implementation on proprietary LLMs. KG-RAG consistently enhanced the performance of LLMs across diverse biomedical prompts by generating responses rooted in established knowledge, accompanied by accurate provenance and statistical evidence (if available) to substantiate the claims. Further benchmarking on human curated datasets, such as biomedical true/false and multiple-choice questions (MCQ), showed a remarkable 71% boost in the performance of the Llama-2 model on the challenging MCQ dataset, demonstrating the framework's capacity to empower open-source models with fewer parameters for domain-specific questions. Furthermore, KG-RAG enhanced the performance of proprietary GPT models, such as GPT-3.5 and GPT-4. In summary, the proposed framework combines explicit and implicit knowledge of KG and LLM in a token optimized fashion, thus enhancing the adaptability of general-purpose LLMs to tackle domain-specific questions in a cost-effective fashion. AVAILABILITY AND IMPLEMENTATION: SPOKE KG can be accessed at https://spoke.rbvi.ucsf.edu/neighborhood.html. It can also be accessed using REST-API (https://spoke.rbvi.ucsf.edu/swagger/). KG-RAG code is made available at https://github.com/BaranziniLab/KG_RAG. Biomedical benchmark datasets used in this study are made available to the research community in the same GitHub repository. Karthik Soman, Peter W. Rose, John Scotter Morris, Rabia E. Akbas, Brett Smith, Braian Peetoom, Catalina Villouta-Reyes, Gabriel Cerono, Yongmei Shi, Angela Rizk-Jackson, Sharat Israni, Charlotte A. Nelson, Sui Huang, Sergio Baranzini |
Bioinform. | 9 |
| 2023 | The scalable precision medicine open knowledge engine (SPOKE): a massive knowledge graph of biomedical informationabstractMOTIVATION: Knowledge graphs (KGs) are being adopted in industry, commerce and academia. Biomedical KG presents a challenge due to the complexity, size and heterogeneity of the underlying information. RESULTS: In this work, we present the Scalable Precision Medicine Open Knowledge Engine (SPOKE), a biomedical KG connecting millions of concepts via semantically meaningful relationships. SPOKE contains 27 million nodes of 21 different types and 53 million edges of 55 types downloaded from 41 databases. The graph is built on the framework of 11 ontologies that maintain its structure, enable mappings and facilitate navigation. SPOKE is built weekly by python scripts which download each resource, check for integrity and completeness, and then create a 'parent table' of nodes and edges. Graph queries are translated by a REST API and users can submit searches directly via an API or a graphical user interface. Conclusions/Significance: SPOKE enables the integration of seemingly disparate information to support precision medicine efforts. AVAILABILITY AND IMPLEMENTATION: The SPOKE neighborhood explorer is available at https://spoke.rbvi.ucsf.edu. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. John Scotter Morris, Karthik Soman, Rabia E. Akbas, Xiaoyuan Zhou, Brett Smith, Elaine C. Meng, Conrad C. Huang, Gabriel Cerono, Gundolf Schenk, Angela Rizk-Jackson, Adil Harroud, Lauren M. Sanders, Sylvain V. Costes, Krish Bharat, Arjun Chakraborty, Alexander R. Pico, Taline Mardirossian, Michael J. Keiser, Alice Tang, Josef Hardi, Yongmei Shi, Mark A. Musen, Sharat Israni, Sui Huang, Peter W. Rose, Charlotte A. Nelson, Sergio Baranzini |
Bioinform. | 21 |
| 2022 | A Unified Blockchain Schema for Chronic Diet ManagementabstractData, information, and technology are improving healthcare aspects such as fitness tracking, patient-doctor communication, and access to health records, etc. Diet management approaches based on a functional estimation involve clinical and computer-supported methods that are utilized in the strategy of treatment. Clinical methods such as dietary records, 24-hour recall, and food frequency are common in practice and majorly rely on a calorie-conscious intake, which though with a cautionary proceeding works as expected, but inherently lags a focus on the processing levels of the ingredients such as the quantifiable content of natural or processed food. Calories consumed from natural sources share contrast to processed food. The exploration, hence, needs a hybrid approach that combines the computer-supported systems to build on a processed food proportion estimation methodology. Emerging augmented reality technologies seamlessly merge real-world environments with computer-supported perceptual information have the potential to solve the problem enabling user decision support through a unified schema of nutritional information. This research presents (i) a blueprint for a unified schema that merges clinical and computer-supported methods that underlie the viable standard for a host of multi-point chronic diet management applications, followed by (ii) a work-in-progress outline of the essential function blocks, resources, clinical considerations, and integrations that form a technology stack for the estimation of processed and non-processed food towards a clinical conscious diet intake. Dali Zhang, Usharani Hareesh Govindarajan, Yongmei Shi, Gagan Narang |
CSCWD | 3 |
| 2022 | Exaflops Biomedical Knowledge Graph AnalyticsabstractWe are motivated by newly proposed methods for mining large-scale corpora of scholarly publications (e.g., full biomedical literature), which consists of tens of millions of papers spanning decades of research. In this setting, analysts seek to discover relationships among concepts. They construct graph representations from annotated text databases and then formulate the relationship-mining problem as an all-pairs shortest paths (APSP) and validate connective paths against curated biomedical knowledge graphs (e.g., Spoke). In this context, we present Coast (Exascale Communication-Optimized All-Pairs Shortest Path) and demonstrate 1.004 EF/s on 9,200 Frontier nodes (73,600 GCDs). We develop hyperbolic performance models (HYPERMOD), which guide optimizations and parametric tuning. The proposed Coast algorithm achieved the memory constant parallel efficiency of 99% in the single-precision tropical semiring. Looking forward, Coast will enable the integration of scholarly corpora like PubMed into the Spoke biomedical knowledge graph. Ramakrishnan Kannan, Piyush Sao, Hao Lu 0001, Jakub Kurzak, Gundolf Schenk, Yongmei Shi, Seung-Hwan Lim, Sharat Israni, Vijay Thakkar, Guojing Cong, Robert M. Patton, Sergio Baranzini, Richard W. Vuduc, Thomas E. Potok |
SC | 6 |
| 2011 | Supporting dictation speech recognition error correction: the impact of external informationabstractAlthough speech recognition technology has made remarkable progress, its wide adoption is still restricted by notable effort made and frustration experienced by users while correcting speech recognition errors. One of the promising ways to improve error correction is by providing user support. Although support mechanisms have been proposed for inline speech recognition error correction, how to develop user support for third-party error correction has been under studied. To address the unique challenges associated with third-party error correction, external information that is obtained outside of an erroneous sentence was employed in this research. Specifically, three types of external information were selected, including word alternative hypotheses, noisy context and accurate context, and their impacts on dictation speech recognition error correction were assessed empirically. As expected, results revealed the importance of context information to improving both the performance and perception of error correction. Additionally, the results also provided insights into possible effects of word error rate and sentence length on error correction. These findings have significant implications to interface design for transcript proofreading systems. Yongmei Shi, Lina Zhou |
Behav. Inf. Technol. | 1 |
| 2010 | Third-party error detection support mechanisms for dictation speech recognitionabstractAlthough speech recognition has improved significantly in recent years, its adoption continues to be limited, in part, by the effort and frustration associated with correcting speech recognition errors. Error detection is a particularly challenging issue in third-party error correction where different individuals are responsible for the original dictation and correcting the resulting text. This research aims to address the difficulty experienced in third-party error detection by developing and evaluating a variety of support mechanisms. Drawing on a growing body of literature on human computer interaction and speech recognition, four support mechanisms were designed and evaluated, namely indexed audio, speech summarization, error prediction, and the presentation of alternative hypotheses. A user study assessed the impact of these support mechanisms on both performance and perceptions during error detection tasks. Performance measures included effectiveness and efficiency, and perception measures included confidence, perceived usefulness, and cognitive workload. The results provide strong support for the use of indexed audio in the context of third-party error detection. The results also confirm that consecutive error rate, or the percentage of recognition errors immediately adjacent to another error, has a negative impact on the effectiveness of third-party error detection. Other support mechanisms failed to improve either effectiveness or perceptions, but they did negate the negative impact as consecutive error rate increased. These findings have significant implications for speech recognition error detection research and the design of error detection support solutions. Lina Zhou, Yongmei Shi, Andrew Sears |
Interact. Comput. | 2 |
| 2008 | A Statistical Language Modeling Approach to Online Deception DetectionabstractOnline deception is disrupting our daily life, organizational process, and even national security. Existing approaches to online deception detection follow a traditional paradigm by using a set of cues as antecedents for deception detection, which may be hindered by ineffective cue identification. Motivated by the strength of statistical language models (SLMs) in capturing the dependency of words in text without explicit feature extraction, we developed SLMs to detect online deception. We also addressed the data sparsity problem in building SLMs in general and in deception detection in specific using smoothing and vocabulary pruning techniques. The developed SLMs were evaluated empirically with diverse datasets. The results showed that the proposed SLM approach to deception detection outperformed a state-of-the-art text categorization method as well as traditional feature-based methods. Lina Zhou, Yongmei Shi, Dongsong Zhang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2006 | Examining knowledge sources for human error correction
Yongmei Shi, Lina Zhou |
INTERSPEECH | 1 |
| 2005 | Data Mining for Detecting Errors in Dictation Speech RecognitionabstractThe efficiency promised by a dictation speech recognition (DSR) system is lessened by the need for correcting recognition errors. Error detection is the precursor of error correction. Developing effective techniques for error detection can thus lead to improved error correction. Current research on error detection has focused mainly on transcription and/or domain-specific speech. Error detection in DSR has been studied less. We propose data mining models for detecting errors in DSR. Instead of relying on internal parameters from DSR systems, we propose a loosely coupled approach to error detection based on features extracted from the DSR output. The features mainly came from two sources: confidence scores and linguistics parsing. Link grammar was innovatively applied to error detection. Three data mining techniques, including Na/spl inodot//spl uml/ve Bayes, neural networks, and Support Vector Machines (SVMs), were evaluated on 5M DSR corpora. The experimental results showed that significant performance was achieved in that F-measures for error detection ranged from 55.3% to 62.5%. This study provided insights into the merit of different data-mining techniques and different types of features in error detection. Lina Zhou, Yongmei Shi, Jinjuan Feng, Andrew Sears |
IEEE Trans. Speech Audio Process. | 2 |