VLDB 2026 Research / reviewers in the wild / expert
Victoria Yao
dblp:180/8016 · also Vicky Yao
· DBLP profile ↗
6ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-3201-9983ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 100% | |
| Artificial intelligence
1 paper |
Information extraction and text analysis · 100% |
Topics — the 7 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › protein analysis › protein-protein interaction
protein-protein interaction network analysis |
1.0 | 1 | 2026 | Splitpea: a Python package for protein-protein interaction network rewiring analysis due to alternative splicing · Bioinform. 2026 |
Bioinformatics and computational biology › statistical genetics
phenotype prediction |
0.9 | 1 | 2025 | ALPINE: An Interpretable Approach for Decoding Phenotypes from Multi-condition Sequencing Data · RECOMB 2025 |
Bioinformatics and computational biology › functional genomics › functional enrichment analysis
gene set enrichment analysis |
0.8 | 1 | 2024 | Enhancing Gene Set Analysis in Embedding Spaces: A Novel Best-Match Approach · RECOMB 2024 |
Natural language and speech › Information extraction and text analysis › named entity recognition
biomedical named entity recognition |
0.7 | 1 | 2023 | Into the Single Cell Multiverse: an End-to-End Dataset for Procedural Knowledge Extraction in Biomedical Texts · NeurIPS 2023 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.7 | 1 | 2023 | Into the Single Cell Multiverse: an End-to-End Dataset for Procedural Knowledge Extraction in Biomedical Texts · NeurIPS 2023 |
Bioinformatics and computational biology › network bioinformatics › biological network analysis › network alignment
biological network alignment |
0.7 | 1 | 2023 | Joint embedding of biological networks for cross-species functional alignment · Bioinform. 2023 |
Bioinformatics and computational biology
transcriptomics |
0.3 | 1 | 2026 | Splitpea: a Python package for protein-protein interaction network rewiring analysis due to alternative splicing · Bioinform. 2026 |
Methods — techniques the papers use, named apart from their topics
expert-curated annotation · 1.3entity disambiguation · 1.3network rewiring mapping · 1.0differential exon usage statistics · 1.0interpretable machine learning · 0.9best-match approach · 0.8network embedding · 0.7natural language processing-inspired cross-training · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Splitpea: a Python package for protein-protein interaction network rewiring analysis due to alternative splicingabstractSUMMARY: Splitpea takes skipped exon event data at the sample or differential expression level from SUPPA2 and rMATS and maps potential changes to protein-protein interaction (PPI) network rewiring events. It handles a variety of input formats via an easy-to-install Python package, including percent spliced in values comparing two conditions, skipped exon counts, or precalculated exon usage statistics between experimental conditions. In each case, Splitpea produces rewired network graphs, edge and gene-level summary statistics, and Cytoscape- or Gephi-ready files for easy visualization, allowing users to find PPIs potentially disrupted or increased by alternative splicing. AVAILABILITY AND IMPLEMENTATION: Source code and accompanying documentation can be found on Github (https://github.com/ylaboratory/splitpea-package), released under a BSD 3-clause license for open-source use, and the Splitpea package is installable via PyPI. Jeffrey Zhong, Alyssa Cantu, Ruth Dannenfelser, Victoria Yao |
Bioinform. | 4 |
| 2025 | ALPINE: An Interpretable Approach for Decoding Phenotypes from Multi-condition Sequencing Data
Wei-Hao Lee, Lechuan Li, Ruth Dannenfelser, Victoria Yao |
RECOMB | 4 |
| 2024 | Enhancing Gene Set Analysis in Embedding Spaces: A Novel Best-Match Approach
Lechuan Li, Ruth Dannenfelser, Charlie Cruz, Victoria Yao |
RECOMB | 4 |
| 2023 | Into the Single Cell Multiverse: an End-to-End Dataset for Procedural Knowledge Extraction in Biomedical TextsabstractMany of the most commonly explored natural language processing (NLP) information extraction tasks can be thought of as evaluations of declarative knowledge, or fact-based information extraction. Procedural knowledge extraction, i.e., breaking down a described process into a series of steps, has received much less attention, perhaps in part due to the lack of structured datasets that capture the knowledge extraction process from end-to-end. To address this unmet need, we present FlaMBé (Flow annotations for Multiverse Biological entities), a collection of expert-curated datasets across a series of complementary tasks that capture procedural knowledge in biomedical texts. This dataset is inspired by the observation that one ubiquitous source of procedural knowledge that is described as unstructured text is within academic papers describing their methodology. The workflows annotated in FlaMBé are from texts in the burgeoning field of single cell research, a research area that has become notorious for the number of software tools and complexity of workflows used. Additionally, FlaMBé provides, to our knowledge, the largest manually curated named entity recognition (NER) and disambiguation (NED) datasets for tissue/cell type, a fundamental biological entity that is critical for knowledge extraction in the biomedical research domain. Beyond providing a valuable dataset to enable further development of NLP models for procedural knowledge extraction, automating the process of workflow mining also has important implications for advancing reproducibility in biomedical research. Ruth Dannenfelser, Jeffrey Zhong, Victoria Yao |
NeurIPS | 4 |
| 2023 | Joint embedding of biological networks for cross-species functional alignmentabstractMOTIVATION: Model organisms are widely used to better understand the molecular causes of human disease. While sequence similarity greatly aids this cross-species transfer, sequence similarity does not imply functional similarity, and thus, several current approaches incorporate protein-protein interactions to help map findings between species. Existing transfer methods either formulate the alignment problem as a matching problem which pits network features against known orthology, or more recently, as a joint embedding problem. RESULTS: We propose a novel state-of-the-art joint embedding solution: Embeddings to Network Alignment (ETNA). ETNA generates individual network embeddings based on network topological structure and then uses a Natural Language Processing-inspired cross-training approach to align the two embeddings using sequence-based orthologs. The final embedding preserves both within and between species gene functional relationships, and we demonstrate that it captures both pairwise and group functional relevance. In addition, ETNA's embeddings can be used to transfer genetic interactions across species and identify phenotypic alignments, laying the groundwork for potential opportunities for drug repurposing and translational studies. AVAILABILITY AND IMPLEMENTATION: https://github.com/ylaboratory/ETNA. Lechuan Li, Ruth Dannenfelser, Yu Zhu 0003, Nathaniel Hejduk, Santiago Segarra, Victoria Yao |
Bioinform. | 6 |
| 2018 | A loop-counting method for covariate-corrected low-rank biclustering of gene-expression and genome-wide association study dataabstractA common goal in data-analysis is to sift through a large data-matrix and detect any significant submatrices (i.e., biclusters) that have a low numerical rank. We present a simple algorithm for tackling this biclustering problem. Our algorithm accumulates information about 2-by-2 submatrices (i.e., 'loops') within the data-matrix, and focuses on rows and columns of the data-matrix that participate in an abundance of low-rank loops. We demonstrate, through analysis and numerical-experiments, that this loop-counting method performs well in a variety of scenarios, outperforming simple spectral methods in many situations of interest. Another important feature of our method is that it can easily be modified to account for aspects of experimental design which commonly arise in practice. For example, our algorithm can be modified to correct for controls, categorical- and continuous-covariates, as well as sparsity within the data. We demonstrate these practical features with two examples; the first drawn from gene-expression analysis and the second drawn from a much larger genome-wide-association-study (GWAS). Aaditya V. Rangan, Caroline C. McGrouther, John Kelsoe, Nicholas J. Schork, Eli Stahl, Qian Zhu 0005, Arjun Krishnan, Victoria Yao, Olga G. Troyanskaya, Seda Bilaloglu, Preeti Raghavan, Sarah Bergen, Anders Juréus, Mikael Landen |
PLoS Comput. Biol. | 8 |