Jeffrey Zhong

dblp:356/2682 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0000-0003-1431-6973ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Information extraction and text analysis · 100%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › protein analysis › protein-protein interaction
protein-protein interaction network analysis
1.012026
Splitpea: a Python package for protein-protein interaction network rewiring analysis due to alternative splicing · Bioinform. 2026
Natural language and speech › Information extraction and text analysis › named entity recognition
biomedical named entity recognition
0.712023
Into the Single Cell Multiverse: an End-to-End Dataset for Procedural Knowledge Extraction in Biomedical Texts · NeurIPS 2023
Natural language and speech › Information extraction and text analysis
named entity recognition
0.712023
Into the Single Cell Multiverse: an End-to-End Dataset for Procedural Knowledge Extraction in Biomedical Texts · NeurIPS 2023
Bioinformatics and computational biology
transcriptomics
0.312026
Splitpea: a Python package for protein-protein interaction network rewiring analysis due to alternative splicing · Bioinform. 2026

Methods — techniques the papers use, named apart from their topics

expert-curated annotation · 1.3entity disambiguation · 1.3network rewiring mapping · 1.0differential exon usage statistics · 1.0
YearPublicationVenuePosition
2026 Splitpea: a Python package for protein-protein interaction network rewiring analysis due to alternative splicing
abstract
SUMMARY: Splitpea takes skipped exon event data at the sample or differential expression level from SUPPA2 and rMATS and maps potential changes to protein-protein interaction (PPI) network rewiring events. It handles a variety of input formats via an easy-to-install Python package, including percent spliced in values comparing two conditions, skipped exon counts, or precalculated exon usage statistics between experimental conditions. In each case, Splitpea produces rewired network graphs, edge and gene-level summary statistics, and Cytoscape- or Gephi-ready files for easy visualization, allowing users to find PPIs potentially disrupted or increased by alternative splicing. AVAILABILITY AND IMPLEMENTATION: Source code and accompanying documentation can be found on Github (https://github.com/ylaboratory/splitpea-package), released under a BSD 3-clause license for open-source use, and the Splitpea package is installable via PyPI.
Jeffrey Zhong, Alyssa Cantu, Ruth Dannenfelser, Victoria Yao
Bioinform.1
2023 Into the Single Cell Multiverse: an End-to-End Dataset for Procedural Knowledge Extraction in Biomedical Texts
abstract
Many of the most commonly explored natural language processing (NLP) information extraction tasks can be thought of as evaluations of declarative knowledge, or fact-based information extraction. Procedural knowledge extraction, i.e., breaking down a described process into a series of steps, has received much less attention, perhaps in part due to the lack of structured datasets that capture the knowledge extraction process from end-to-end. To address this unmet need, we present FlaMBé (Flow annotations for Multiverse Biological entities), a collection of expert-curated datasets across a series of complementary tasks that capture procedural knowledge in biomedical texts. This dataset is inspired by the observation that one ubiquitous source of procedural knowledge that is described as unstructured text is within academic papers describing their methodology. The workflows annotated in FlaMBé are from texts in the burgeoning field of single cell research, a research area that has become notorious for the number of software tools and complexity of workflows used. Additionally, FlaMBé provides, to our knowledge, the largest manually curated named entity recognition (NER) and disambiguation (NED) datasets for tissue/cell type, a fundamental biological entity that is critical for knowledge extraction in the biomedical research domain. Beyond providing a valuable dataset to enable further development of NLP models for procedural knowledge extraction, automating the process of workflow mining also has important implications for advancing reproducibility in biomedical research.
Ruth Dannenfelser, Jeffrey Zhong, Victoria Yao
NeurIPS2