VLDB 2026 Research / reviewers in the wild / expert
James P. Balhoff
dblp:90/11448
· DBLP profile ↗
12ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0002-8688-6599ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 3 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Vision and language · 40% Generative modeling · 37% Representation and self-supervised learning · 19% | |
| Interdisciplinary, comprehensive, and emerging computing
6 papers |
Bioinformatics and computational biology · 91% Computational science and engineering · 9% | |
| Databases, data mining, and information retrieval
1 paper |
Knowledge graphs · 75% Information retrieval · 25% |
Topics — the 18 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › contrastive learning
hierarchical contrastive learning |
0.9 | 1 | 2025 | BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
conditional diffusion model |
0.8 | 1 | 2024 | Hierarchical Conditioning of Diffusion Models Using Tree-of-Life for Studying Species Evolution · ECCV (89) 2024 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Hierarchical Conditioning of Diffusion Models Using Tree-of-Life for Studying Species Evolution · ECCV (89) 2024 |
Computer vision › Vision and language › vision-language model
vision-language model evaluation |
0.8 | 1 | 2024 | VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images · NeurIPS 2024 |
Bioinformatics and computational biology
evolutionary biology |
0.8 | 1 | 2024 | Hierarchical Conditioning of Diffusion Models Using Tree-of-Life for Studying Species Evolution · ECCV (89) 2024 |
Bioinformatics and computational biology › network bioinformatics › biological network analysis
biological knowledge graph |
0.7 | 1 | 2023 | KG-Hub - building and exchanging biological knowledge graphs · Bioinform. 2023 |
Bioinformatics and computational biology
phylogenetics |
0.7 | 1 | 2023 | Discovering Novel Biological Traits From Images Using Phylogeny-Guided Neural Networks · KDD 2023 |
Bioinformatics and computational biology › biological database
pathway database |
0.5 | 1 | 2021 | Reactome and the Gene Ontology: digital convergence of data resources · Bioinform. 2021 |
Knowledge graphs
knowledge graph exploration |
0.4 | 1 | 2019 | ROBOKOP: an abstraction layer and user interface for knowledge graphs to support question answering · Bioinform. 2019 |
Knowledge graphs
knowledge graph querying |
0.4 | 1 | 2019 | ROBOKOP: an abstraction layer and user interface for knowledge graphs to support question answering · Bioinform. 2019 |
Knowledge graphs › knowledge graph querying
knowledge graph question answering |
0.4 | 1 | 2019 | ROBOKOP: an abstraction layer and user interface for knowledge graphs to support question answering · Bioinform. 2019 |
Information retrieval › ranking › graph-based ranking
subgraph ranking |
0.4 | 1 | 2019 | ROBOKOP: an abstraction layer and user interface for knowledge graphs to support question answering · Bioinform. 2019 |
Computer vision › Vision and language › vision-language model
pre-trained vision-language model |
0.2 | 1 | 2024 | VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › hallucination
vision-language model hallucination |
0.2 | 1 | 2024 | VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images · NeurIPS 2024 |
Machine learning › Generative modeling › generative adversarial network
image-to-image translation |
0.2 | 1 | 2023 | Discovering Novel Biological Traits From Images Using Phylogeny-Guided Neural Networks · KDD 2023 |
Computational science and engineering
graph learning |
0.2 | 1 | 2023 | KG-Hub - building and exchanging biological knowledge graphs · Bioinform. 2023 |
Computational science and engineering › graph learning
link prediction |
0.2 | 1 | 2023 | KG-Hub - building and exchanging biological knowledge graphs · Bioinform. 2023 |
Bioinformatics and computational biology › knowledge representation in biology
biomedical knowledge graph |
0.1 | 1 | 2019 | ROBOKOP: an abstraction layer and user interface for knowledge graphs to support question answering · Bioinform. 2019 |
Methods — techniques the papers use, named apart from their topics
zero-shot evaluation · 1.5prompting techniques · 1.5phylogenetic embeddings · 1.5diffusion model · 1.5quantization · 1.3phylogeny encoding · 1.3neural network · 1.3embedding space analysis · 0.9contrastive learning · 0.9node embedding · 0.7graph machine learning · 0.7extract-transform-load · 0.7graph database querying · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive LearningabstractFoundation models trained at scale exhibit remarkable emergent behaviors, learning new capabilities beyond their initial training objectives. We find such emergent behaviors in biological vision models via large-scale contrastive vision-language training. To achieve this, we first curate TreeOfLife-200M, comprising 214 million images of living organisms, the largest and most diverse biological organism image dataset to date. We then train BioCLIP 2 on TreeOfLife-200M to distinguish different species. Despite the narrow training objective, BioCLIP 2 yields extraordinary accuracy when applied to various biological visual tasks such as habitat classification and trait prediction. We identify emergent properties in the learned embedding space of BioCLIP 2. At the inter-species level, the embedding distribution of different species aligns closely with functional and ecological meanings (e.g., beak sizes and habitats). At the intra-species level, instead of being diminished, the intra-species variations (e.g., life stages and sexes) are preserved and better separated in subspaces orthogonal to inter-species distinctions. We provide formal proof and analyses to explain why hierarchical supervision and contrastive objectives encourage these emergent properties. Crucially, our results reveal that these properties become increasingly significant with larger-scale training data, leading to a biologically meaningful embedding space. Jianyang Gu, Samuel Stevens 0001, Elizabeth G. Campolongo, Matthew J. Thompson, Net Zhang, Jiaman Wu, Andrei Kopanev, Zheda Mai, Alexander E. White, James P. Balhoff, Wasila M. Dahdul, Daniel I. Rubenstein, Hilmar Lapp, Tanya Y. Berger-Wolf, Wei-Lun Chao, Yu Su 0001 |
NeurIPS | 10 |
| 2024 | Hierarchical Conditioning of Diffusion Models Using Tree-of-Life for Studying Species EvolutionabstractAbstract A central problem in biology is to understand how organisms evolve and adapt to their environment by acquiring variations in the observable characteristics or traits of species across the tree of life. With the growing availability of large-scale image repositories in biology and recent advances in generative modeling, there is an opportunity to accelerate the discovery of evolutionary traits automatically from images. Toward this goal, we introduce Phylo-Diffusion, a novel framework for conditioning diffusion models with phylogenetic knowledge represented in the form of HIERarchical Embeddings (HIER-Embeds). We also propose two new experiments for perturbing the embedding space of Phylo-Diffusion: trait masking and trait swapping, inspired by counterpart experiments of gene knockout and gene editing/swapping. Our work represents a novel methodological advance in generative modeling to structure the embedding space of diffusion models using tree-based knowledge. Our work also opens a new chapter of research in evolutionary biology by using generative models to visualize evolutionary changes directly from images. We empirically demonstrate the usefulness of Phylo-Diffusion in capturing meaningful trait variations for fishes and birds, revealing novel insights about the biological mechanisms of their evolution. (Model and code can be found at imageomics.github.io/phylo-diffusion ) Mridul Khurana, Arka Daw, M. Maruf, Josef C. Uyeda, Wasila M. Dahdul, Caleb Charpentier, Yasin Bakis, Henry L. Bart Jr., Paula M. Mabee, Hilmar Lapp, James P. Balhoff, Wei-Lun Chao, Charles V. Stewart, Tanya Y. Berger-Wolf, Anuj Karpatne |
ECCV (89) | 11 |
| 2024 | VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological ImagesabstractImages are increasingly becoming the currency for documenting biodiversity on the planet, providing novel opportunities for accelerating scientific discoveries in the field of organismal biology, especially with the advent of large vision-language models (VLMs). We ask if pre-trained VLMs can aid scientists in answering a range of biologically relevant questions without any additional fine-tuning. In this paper, we evaluate the effectiveness of $12$ state-of-the-art (SOTA) VLMs in the field of organismal biology using a novel dataset, VLM4Bio, consisting of $469K$ question-answer pairs involving $30K$ images from three groups of organisms: fishes, birds, and butterflies, covering five biologically relevant tasks. We also explore the effects of applying prompting techniques and tests for reasoning hallucination on the performance of VLMs, shedding new light on the capabilities of current SOTA VLMs in answering biologically relevant questions using images. M. Maruf, Arka Daw, Kazi Sajeed Mehrab, Harish Babu Manogaran, Abhilash Neog, Medha Sawhney, Mridul Khurana, James P. Balhoff, Yasin Bakis, Bahadir Altintas, Matthew J. Thompson, Elizabeth G. Campolongo, Josef C. Uyeda, Hilmar Lapp, Henry L. Bart Jr., Paula M. Mabee, Yu Su 0001, Wei-Lun Chao, Charles V. Stewart, Tanya Y. Berger-Wolf, Wasila M. Dahdul, Anuj Karpatne |
NeurIPS | 8 |
| 2023 | Discovering Novel Biological Traits From Images Using Phylogeny-Guided Neural NetworksabstractDiscovering evolutionary traits that are heritable across species on the tree of life (also referred to as a phylogenetic tree) is of great interest to biologists to understand how organisms diversify and evolve. However, the measurement of traits is often a subjective and labor-intensive process, making trait discovery a highly label-scarce problem. We present a novel approach for discovering evolutionary traits directly from images without relying on trait labels. Our proposed approach, Phylo-NN, encodes the image of an organism into a sequence of quantized feature vectors -or codes- where different segments of the sequence capture evolutionary signals at varying ancestry levels in the phylogeny. We demonstrate the effectiveness of our approach in producing biologically meaningful results in a number of downstream tasks including species image generation and species-to-species image translation, using fish species as a target example Mohannad Elhamod, Mridul Khurana, Harish Babu Manogaran, Josef C. Uyeda, Meghan A. Balk, Wasila M. Dahdul, Yasin Bakis, Henry L. Bart Jr., Paula M. Mabee, Hilmar Lapp, James P. Balhoff, Caleb Charpentier, David Carlyn, Wei-Lun Chao, Charles V. Stewart, Daniel I. Rubenstein, Tanya Y. Berger-Wolf, Anuj Karpatne |
KDD | 11 |
| 2023 | KG-Hub - building and exchanging biological knowledge graphsabstractMOTIVATION: Knowledge graphs (KGs) are a powerful approach for integrating heterogeneous data and making inferences in biology and many other domains, but a coherent solution for constructing, exchanging, and facilitating the downstream use of KGs is lacking. RESULTS: Here we present KG-Hub, a platform that enables standardized construction, exchange, and reuse of KGs. Features include a simple, modular extract-transform-load pattern for producing graphs compliant with Biolink Model (a high-level data model for standardizing biological data), easy integration of any OBO (Open Biological and Biomedical Ontologies) ontology, cached downloads of upstream data sources, versioned and automatically updated builds with stable URLs, web-browsable storage of KG artifacts on cloud infrastructure, and easy reuse of transformed subgraphs across projects. Current KG-Hub projects span use cases including COVID-19 research, drug repurposing, microbial-environmental interactions, and rare disease research. KG-Hub is equipped with tooling to easily analyze and manipulate KGs. KG-Hub is also tightly integrated with graph machine learning (ML) tools which allow automated graph ML, including node embeddings and training of models for link prediction and node classification. AVAILABILITY AND IMPLEMENTATION: https://kghub.org. J. Harry Caufield, Tim E. Putman, Kevin Schaper, Deepak R. Unni, Harshad Hegde, Tiffany Callahan, Luca Cappelletti, Sierra A. T. Moxon, Vida Ravanmehr, Seth Carbon, Lauren E. Chan, Katherina G. Cortes, Kent A. Shefchek, Glass Elsarboukh, James P. Balhoff, Tommaso Fontana, Nicolas Matentzoglu, Richard M. Bruskiewich, Anne E. Thessen, Nomi L. Harris, Monica C. Munoz-Torres, Melissa A. Haendel, Peter N. Robinson, Marcin P. Joachimiak, Chris Mungall, Justin T. Reese |
Bioinform. | 15 |
| 2022 | Maximizing Interoperability, Enriching EHR Data: Transforming HL7 FHIR Data to RDF Using the FHIR RDF Playground
James Champion, Eric Prud'hommeaux, David Booth, Gaurav Vaidya, James P. Balhoff, Deepak K. Sharma, Guoqian Jiang, Emily R. Pfaff |
AMIA | 5 |
| 2021 | Reactome and the Gene Ontology: digital convergence of data resourcesabstractMOTIVATION: Gene Ontology Causal Activity Models (GO-CAMs) assemble individual associations of gene products with cellular components, molecular functions and biological processes into causally linked activity flow models. Pathway databases such as the Reactome Knowledgebase create detailed molecular process descriptions of reactions and assemble them, based on sharing of entities between individual reactions into pathway descriptions. RESULTS: To convert the rich content of Reactome into GO-CAMs, we have developed a software tool, Pathways2GO, to convert the entire set of normal human Reactome pathways into GO-CAMs. This conversion yields standard GO annotations from Reactome content and supports enhanced quality control for both Reactome and GO, yielding a nearly seamless conversion between these two resources for the bioinformatics community. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Benjamin M. Good, Kimberly Van Auken, David P. Hill, Huaiyu Mi, Seth Carbon, James P. Balhoff, Laurent-Philippe Albou, Paul D. Thomas, Chris Mungall, Judith A. Blake, Peter D'Eustachio |
Bioinform. | 6 |
| 2020 | Transforming the study of organisms: Phenomic data models and knowledge basesabstractThe rapidly decreasing cost of gene sequencing has resulted in a deluge of genomic data from across the tree of life; however, outside a few model organism databases, genomic data are limited in their scientific impact because they are not accompanied by computable phenomic data. The majority of phenomic data are contained in countless small, heterogeneous phenotypic data sets that are very difficult or impossible to integrate at scale because of variable formats, lack of digitization, and linguistic problems. One powerful solution is to represent phenotypic data using data models with precise, computable semantics, but adoption of semantic standards for representing phenotypic data has been slow, especially in biodiversity and ecology. Some phenotypic and trait data are available in a semantic language from knowledge bases, but these are often not interoperable. In this review, we will compare and contrast existing ontology and data models, focusing on nonhuman phenotypes and traits. We discuss barriers to integration of phenotypic data and make recommendations for developing an operationally useful, semantically interoperable phenotypic data ecosystem. Anne E. Thessen, Ramona L. Walls, Lars Vogt, Jessica Singer, Robert Warren, Pier Luigi Buttigieg, James P. Balhoff, Chris Mungall, Deborah L. McGuinness, Brian J. Stucky, Matthew J. Yoder, Melissa A. Haendel |
PLoS Comput. Biol. | 7 |
| 2019 | ROBOKOP: an abstraction layer and user interface for knowledge graphs to support question answeringabstractSUMMARY: Knowledge graphs (KGs) are quickly becoming a common-place tool for storing relationships between entities from which higher-level reasoning can be conducted. KGs are typically stored in a graph-database format, and graph-database queries can be used to answer questions of interest that have been posed by users such as biomedical researchers. For simple queries, the inclusion of direct connections in the KG and the storage and analysis of query results are straightforward; however, for complex queries, these capabilities become exponentially more challenging with each increase in complexity of the query. For instance, one relatively complex query can yield a KG with hundreds of thousands of query results. Thus, the ability to efficiently query, store, rank and explore sub-graphs of a complex KG represents a major challenge to any effort designed to exploit the use of KGs for applications in biomedical research and other domains. We present Reasoning Over Biomedical Objects linked in Knowledge Oriented Pathways as an abstraction layer and user interface to more easily query KGs and store, rank and explore query results. AVAILABILITY AND IMPLEMENTATION: An instance of the ROBOKOP UI for exploration of the ROBOKOP Knowledge Graph can be found at http://robokop.renci.org. The ROBOKOP Knowledge Graph can be accessed at http://robokopkg.renci.org. Code and instructions for building and deploying ROBOKOP are available under the MIT open software license from https://github.com/NCATS-Gamma/robokop. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kenneth Morton, Patrick Wang 0003, Chris Bizon, Steven Cox 0001, James P. Balhoff, Yaphet Kebede, Karamarie Fecho, Alexander Tropsha |
Bioinform. | 5 |
| 2019 | ROBOT: A Tool for Automating Ontology WorkflowsabstractBACKGROUND: Ontologies are invaluable in the life sciences, but building and maintaining ontologies often requires a challenging number of distinct tasks such as running automated reasoners and quality control checks, extracting dependencies and application-specific subsets, generating standard reports, and generating release files in multiple formats. Similar to more general software development, automation is the key to executing and managing these tasks effectively and to releasing more robust products in standard forms. For ontologies using the Web Ontology Language (OWL), the OWL API Java library is the foundation for a range of software tools, including the Protégé ontology editor. In the Open Biological and Biomedical Ontologies (OBO) community, we recognized the need to package a wide range of low-level OWL API functionality into a library of common higher-level operations and to make those operations available as a command-line tool. RESULTS: ROBOT (a recursive acronym for "ROBOT is an OBO Tool") is an open source library and command-line tool for automating ontology development tasks. The library can be called from any programming language that runs on the Java Virtual Machine (JVM). Most usage is through the command-line tool, which runs on macOS, Linux, and Windows. ROBOT provides ontology processing commands for a variety of tasks, including commands for converting formats, running a reasoner, creating import modules, running reports, and various other tasks. These commands can be combined into larger workflows using a separate task execution system such as GNU Make, and workflows can be automatically executed within continuous integration systems. CONCLUSIONS: ROBOT supports automation of a wide range of ontology development tasks, focusing on OBO conventions. It packages common high-level ontology development functionality into a convenient library, and makes it easy to configure, combine, and execute individual tasks in comprehensive, automated workflows. This helps ontology developers to efficiently create, maintain, and release high-quality ontologies, so that they can spend more time focusing on development tasks. It also helps guarantee that released ontologies are free of certain types of logical errors and conform to standard quality control checks, increasing the overall robustness and efficiency of the ontology development lifecycle. Rebecca C. Jackson, James P. Balhoff, Eric Douglass, Nomi L. Harris, Chris Mungall, James A. Overton |
BMC Bioinform. | 2 |
| 2017 | A generic bioinformatics pipeline to integrate large-scale trait data with large phylogeniesabstractAncestral state reconstructions are used to infer the evolutionary history of phenotypic traits to understand their evolution. Current ancestral state reconstructions require integration of larger trait matrices with larger phylogenies, which introduces several challenges. We identified these challenges and developed a generic pipeline that uses the Phenoscape Knowledgebase (Phenoscape KB) to retrieve raw trait matrices and convert them to an efficient version that can be easily integrated with large phylogenies downloaded from the Open Tree of Life (Open Tree). We demonstrate the performance of this pipeline using the evolution of the pectoral and pelvic fins as the use case, which involves integration of a large-scale phylogeny containing over 38,000 taxa. Pasan C. Fernando, Laura M. Jackson, Erliang Zeng, Paula M. Mabee, James P. Balhoff |
BIBM | 5 |
| 2013 | Phylotastic! Making tree-of-life knowledge accessible, reusable and convenientabstractBACKGROUND: Scientists rarely reuse expert knowledge of phylogeny, in spite of years of effort to assemble a great "Tree of Life" (ToL). A notable exception involves the use of Phylomatic, which provides tools to generate custom phylogenies from a large, pre-computed, expert phylogeny of plant taxa. This suggests great potential for a more generalized system that, starting with a query consisting of a list of any known species, would rectify non-standard names, identify expert phylogenies containing the implicated taxa, prune away unneeded parts, and supply branch lengths and annotations, resulting in a custom phylogeny suited to the user's needs. Such a system could become a sustainable community resource if implemented as a distributed system of loosely coupled parts that interact through clearly defined interfaces. RESULTS: With the aim of building such a "phylotastic" system, the NESCent Hackathons, Interoperability, Phylogenies (HIP) working group recruited 2 dozen scientist-programmers to a weeklong programming hackathon in June 2012. During the hackathon (and a three-month follow-up period), 5 teams produced designs, implementations, documentation, presentations, and tests including: (1) a generalized scheme for integrating components; (2) proof-of-concept pruners and controllers; (3) a meta-API for taxonomic name resolution services; (4) a system for storing, finding, and retrieving phylogenies using semantic web technologies for data exchange, storage, and querying; (5) an innovative new service, DateLife.org, which synthesizes pre-computed, time-calibrated phylogenies to assign ages to nodes; and (6) demonstration projects. These outcomes are accessible via a public code repository (GitHub.com), a website (http://www.phylotastic.org), and a server image. CONCLUSIONS: Approximately 9 person-months of effort (centered on a software development hackathon) resulted in the design and implementation of proof-of-concept software for 4 core phylotastic components, 3 controllers, and 3 end-user demonstration tools. While these products have substantial limitations, they suggest considerable potential for a distributed system that makes phylogenetic knowledge readily accessible in computable form. Widespread use of phylotastic systems will create an electronic marketplace for sharing phylogenetic knowledge that will spur innovation in other areas of the ToL enterprise, such as annotation of sources and methods and third-party methods of quality assessment. Arlin Stoltzfus, Hilmar Lapp, Naim Matasci, Helena F. Deus, Brian Sidlauskas, Christian M. Zmasek, Gaurav Vaidya, Enrico Pontelli, Karen Cranston, Rutger A. Vos, Campbell O. Webb, Luke J. Harmon, Meg Pirrung, Brian C. O'Meara, Matthew W. Pennell, Siavash Mirarab, Michael S. Rosenberg, James P. Balhoff, Holly M. Bik, Tracy A. Heath, Peter E. Midford, Joseph W. Brown, Emily Jane McTavish, Jeet Sukumaran, Mark Westneat, Michael E. Alfaro, Aaron Steele, Greg Jordan |
BMC Bioinform. | 18 |