Henry L. Bart Jr.

dblp:56/399 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0002-5662-9444ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 36% Vision and language · 21% Segmentation and scene understanding · 18%
Interdisciplinary, comprehensive, and emerging computing
6 papers
Bioinformatics and computational biology · 79% Environmental and earth informatics · 21%
Databases, data mining, and information retrieval
3 papers
Data mining · 100%

Topics — the 24 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Environmental and earth informatics
biodiversity informatics
0.912025
Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from Images · CVPR 2025
Bioinformatics and computational biology
species classification
0.912025
Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from Images · CVPR 2025
Machine learning › Generative modeling › diffusion model
conditional diffusion model
0.812024
Hierarchical Conditioning of Diffusion Models Using Tree-of-Life for Studying Species Evolution · ECCV (89) 2024
Machine learning › Generative modeling
diffusion model
0.812024
Hierarchical Conditioning of Diffusion Models Using Tree-of-Life for Studying Species Evolution · ECCV (89) 2024
Computer vision › Vision and language › vision-language model
vision-language model evaluation
0.812024
VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images · NeurIPS 2024
Bioinformatics and computational biology
evolutionary biology
0.812024
Hierarchical Conditioning of Diffusion Models Using Tree-of-Life for Studying Species Evolution · ECCV (89) 2024
Bioinformatics and computational biology
phylogenetics
0.712023
Discovering Novel Biological Traits From Images Using Phylogeny-Guided Neural Networks · KDD 2023
Data mining
anomaly detection
0.322015
Robust Model-Based Learning via Spatial-EM Algorithm · IEEE Trans. Knowl. Data Eng. 2015
Depth-Based Novelty Detection and Its Application to Taxonomic Research · ICDM 2007
Machine learning › Trustworthy machine learning › interpretability
explainable AI
0.312025
Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from Images · CVPR 2025
Machine learning › Learning paradigms
long-tailed recognition
0.312025
Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from Images · CVPR 2025
Computer vision › Vision and language › vision-language model
pre-trained vision-language model
0.212024
VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images · NeurIPS 2024
Machine learning › Trustworthy machine learning › hallucination
vision-language model hallucination
0.212024
VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
mixture model learning
0.212015
Robust Model-Based Learning via Spatial-EM Algorithm · IEEE Trans. Knowl. Data Eng. 2015
Data mining
clustering
0.212015
Robust Model-Based Learning via Spatial-EM Algorithm · IEEE Trans. Knowl. Data Eng. 2015
Data mining › anomaly detection
outlier detection
0.212015
Robust Model-Based Learning via Spatial-EM Algorithm · IEEE Trans. Knowl. Data Eng. 2015
Data mining › clustering
robust clustering
0.212015
Robust Model-Based Learning via Spatial-EM Algorithm · IEEE Trans. Knowl. Data Eng. 2015
Machine learning › Generative modeling › generative adversarial network
image-to-image translation
0.212023
Discovering Novel Biological Traits From Images Using Phylogeny-Guided Neural Networks · KDD 2023
Bioinformatics and computational biology
species identification
0.122007
Depth-Based Novelty Detection and Its Application to Taxonomic Research · ICDM 2007
A Computational Framework for Taxonomic Research: Diagnosing Body Shape within Fish Species Complexes · ICDM 2005
Bioinformatics and computational biology
taxonomy
0.122007
Depth-Based Novelty Detection and Its Application to Taxonomic Research · ICDM 2007
A Computational Framework for Taxonomic Research: Diagnosing Body Shape within Fish Species Complexes · ICDM 2005
Machine learning › Time series and sequential data
anomaly detection
0.112009
Outlier Detection with the Kernelized Spatial Depth Function · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Machine learning › Trustworthy machine learning
statistical depth
0.112009
Outlier Detection with the Kernelized Spatial Depth Function · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Data mining › anomaly detection
novelty detection
0.112007
Depth-Based Novelty Detection and Its Application to Taxonomic Research · ICDM 2007
Data mining › predictive modeling
classification
0.112005
A Computational Framework for Taxonomic Research: Diagnosing Body Shape within Fish Species Complexes · ICDM 2005
Data mining › dimensionality reduction
feature selection
0.112005
A Computational Framework for Taxonomic Research: Diagnosing Body Shape within Fish Species Complexes · ICDM 2005

Methods — techniques the papers use, named apart from their topics

machine learning · 1.7computer vision · 1.7zero-shot evaluation · 1.5prompting techniques · 1.5phylogenetic embeddings · 1.5diffusion model · 1.5quantization · 1.3phylogeny encoding · 1.3neural network · 1.3median-based location estimation · 0.4rank-based scatter estimation · 0.2kernelized spatial depth · 0.1statistical depth functions · 0.1classification · 0.1automated feature selection · 0.1
YearPublicationVenuePosition
2025 Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from Images
abstract
We introduce Fish-Visual Trait Analysis (Fish-Vista), the first organismal image dataset designed for the analysis of visual traits of aquatic species directly from images using machine learning and computer vision methods. Fish-Vista contains 69,269 annotated images spanning 4,316 fish species, curated and organized to serve three downstream tasks: species classification, trait identification, and trait segmentation. Our work makes two key contributions. First, we provide a fully reproducible data processing pipeline to process fish images sourced from various museum collections, contributing to the advancement of AI in biodiversity science. We annotate the images with carefully curated labels from biological databases and manual annotations to create an AI-ready dataset of visual traits. Second, our work offers fertile grounds for researchers to develop novel methods for a variety of problems in computer vision such as handling long-tailed distributions, out-of-distribution generalization, learning with weak labels, explainable AI, and segmenting small objects. Dataset and code for Fish-Vista are available at https://github.com/Imageomics/Fish-Vista
Kazi Sajeed Mehrab, M. Maruf, Arka Daw, Abhilash Neog, Harish Babu Manogaran, Mridul Khurana, Zhenyang Feng, Bahadir Altintas, Yasin Bakis, Elizabeth G. Campolongo, Matthew J. Thompson, Hilmar Lapp, Tanya Y. Berger-Wolf, Paula M. Mabee, Henry L. Bart Jr., Wei-Lun Chao, Wasila M. Dahdul, Anuj Karpatne
CVPR16
2024 Hierarchical Conditioning of Diffusion Models Using Tree-of-Life for Studying Species Evolution
abstract
Abstract A central problem in biology is to understand how organisms evolve and adapt to their environment by acquiring variations in the observable characteristics or traits of species across the tree of life. With the growing availability of large-scale image repositories in biology and recent advances in generative modeling, there is an opportunity to accelerate the discovery of evolutionary traits automatically from images. Toward this goal, we introduce Phylo-Diffusion, a novel framework for conditioning diffusion models with phylogenetic knowledge represented in the form of HIERarchical Embeddings (HIER-Embeds). We also propose two new experiments for perturbing the embedding space of Phylo-Diffusion: trait masking and trait swapping, inspired by counterpart experiments of gene knockout and gene editing/swapping. Our work represents a novel methodological advance in generative modeling to structure the embedding space of diffusion models using tree-based knowledge. Our work also opens a new chapter of research in evolutionary biology by using generative models to visualize evolutionary changes directly from images. We empirically demonstrate the usefulness of Phylo-Diffusion in capturing meaningful trait variations for fishes and birds, revealing novel insights about the biological mechanisms of their evolution. (Model and code can be found at imageomics.github.io/phylo-diffusion )
Mridul Khurana, Arka Daw, M. Maruf, Josef C. Uyeda, Wasila M. Dahdul, Caleb Charpentier, Yasin Bakis, Henry L. Bart Jr., Paula M. Mabee, Hilmar Lapp, James P. Balhoff, Wei-Lun Chao, Charles V. Stewart, Tanya Y. Berger-Wolf, Anuj Karpatne
ECCV (89)8
2024 VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images
abstract
Images are increasingly becoming the currency for documenting biodiversity on the planet, providing novel opportunities for accelerating scientific discoveries in the field of organismal biology, especially with the advent of large vision-language models (VLMs). We ask if pre-trained VLMs can aid scientists in answering a range of biologically relevant questions without any additional fine-tuning. In this paper, we evaluate the effectiveness of $12$ state-of-the-art (SOTA) VLMs in the field of organismal biology using a novel dataset, VLM4Bio, consisting of $469K$ question-answer pairs involving $30K$ images from three groups of organisms: fishes, birds, and butterflies, covering five biologically relevant tasks. We also explore the effects of applying prompting techniques and tests for reasoning hallucination on the performance of VLMs, shedding new light on the capabilities of current SOTA VLMs in answering biologically relevant questions using images.
M. Maruf, Arka Daw, Kazi Sajeed Mehrab, Harish Babu Manogaran, Abhilash Neog, Medha Sawhney, Mridul Khurana, James P. Balhoff, Yasin Bakis, Bahadir Altintas, Matthew J. Thompson, Elizabeth G. Campolongo, Josef C. Uyeda, Hilmar Lapp, Henry L. Bart Jr., Paula M. Mabee, Yu Su 0001, Wei-Lun Chao, Charles V. Stewart, Tanya Y. Berger-Wolf, Wasila M. Dahdul, Anuj Karpatne
NeurIPS15
2023 Discovering Novel Biological Traits From Images Using Phylogeny-Guided Neural Networks
abstract
Discovering evolutionary traits that are heritable across species on the tree of life (also referred to as a phylogenetic tree) is of great interest to biologists to understand how organisms diversify and evolve. However, the measurement of traits is often a subjective and labor-intensive process, making trait discovery a highly label-scarce problem. We present a novel approach for discovering evolutionary traits directly from images without relying on trait labels. Our proposed approach, Phylo-NN, encodes the image of an organism into a sequence of quantized feature vectors -or codes- where different segments of the sequence capture evolutionary signals at varying ancestry levels in the phylogeny. We demonstrate the effectiveness of our approach in producing biologically meaningful results in a number of downstream tasks including species image generation and species-to-species image translation, using fish species as a target example
Mohannad Elhamod, Mridul Khurana, Harish Babu Manogaran, Josef C. Uyeda, Meghan A. Balk, Wasila M. Dahdul, Yasin Bakis, Henry L. Bart Jr., Paula M. Mabee, Hilmar Lapp, James P. Balhoff, Caleb Charpentier, David Carlyn, Wei-Lun Chao, Charles V. Stewart, Daniel I. Rubenstein, Tanya Y. Berger-Wolf, Anuj Karpatne
KDD8
2015 Robust Model-Based Learning via Spatial-EM Algorithm
abstract
This paper presents a new robust EM algorithm for the finite mixture learning procedures. The proposed Spatial-EM algorithm utilizes median-based location and rank-based scatter estimators to replace sample mean and sample covariance matrix in each M step, hence enhancing stability and robustness of the algorithm. It is robust to outliers and initial values. Compared with many robust mixture learning methods, the Spatial-EM has the advantages of simplicity in implementation and statistical efficiency. We apply Spatial-EM to supervised and unsupervised learning scenarios. More specifically, robust clustering and outlier detection methods based on Spatial-EM have been proposed. We apply the outlier detection to taxonomic research on fish species novelty discovery. Two real datasets are used for clustering analysis. Compared with the regular EM and many other existing methods such as K-median, X-EM and SVM, our method demonstrates superior performance and high robustness.
Xin Dang, Henry L. Bart Jr., Yixin Chen 0002
IEEE Trans. Knowl. Data Eng.3
2010 Joint feature selection and classification for taxonomic problems within fish species complexes
Yixin Chen 0002, Shuqing Huang, Henry L. Bart Jr.
Pattern Anal. Appl.4
2009 Outlier Detection with the Kernelized Spatial Depth Function
abstract
Statistical depth functions provide from the "deepest" point a "center-outward ordering" of multidimensional data. In this sense, depth functions can measure the "extremeness" or "outlyingness" of a data point with respect to a given data set. Hence, they can detect outliers--observations that appear extreme relative to the rest of the observations. Of the various statistical depths, the spatial depth is especially appealing because of its computational efficiency and mathematical tractability. In this article, we propose a novel statistical depth, the kernelized spatial depth (KSD), which generalizes the spatial depth via positive definite kernels. By choosing a proper kernel, the KSD can capture the local structure of a data set while the spatial depth fails. We demonstrate this by the half-moon data and the ring-shaped data. Based on the KSD, we propose a novel outlier detection algorithm, by which an observation with a depth value less than a threshold is declared as an outlier. The proposed algorithm is simple in structure: the threshold is the only one parameter for a given kernel. It applies to a one-class learning setting, in which "normal" observations are given as the training data, as well as to a missing label scenario, where the training set consists of a mixture of normal observations and outliers with unknown labels. We give upper bounds on the false alarm probability of a depth-based detector. These upper bounds can be used to determine the threshold. We perform extensive experiments on synthetic data and data sets from real applications. The proposed outlier detector is compared with existing methods. The KSD outlier detector demonstrates a competitive performance.
Yixin Chen 0002, Xin Dang, Hanxiang Peng, Henry L. Bart Jr.
IEEE Trans. Pattern Anal. Mach. Intell.4
2007 Depth-Based Novelty Detection and Its Application to Taxonomic Research
abstract
It is estimated that less than 10 percent of the world's species have been described, yet species are being lost daily due to human destruction of natural habitats. The job of describing the earth's remaining species is exacerbated by the shrinking number of practicing taxonomists and the very slow pace of traditional taxonomic research. In this article, we tackle, from a novelty detection perspective, one of the most important and challenging research objectives in taxonomy ­ new species identification. We propose a unique and efficient novelty detection framework based on statistical depth functions. Statistical depth functions provide from the "deepest" point a "center-outward ordering" of multidimensional data. In this sense, they can detect observations that appear extreme relative to the rest of the observations, i.e., novelty. Of the various statistical depths, the spatial depth is especially appealing because of its computational efficiency and mathematical tractability. We propose a novel statistical depth, the kernelized spatial depth (KSD) that generalizes the spatial depth via positive definite kernels. By choosing a proper kernel, the KSD can capture the local structure of a data set while the spatial depth fails. Observations with depth values less than a threshold are declared as novel. The proposed algorithm is simple in structure: the threshold is the only one parameter for a given kernel. We give an upper bound on the false alarm probability of a depth-based detector, which can be used to determine the threshold. Experimental study demonstrates its excellent potential in new species discovery.
Yixin Chen 0002, Henry L. Bart Jr., Xin Dang, Hanxiang Peng
ICDM2
2007 Integrated Feature Selection and Clustering from Multiple Views for a Taxonomic Problem
abstract
As computer and database technologies advance rapidly, biologists all over the world can share biologically meaningful data from images of specimens and use the data to classify the specimens taxonomically. Accurate shape analysis of a specimen from multiple views of 2D images is crucial for finding diagnostic features using geometric morphometric techniques. We propose an integrated feature selection and clustering framework that automatically identifies a set of feature variables to group specimens into a binary cluster tree. The candidate features are generated from reconstructed 3D shape and local saliency characteristics from 2D images of the specimen. We use a mixture model to estimate the significance value of each feature and control the false discovery rate in the feature selection process so that the clustering algorithm can efficiently partition the specimen samples into clusters that may correspond to different species. The experiments on a taxonomic problem involving species of suckers in the genus Carpiodes demonstrate promising results using the proposed framework with small sample size.
Henry L. Bart Jr., Shuqing Huang
MMSP2
2006 Taxonomy in Fish Species Complexes: A Role for Multimedia Information
abstract
Biologists could make valuable use of the wealth of specimen information in natural history museum databases. "Taxonomy via the Internet" aims to build a centralized database where biologists can store, manipulate and retrieve biologically meaningful data from images of specimens and use the data to classify the specimens taxonomically. Multimedia information representation provides a new computational tool for extracting useful features from large databases of specimen images and has potential to expedite the pace of taxonomic research. In this paper, we use a taxonomic problem involving species of suckers in the genus Carpiodes to demonstrate the utility of this method. Logistic regression classifier with fully automated feature selection procedure is compared with the best landmark based classifier to illustrate how image quality affects classification accuracy. We discuss the need of creating a multimedia database using images of specimens from a fish collection
Shuqing Huang, Henry L. Bart Jr.
MMSP3
2005 A Computational Framework for Taxonomic Research: Diagnosing Body Shape within Fish Species Complexes
abstract
It is estimated that ninety percent of the world's species have yet to be discovered and described. The main reason for the slow pace of new species description is that the science of taxonomy, as traditionally practiced, can be very laborious. To formally describe a new species, taxonomists have to manually gather and analyze data from large numbers of specimens, often from broad geographic areas, and identify the smallest subset of external body characters that uniquely diagnoses the new species as distinct from all its known relatives. In this paper, we use an automated feature selection and classification approach to address the taxonomic impediment in new species discovery. The experiments on a taxonomic problem involving species of suckers in the genus Carpiodes demonstrate promising results.
Yixin Chen 0002, Henry L. Bart Jr., Shuqing Huang
ICDM2