EDBT 2026 Demo / reviewers in the wild / expert
Henry L. Bart Jr.
dblp:56/399
· DBLP profile ↗
11ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0002-5662-9444ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Generative modeling · 36% Vision and language · 21% Segmentation and scene understanding · 18% | |
| Interdisciplinary, comprehensive, and emerging computing
6 papers |
Bioinformatics and computational biology · 79% Environmental and earth informatics · 21% | |
| Databases, data mining, and information retrieval
3 papers |
Data mining · 100% |
Topics — the 24 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Environmental and earth informatics
biodiversity informatics |
0.9 | 1 | 2025 | Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from Images · CVPR 2025 |
Bioinformatics and computational biology
species classification |
0.9 | 1 | 2025 | Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from Images · CVPR 2025 |
Machine learning › Generative modeling › diffusion model
conditional diffusion model |
0.8 | 1 | 2024 | Hierarchical Conditioning of Diffusion Models Using Tree-of-Life for Studying Species Evolution · ECCV (89) 2024 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Hierarchical Conditioning of Diffusion Models Using Tree-of-Life for Studying Species Evolution · ECCV (89) 2024 |
Computer vision › Vision and language › vision-language model
vision-language model evaluation |
0.8 | 1 | 2024 | VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images · NeurIPS 2024 |
Bioinformatics and computational biology
evolutionary biology |
0.8 | 1 | 2024 | Hierarchical Conditioning of Diffusion Models Using Tree-of-Life for Studying Species Evolution · ECCV (89) 2024 |
Bioinformatics and computational biology
phylogenetics |
0.7 | 1 | 2023 | Discovering Novel Biological Traits From Images Using Phylogeny-Guided Neural Networks · KDD 2023 |
Data mining
anomaly detection |
0.3 | 2 | 2015 | Robust Model-Based Learning via Spatial-EM Algorithm · IEEE Trans. Knowl. Data Eng. 2015 Depth-Based Novelty Detection and Its Application to Taxonomic Research · ICDM 2007 |
Machine learning › Trustworthy machine learning › interpretability
explainable AI |
0.3 | 1 | 2025 | Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from Images · CVPR 2025 |
Machine learning › Learning paradigms
long-tailed recognition |
0.3 | 1 | 2025 | Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from Images · CVPR 2025 |
Computer vision › Vision and language › vision-language model
pre-trained vision-language model |
0.2 | 1 | 2024 | VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › hallucination
vision-language model hallucination |
0.2 | 1 | 2024 | VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
mixture model learning |
0.2 | 1 | 2015 | Robust Model-Based Learning via Spatial-EM Algorithm · IEEE Trans. Knowl. Data Eng. 2015 |
Data mining
clustering |
0.2 | 1 | 2015 | Robust Model-Based Learning via Spatial-EM Algorithm · IEEE Trans. Knowl. Data Eng. 2015 |
Data mining › anomaly detection
outlier detection |
0.2 | 1 | 2015 | Robust Model-Based Learning via Spatial-EM Algorithm · IEEE Trans. Knowl. Data Eng. 2015 |
Data mining › clustering
robust clustering |
0.2 | 1 | 2015 | Robust Model-Based Learning via Spatial-EM Algorithm · IEEE Trans. Knowl. Data Eng. 2015 |
Machine learning › Generative modeling › generative adversarial network
image-to-image translation |
0.2 | 1 | 2023 | Discovering Novel Biological Traits From Images Using Phylogeny-Guided Neural Networks · KDD 2023 |
Bioinformatics and computational biology
species identification |
0.1 | 2 | 2007 | Depth-Based Novelty Detection and Its Application to Taxonomic Research · ICDM 2007 A Computational Framework for Taxonomic Research: Diagnosing Body Shape within Fish Species Complexes · ICDM 2005 |
Bioinformatics and computational biology
taxonomy |
0.1 | 2 | 2007 | Depth-Based Novelty Detection and Its Application to Taxonomic Research · ICDM 2007 A Computational Framework for Taxonomic Research: Diagnosing Body Shape within Fish Species Complexes · ICDM 2005 |
Machine learning › Time series and sequential data
anomaly detection |
0.1 | 1 | 2009 | Outlier Detection with the Kernelized Spatial Depth Function · IEEE Trans. Pattern Anal. Mach. Intell. 2009 |
Machine learning › Trustworthy machine learning
statistical depth |
0.1 | 1 | 2009 | Outlier Detection with the Kernelized Spatial Depth Function · IEEE Trans. Pattern Anal. Mach. Intell. 2009 |
Data mining › anomaly detection
novelty detection |
0.1 | 1 | 2007 | Depth-Based Novelty Detection and Its Application to Taxonomic Research · ICDM 2007 |
Data mining › predictive modeling
classification |
0.1 | 1 | 2005 | A Computational Framework for Taxonomic Research: Diagnosing Body Shape within Fish Species Complexes · ICDM 2005 |
Data mining › dimensionality reduction
feature selection |
0.1 | 1 | 2005 | A Computational Framework for Taxonomic Research: Diagnosing Body Shape within Fish Species Complexes · ICDM 2005 |
Methods — techniques the papers use, named apart from their topics
machine learning · 1.7computer vision · 1.7zero-shot evaluation · 1.5prompting techniques · 1.5phylogenetic embeddings · 1.5diffusion model · 1.5quantization · 1.3phylogeny encoding · 1.3neural network · 1.3median-based location estimation · 0.4rank-based scatter estimation · 0.2kernelized spatial depth · 0.1statistical depth functions · 0.1classification · 0.1automated feature selection · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from ImagesabstractWe introduce Fish-Visual Trait Analysis (Fish-Vista), the first organismal image dataset designed for the analysis of visual traits of aquatic species directly from images using machine learning and computer vision methods. Fish-Vista contains 69,269 annotated images spanning 4,316 fish species, curated and organized to serve three downstream tasks: species classification, trait identification, and trait segmentation. Our work makes two key contributions. First, we provide a fully reproducible data processing pipeline to process fish images sourced from various museum collections, contributing to the advancement of AI in biodiversity science. We annotate the images with carefully curated labels from biological databases and manual annotations to create an AI-ready dataset of visual traits. Second, our work offers fertile grounds for researchers to develop novel methods for a variety of problems in computer vision such as handling long-tailed distributions, out-of-distribution generalization, learning with weak labels, explainable AI, and segmenting small objects. Dataset and code for Fish-Vista are available at https://github.com/Imageomics/Fish-Vista Kazi Sajeed Mehrab, M. Maruf, Arka Daw, Abhilash Neog, Harish Babu Manogaran, Mridul Khurana, Zhenyang Feng, Bahadir Altintas, Yasin Bakis, Elizabeth G. Campolongo, Matthew J. Thompson, Hilmar Lapp, Tanya Y. Berger-Wolf, Paula M. Mabee, Henry L. Bart Jr., Wei-Lun Chao, Wasila M. Dahdul, Anuj Karpatne |
CVPR | 16 |
| 2024 | Hierarchical Conditioning of Diffusion Models Using Tree-of-Life for Studying Species EvolutionabstractAbstract A central problem in biology is to understand how organisms evolve and adapt to their environment by acquiring variations in the observable characteristics or traits of species across the tree of life. With the growing availability of large-scale image repositories in biology and recent advances in generative modeling, there is an opportunity to accelerate the discovery of evolutionary traits automatically from images. Toward this goal, we introduce Phylo-Diffusion, a novel framework for conditioning diffusion models with phylogenetic knowledge represented in the form of HIERarchical Embeddings (HIER-Embeds). We also propose two new experiments for perturbing the embedding space of Phylo-Diffusion: trait masking and trait swapping, inspired by counterpart experiments of gene knockout and gene editing/swapping. Our work represents a novel methodological advance in generative modeling to structure the embedding space of diffusion models using tree-based knowledge. Our work also opens a new chapter of research in evolutionary biology by using generative models to visualize evolutionary changes directly from images. We empirically demonstrate the usefulness of Phylo-Diffusion in capturing meaningful trait variations for fishes and birds, revealing novel insights about the biological mechanisms of their evolution. (Model and code can be found at imageomics.github.io/phylo-diffusion ) Mridul Khurana, Arka Daw, M. Maruf, Josef C. Uyeda, Wasila M. Dahdul, Caleb Charpentier, Yasin Bakis, Henry L. Bart Jr., Paula M. Mabee, Hilmar Lapp, James P. Balhoff, Wei-Lun Chao, Charles V. Stewart, Tanya Y. Berger-Wolf, Anuj Karpatne |
ECCV (89) | 8 |
| 2024 | VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological ImagesabstractImages are increasingly becoming the currency for documenting biodiversity on the planet, providing novel opportunities for accelerating scientific discoveries in the field of organismal biology, especially with the advent of large vision-language models (VLMs). We ask if pre-trained VLMs can aid scientists in answering a range of biologically relevant questions without any additional fine-tuning. In this paper, we evaluate the effectiveness of $12$ state-of-the-art (SOTA) VLMs in the field of organismal biology using a novel dataset, VLM4Bio, consisting of $469K$ question-answer pairs involving $30K$ images from three groups of organisms: fishes, birds, and butterflies, covering five biologically relevant tasks. We also explore the effects of applying prompting techniques and tests for reasoning hallucination on the performance of VLMs, shedding new light on the capabilities of current SOTA VLMs in answering biologically relevant questions using images. M. Maruf, Arka Daw, Kazi Sajeed Mehrab, Harish Babu Manogaran, Abhilash Neog, Medha Sawhney, Mridul Khurana, James P. Balhoff, Yasin Bakis, Bahadir Altintas, Matthew J. Thompson, Elizabeth G. Campolongo, Josef C. Uyeda, Hilmar Lapp, Henry L. Bart Jr., Paula M. Mabee, Yu Su 0001, Wei-Lun Chao, Charles V. Stewart, Tanya Y. Berger-Wolf, Wasila M. Dahdul, Anuj Karpatne |
NeurIPS | 15 |
| 2023 | Discovering Novel Biological Traits From Images Using Phylogeny-Guided Neural NetworksabstractDiscovering evolutionary traits that are heritable across species on the tree of life (also referred to as a phylogenetic tree) is of great interest to biologists to understand how organisms diversify and evolve. However, the measurement of traits is often a subjective and labor-intensive process, making trait discovery a highly label-scarce problem. We present a novel approach for discovering evolutionary traits directly from images without relying on trait labels. Our proposed approach, Phylo-NN, encodes the image of an organism into a sequence of quantized feature vectors -or codes- where different segments of the sequence capture evolutionary signals at varying ancestry levels in the phylogeny. We demonstrate the effectiveness of our approach in producing biologically meaningful results in a number of downstream tasks including species image generation and species-to-species image translation, using fish species as a target example Mohannad Elhamod, Mridul Khurana, Harish Babu Manogaran, Josef C. Uyeda, Meghan A. Balk, Wasila M. Dahdul, Yasin Bakis, Henry L. Bart Jr., Paula M. Mabee, Hilmar Lapp, James P. Balhoff, Caleb Charpentier, David Carlyn, Wei-Lun Chao, Charles V. Stewart, Daniel I. Rubenstein, Tanya Y. Berger-Wolf, Anuj Karpatne |
KDD | 8 |
| 2015 | Robust Model-Based Learning via Spatial-EM AlgorithmabstractThis paper presents a new robust EM algorithm for the finite mixture learning procedures. The proposed Spatial-EM algorithm utilizes median-based location and rank-based scatter estimators to replace sample mean and sample covariance matrix in each M step, hence enhancing stability and robustness of the algorithm. It is robust to outliers and initial values. Compared with many robust mixture learning methods, the Spatial-EM has the advantages of simplicity in implementation and statistical efficiency. We apply Spatial-EM to supervised and unsupervised learning scenarios. More specifically, robust clustering and outlier detection methods based on Spatial-EM have been proposed. We apply the outlier detection to taxonomic research on fish species novelty discovery. Two real datasets are used for clustering analysis. Compared with the regular EM and many other existing methods such as K-median, X-EM and SVM, our method demonstrates superior performance and high robustness. Xin Dang, Henry L. Bart Jr., Yixin Chen 0002 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2010 | Joint feature selection and classification for taxonomic problems within fish species complexes
Yixin Chen 0002, Shuqing Huang, Henry L. Bart Jr. |
Pattern Anal. Appl. | 4 |
| 2009 | Outlier Detection with the Kernelized Spatial Depth FunctionabstractStatistical depth functions provide from the "deepest" point a "center-outward ordering" of multidimensional data. In this sense, depth functions can measure the "extremeness" or "outlyingness" of a data point with respect to a given data set. Hence, they can detect outliers--observations that appear extreme relative to the rest of the observations. Of the various statistical depths, the spatial depth is especially appealing because of its computational efficiency and mathematical tractability. In this article, we propose a novel statistical depth, the kernelized spatial depth (KSD), which generalizes the spatial depth via positive definite kernels. By choosing a proper kernel, the KSD can capture the local structure of a data set while the spatial depth fails. We demonstrate this by the half-moon data and the ring-shaped data. Based on the KSD, we propose a novel outlier detection algorithm, by which an observation with a depth value less than a threshold is declared as an outlier. The proposed algorithm is simple in structure: the threshold is the only one parameter for a given kernel. It applies to a one-class learning setting, in which "normal" observations are given as the training data, as well as to a missing label scenario, where the training set consists of a mixture of normal observations and outliers with unknown labels. We give upper bounds on the false alarm probability of a depth-based detector. These upper bounds can be used to determine the threshold. We perform extensive experiments on synthetic data and data sets from real applications. The proposed outlier detector is compared with existing methods. The KSD outlier detector demonstrates a competitive performance. Yixin Chen 0002, Xin Dang, Hanxiang Peng, Henry L. Bart Jr. |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2007 | Depth-Based Novelty Detection and Its Application to Taxonomic ResearchabstractIt is estimated that less than 10 percent of the world's species have been described, yet species are being lost daily due to human destruction of natural habitats. The job of describing the earth's remaining species is exacerbated by the shrinking number of practicing taxonomists and the very slow pace of traditional taxonomic research. In this article, we tackle, from a novelty detection perspective, one of the most important and challenging research objectives in taxonomy new species identification. We propose a unique and efficient novelty detection framework based on statistical depth functions. Statistical depth functions provide from the "deepest" point a "center-outward ordering" of multidimensional data. In this sense, they can detect observations that appear extreme relative to the rest of the observations, i.e., novelty. Of the various statistical depths, the spatial depth is especially appealing because of its computational efficiency and mathematical tractability. We propose a novel statistical depth, the kernelized spatial depth (KSD) that generalizes the spatial depth via positive definite kernels. By choosing a proper kernel, the KSD can capture the local structure of a data set while the spatial depth fails. Observations with depth values less than a threshold are declared as novel. The proposed algorithm is simple in structure: the threshold is the only one parameter for a given kernel. We give an upper bound on the false alarm probability of a depth-based detector, which can be used to determine the threshold. Experimental study demonstrates its excellent potential in new species discovery. Yixin Chen 0002, Henry L. Bart Jr., Xin Dang, Hanxiang Peng |
ICDM | 2 |
| 2007 | Integrated Feature Selection and Clustering from Multiple Views for a Taxonomic ProblemabstractAs computer and database technologies advance rapidly, biologists all over the world can share biologically meaningful data from images of specimens and use the data to classify the specimens taxonomically. Accurate shape analysis of a specimen from multiple views of 2D images is crucial for finding diagnostic features using geometric morphometric techniques. We propose an integrated feature selection and clustering framework that automatically identifies a set of feature variables to group specimens into a binary cluster tree. The candidate features are generated from reconstructed 3D shape and local saliency characteristics from 2D images of the specimen. We use a mixture model to estimate the significance value of each feature and control the false discovery rate in the feature selection process so that the clustering algorithm can efficiently partition the specimen samples into clusters that may correspond to different species. The experiments on a taxonomic problem involving species of suckers in the genus Carpiodes demonstrate promising results using the proposed framework with small sample size. Henry L. Bart Jr., Shuqing Huang |
MMSP | 2 |
| 2006 | Taxonomy in Fish Species Complexes: A Role for Multimedia InformationabstractBiologists could make valuable use of the wealth of specimen information in natural history museum databases. "Taxonomy via the Internet" aims to build a centralized database where biologists can store, manipulate and retrieve biologically meaningful data from images of specimens and use the data to classify the specimens taxonomically. Multimedia information representation provides a new computational tool for extracting useful features from large databases of specimen images and has potential to expedite the pace of taxonomic research. In this paper, we use a taxonomic problem involving species of suckers in the genus Carpiodes to demonstrate the utility of this method. Logistic regression classifier with fully automated feature selection procedure is compared with the best landmark based classifier to illustrate how image quality affects classification accuracy. We discuss the need of creating a multimedia database using images of specimens from a fish collection Shuqing Huang, Henry L. Bart Jr. |
MMSP | 3 |
| 2005 | A Computational Framework for Taxonomic Research: Diagnosing Body Shape within Fish Species ComplexesabstractIt is estimated that ninety percent of the world's species have yet to be discovered and described. The main reason for the slow pace of new species description is that the science of taxonomy, as traditionally practiced, can be very laborious. To formally describe a new species, taxonomists have to manually gather and analyze data from large numbers of specimens, often from broad geographic areas, and identify the smallest subset of external body characters that uniquely diagnoses the new species as distinct from all its known relatives. In this paper, we use an automated feature selection and classification approach to address the taxonomic impediment in new species discovery. The experiments on a taxonomic problem involving species of suckers in the genus Carpiodes demonstrate promising results. Yixin Chen 0002, Henry L. Bart Jr., Shuqing Huang |
ICDM | 2 |