David Carlyn

dblp:344/1011 · also David Edward Carlyn · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0002-8323-0359ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 60% Image recognition and object detection · 19% Representation and self-supervised learning · 14%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 77% Rendering · 23%

Topics — the 9 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
1.622025
Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation · CVPR 2025
A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis · ICLR 2024
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
1.022025
A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis · ICLR 2024
Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation · CVPR 2025
Machine learning › Trustworthy machine learning › interpretability › visual explanation
class activation map
0.912025
Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation · CVPR 2025
Machine learning › Trustworthy machine learning › interpretability › attention analysis
attention-based explanation
0.812024
A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis · ICLR 2024
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
vision foundation model
0.812024
BioCLIP: A Vision Foundation Model for the Tree of Life · CVPR 2024
Bioinformatics and computational biology
phylogenetics
0.712023
Discovering Novel Biological Traits From Images Using Phylogeny-Guided Neural Networks · KDD 2023
Machine learning › Deep learning architectures and training
transformer
0.212024
A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis · ICLR 2024
Machine learning › Generative modeling › generative adversarial network
image-to-image translation
0.212023
Discovering Novel Biological Traits From Images Using Phylogeny-Guided Neural Networks · KDD 2023
Rendering
inverse rendering
0.212023
Learning Fractals by Gradient Descent · AAAI 2023

Methods — techniques the papers use, named apart from their topics

fine-grained classification · 1.5contrastive learning · 1.5quantization · 1.3phylogeny encoding · 1.3neural network · 1.3discriminative region localization · 0.9class activation mapping · 0.9cross-attention · 0.8class-specific queries · 0.8loss function optimization · 0.7gradient descent · 0.7
YearPublicationVenuePosition
2025 Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation
abstract
Class activation map (CAM) has been widely used to highlight image regions that contribute to class predictions. Despite its simplicity and computational efficiency, CAM often struggles to identify discriminative regions that distinguish visually similar fine-grained classes. Prior efforts address this limitation by introducing more sophisticated explanation processes, but at the cost of extra complexity. In this paper, we propose Finer-CAM, a method that retains CAM’s efficiency while achieving precise localization of discriminative regions. Our key insight is that the deficiency of CAM lies not in "how" it explains, but in "what" it explains. Specifically, previous methods attempt to identify all cues contributing to the target class’s logit value, which inadvertently also activates regions predictive of visually similar classes. By explicitly comparing the target class with similar classes and spotting their differences, Finer-CAM suppresses features shared with other classes and emphasizes the unique, discriminative details of the target class. Finer-CAM is easy to implement, compatible with various CAM methods, and can be extended to multi-modal models for accurate localization of specific concepts. Additionally, Finer-CAM allows adjustable comparison strength, enabling users to selectively highlight coarse object contours or fine discriminative details. Quantitatively, we show that masking out the top 5% of activated pixels by Finer-CAM results in a larger relative confidence drop compared to baselines. The source code and demo are available at https://github.com/Imageomics/Finer-CAM.
Jianyang Gu, Arpita Chowdhury, Zheda Mai, David Carlyn, Tanya Y. Berger-Wolf, Yu Su 0001, Wei-Lun Chao
CVPR5
2024 BioCLIP: A Vision Foundation Model for the Tree of Life
abstract
Images of the natural world, collected by a variety of cameras, from drones to individual phones, are increasingly abundant sources of biological information. There is an ex-plosion of computational methods and tools, particularly computer vision, for extracting biologically relevant information from images for science and conservation. Yet most of these are bespoke approaches designed for a specific task and are not easily adaptable or extendable to new questions, contexts, and datasets. A vision model for general or-ganismal biology questions on images is of timely need. To approach this, we curate and release Tree Of Life-10m, the largest and most diverse ML-ready dataset of biology images. We then develop Bioclip, a foundation model for the tree of life, leveraging the unique properties of bi-ology captured by Treeoflife-10m, namely the abun-dance and variety of images of plants, animals, and fungi, together with the availability of rich structured biological knowledge. We rigorously benchmark our approach on di-verse fine-grained biology classification tasks and find that BloCLIP consistently and substantially outperforms existing baselines (by 16% to 17% absolute). Intrinsic evaluation reveals that BloCLIP has learned a hierarchical representation conforming to the tree of life, shedding light on its strong generalizability.11imageomics.github.io/bioclip has models, data and code.
Samuel Stevens 0001, Jiaman Wu, Matthew J. Thompson, Elizabeth G. Campolongo, Chan Hee Song, David Carlyn, Wasila M. Dahdul, Charles V. Stewart, Tanya Y. Berger-Wolf, Wei-Lun Chao, Yu Su 0001
CVPR6
2024 A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis
abstract
We present a novel usage of Transformers to make image classification interpretable. Unlike mainstream classifiers that wait until the last fully connected layer to incorporate class information to make predictions, we investigate a proactive approach, asking each class to search for itself in an image. We realize this idea via a Transformer encoder-decoder inspired by DEtection TRansformer (DETR). We learn ''class-specific'' queries (one for each class) as input to the decoder, enabling each class to localize its patterns in an image via cross-attention. We name our approach INterpretable TRansformer (INTR), which is fairly easy to implement and exhibits several compelling properties. We show that INTR intrinsically encourages each class to attend distinctively; the cross-attention weights thus provide a faithful interpretation of the prediction. Interestingly, via ''multi-head'' cross-attention, INTR could identify different ''attributes'' of a class, making it particularly suitable for fine-grained classification and analysis, which we demonstrate on eight datasets. Our code and pre-trained models are publicly accessible at the Imageomics Institute GitHub site: https://github.com/Imageomics/INTR.
Dipanjyoti Paul, Arpita Chowdhury, Xinqi Xiong, Feng-Ju Chang, David Carlyn, Samuel Stevens 0001, Kaiya Provost, Anuj Karpatne, Bryan Carstens, Daniel I. Rubenstein, Charles V. Stewart, Tanya Y. Berger-Wolf, Yu Su 0001, Wei-Lun Chao
ICLR5
2023 Learning Fractals by Gradient Descent
abstract
Fractals are geometric shapes that can display complex and self-similar patterns found in nature (e.g., clouds and plants). Recent works in visual recognition have leveraged this property to create random fractal images for model pre-training. In this paper, we study the inverse problem --- given a target image (not necessarily a fractal), we aim to generate a fractal image that looks like it. We propose a novel approach that learns the parameters underlying a fractal image via gradient descent. We show that our approach can find fractal parameters of high visual quality and be compatible with different loss functions, opening up several potentials, e.g., learning fractals for downstream tasks, scientific understanding, etc.
Cheng-Hao Tu 0001, Hong-You Chen, David Carlyn, Wei-Lun Chao
AAAI3
2023 Discovering Novel Biological Traits From Images Using Phylogeny-Guided Neural Networks
abstract
Discovering evolutionary traits that are heritable across species on the tree of life (also referred to as a phylogenetic tree) is of great interest to biologists to understand how organisms diversify and evolve. However, the measurement of traits is often a subjective and labor-intensive process, making trait discovery a highly label-scarce problem. We present a novel approach for discovering evolutionary traits directly from images without relying on trait labels. Our proposed approach, Phylo-NN, encodes the image of an organism into a sequence of quantized feature vectors -or codes- where different segments of the sequence capture evolutionary signals at varying ancestry levels in the phylogeny. We demonstrate the effectiveness of our approach in producing biologically meaningful results in a number of downstream tasks including species image generation and species-to-species image translation, using fish species as a target example
Mohannad Elhamod, Mridul Khurana, Harish Babu Manogaran, Josef C. Uyeda, Meghan A. Balk, Wasila M. Dahdul, Yasin Bakis, Henry L. Bart Jr., Paula M. Mabee, Hilmar Lapp, James P. Balhoff, Caleb Charpentier, David Carlyn, Wei-Lun Chao, Charles V. Stewart, Daniel I. Rubenstein, Tanya Y. Berger-Wolf, Anuj Karpatne
KDD13