Kelly O. Marshall

dblp:349/4918 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0005-9771-947XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Generative modeling · 32% Image recognition and object detection · 32% Trustworthy machine learning · 16%
Computer graphics and multimedia
1 paper
Computational fabrication · 77% Multimedia analysis and retrieval · 23%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Environmental and earth informatics · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
code generation
0.812024
Slice-100K: A Multimodal Dataset for Extrusion-based 3D Printing · NeurIPS 2024
Machine learning › Generative modeling
concept erasure
0.812024
Circumventing Concept Erasure Methods For Text-To-Image Generative Models · ICLR 2024
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
0.812024
BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity · NeurIPS 2024
Machine learning › Trustworthy machine learning
robustness
0.812024
Circumventing Concept Erasure Methods For Text-To-Image Generative Models · ICLR 2024
Computer vision › Image recognition and object detection › image classification › fine-grained image classification
species recognition
0.812024
BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity · NeurIPS 2024
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.812024
Circumventing Concept Erasure Methods For Text-To-Image Generative Models · ICLR 2024
Environmental and earth informatics
biodiversity informatics
0.812024
BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity · NeurIPS 2024
Computational fabrication
additive manufacturing
0.812024
Slice-100K: A Multimodal Dataset for Extrusion-based 3D Printing · NeurIPS 2024
Computer vision › Vision and language
vision-language pretraining
0.212024
BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity · NeurIPS 2024
Multimedia analysis and retrieval › multimedia dataset construction
multimodal dataset
0.212024
Slice-100K: A Multimodal Dataset for Extrusion-based 3D Printing · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

zero-shot learning · 1.5fine-tuning · 1.5GPT-2 · 1.5CLIP · 1.5learned word embeddings · 0.8
YearPublicationVenuePosition
2025 PITCH: AI-assisted Tagging of Deepfake Audio Calls using Challenge-Response
Govind Mittal, Arthur Jakobsson, Kelly O. Marshall, Chinmay Hegde, Nasir Memon
AsiaCCS3
2024 Circumventing Concept Erasure Methods For Text-To-Image Generative Models
abstract
Text-to-image generative models can produce photo-realistic images for an extremely broad range of concepts, and their usage has proliferated widely among the general public. On the flip side, these models have numerous drawbacks, including their potential to generate images featuring sexually explicit content, mirror artistic styles without permission, or even hallucinate (or deepfake) the likenesses of celebrities. Consequently, various methods have been proposed in order to "erase" sensitive concepts from text-to-image models. In this work, we examine seven recently proposed concept erasure methods, and show that targeted concepts are not fully excised from any of these methods. Specifically, we leverage the existence of special learned word embeddings that can retrieve "erased" concepts from the sanitized models with no alterations to their weights. Our results highlight the brittleness of post hoc concept erasure methods, and call into question their use in the algorithmic toolkit for AI safety.
Minh Pham 0005, Kelly O. Marshall, Niv Cohen, Govind Mittal, Chinmay Hegde
ICLR2
2024 Slice-100K: A Multimodal Dataset for Extrusion-based 3D Printing
abstract
G-code (Geometric code) or RS-274 is the most widely used computer numerical control (CNC) and 3D printing programming language. G-code provides machine instructions for the movement of the 3D printer, especially for the nozzle, stage, and extrusion of material for extrusion-based additive manufacturing. Currently, there does not exist a large repository of curated CAD models along with their corresponding G-code files for additive manufacturing. To address this issue, we present Slice-100K, a first-of-its-kind dataset of over 100,000 G-code files, along with their tessellated CAD model, LVIS (Large Vocabulary Instance Segmentation) categories, geometric properties, and renderings. We build our dataset from triangulated meshes derived from Objaverse-XL and Thingi10K datasets. We demonstrate the utility of this dataset by finetuning GPT-2 on a subset of the dataset for G-code translation from a legacy G-code format (Sailfish) to a more modern, widely used format (Marlin). Our dataset can be found here. Slice-100K will be the first step in developing a multimodal foundation model for digital manufacturing.
Anushrut Jignasu, Kelly O. Marshall, Ankush Kumar Mishra, Lucas Nerone Rillo, Baskar Ganapathysubramanian, Aditya Balu, Chinmay Hegde, Adarsh Krishnamurthy
NeurIPS2
2024 BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity
abstract
We introduce BioTrove, the largest publicly accessible dataset designed to advance AI applications in biodiversity. Curated from the iNaturalist platform and vetted to include only research-grade data, BioTrove contains 161.9 million images, offering unprecedented scale and diversity from three primary kingdoms: Animalia ("animals"), Fungi ("fungi"), and Plantae ("plants"), spanning approximately 366.6K species. Each image is annotated with scientific names, taxonomic hierarchies, and common names, providing rich metadata to support accurate AI model development across diverse species and ecosystems.We demonstrate the value of BioTrove by releasing a suite of CLIP models trained using a subset of 40 million captioned images, known as BioTrove-Train. This subset focuses on seven categories within the dataset that are underrepresented in standard image recognition models, selected for their critical role in biodiversity and agriculture: Aves ("birds"), Arachnida} ("spiders/ticks/mites"), Insecta ("insects"), Plantae ("plants"), Fungi ("fungi"), Mollusca ("snails"), and Reptilia ("snakes/lizards"). To support rigorous assessment, we introduce several new benchmarks and report model accuracy for zero-shot learning across life stages, rare species, confounding species, and multiple taxonomic levels.We anticipate that BioTrove will spur the development of AI models capable of supporting digital tools for pest control, crop monitoring, biodiversity assessment, and environmental conservation. These advancements are crucial for ensuring food security, preserving ecosystems, and mitigating the impacts of climate change. BioTrove is publicly available, easily accessible, and ready for immediate use.
Chih-Hsuan Yang, Benjamin Feuer, Talukder Z. Jubery, Zi K. Deng, Andre Nakkab, Md. Zahid Hasan, Shivani Chiranjeevi, Kelly O. Marshall, Nirmal Baishnab, Asheesh Kumar Singh, Arti Singh, Soumik Sarkar, Nirav C. Merchant, Chinmay Hegde, Baskar Ganapathysubramanian
NeurIPS8