Micah Bowles

dblp:280/0063 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Representation and self-supervised learning · 39% Deep learning architectures and training · 31% Vision and language · 31%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational science and engineering · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
foundation model
0.912025
AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked modeling
0.912025
AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025
Computational science and engineering › astronomy
astrophysics
0.812024
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100 TB of Astronomical Scientific Data · NeurIPS 2024
Computational science and engineering › astronomy
astronomical data analysis
0.312025
AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025
Computational science and engineering
astronomy
0.312025
AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025
Machine learning › Representation and self-supervised learning
multimodal representation learning
0.212024
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100 TB of Astronomical Scientific Data · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

transformer · 1.7tokenization · 1.7masked modeling · 1.7multimodal machine learning · 1.5
YearPublicationVenuePosition
2025 AION-1: Omnimodal Foundation Model for Astronomical Sciences
abstract
While foundation models have shown promise across a variety of fields, astronomy lacks a unified framework for joint modeling across its highly diverse data modalities. In this paper, we present AION-1, the first large-scale multimodal foundation family of models for astronomy. AION-1 enables arbitrary transformations between heterogeneous data types using a two-stage architecture: modality-specific tokenization followed by transformer-based masked modeling of cross-modal token sequences. Trained on over 200M astronomical objects, AION-1 demonstrates strong performance across regression, classification, generation, and object retrieval tasks. Beyond astronomy, AION-1 provides a scalable blueprint for multimodal scientific foundation models that can seamlessly integrate heterogeneous combinations of real-world observations. Our model release is entirely open source, including the dataset, training script, and weights.
Liam Holden Parker, François Lanusse, Jeff Shen, Ollie Liu, Tom Hehir, Leopoldo Sarra, Lucas Meyer, Micah Bowles, Sebastian Wagner-Carena, Helen Qu, Siavash Golkar, Alberto Bietti, Hatim Bourfoune, Pierre Cornette, Keiya Hirashima, Géraud Krawezik, Ruben Ohana, Nicholas Lourie, Michael McCabe, Rudy Morel, Payel Mukhopadhyay, Mariel Pettee, Kyunghyun Cho, Miles D. Cranmer, Shirley Ho
NeurIPS8
2024 The Multimodal Universe: Enabling Large-Scale Machine Learning with 100 TB of Astronomical Scientific Data
abstract
We present the Multimodal Universe, a large-scale multimodal dataset of scientific astronomical data, compiled specifically to facilitate machine learning research. Overall, our dataset contains hundreds of millions of astronomical observations, constituting 100TB of multi-channel and hyper-spectral images, spectra, multivariate time series, as well as a wide variety of associated scientific measurements and metadata. In addition, we include a range of benchmark tasks representative of standard practices for machine learning methods in astrophysics. This massive dataset will enable the development of large multi-modal models specifically targeted towards scientific applications. All codes used to compile the dataset, and a description of how to access the data is available at https://github.com/MultimodalUniverse/MultimodalUniverse
Eirini Angeloudi, Jeroen Audenaert, Micah Bowles, Benjamin M. Boyd, David Chemaly, Brian Cherinka, Ioana Ciuca, Miles D. Cranmer, Aaron Do, Matthew Grayling, Erin E. Hayes, Tom Hehir, Shirley Ho, Marc Huertas-Company, Kartheik Iyer, Maja Jablonska, François Lanusse, Kaisey Mandel, Rafael Martínez-Galarza, Peter Melchior, Lucas Meyer, Liam Holden Parker, Helen Qu, Jeff Shen, Michael J. Smith 0013, Connor Stone, Mike Walmsley, John F. Wu
NeurIPS3