Keiya Hirashima

dblp:340/7645 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0002-1972-2674ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Computational science and engineering · 100%
Artificial intelligence
2 papers
Deep learning architectures and training · 30% Vision and language · 30% Representation and self-supervised learning · 30%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 9 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
foundation model
0.912025
AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked modeling
0.912025
AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025
High-performance computing
scientific computing systems
0.912025
The First Star-by-star $N$-body/Hydrodynamics Simulation of Our Galaxy Coupling with a Surrogate Model · SC 2025
Computational science and engineering › computational physics
physics simulation
0.812024
The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine Learning · NeurIPS 2024
Computational science and engineering › scientific machine learning
surrogate modeling
0.812024
The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine Learning · NeurIPS 2024
Information retrieval › evaluation
benchmark dataset
0.812024
The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine Learning · NeurIPS 2024
Computational science and engineering › astronomy
astronomical data analysis
0.312025
AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025
Computational science and engineering
astronomy
0.312025
AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

surrogate model · 2.6machine learning · 2.6transformer · 1.7tokenization · 1.7masked modeling · 1.7pytorch · 1.5
YearPublicationVenuePosition
2025 AION-1: Omnimodal Foundation Model for Astronomical Sciences
abstract
While foundation models have shown promise across a variety of fields, astronomy lacks a unified framework for joint modeling across its highly diverse data modalities. In this paper, we present AION-1, the first large-scale multimodal foundation family of models for astronomy. AION-1 enables arbitrary transformations between heterogeneous data types using a two-stage architecture: modality-specific tokenization followed by transformer-based masked modeling of cross-modal token sequences. Trained on over 200M astronomical objects, AION-1 demonstrates strong performance across regression, classification, generation, and object retrieval tasks. Beyond astronomy, AION-1 provides a scalable blueprint for multimodal scientific foundation models that can seamlessly integrate heterogeneous combinations of real-world observations. Our model release is entirely open source, including the dataset, training script, and weights.
Liam Holden Parker, François Lanusse, Jeff Shen, Ollie Liu, Tom Hehir, Leopoldo Sarra, Lucas Meyer, Micah Bowles, Sebastian Wagner-Carena, Helen Qu, Siavash Golkar, Alberto Bietti, Hatim Bourfoune, Pierre Cornette, Keiya Hirashima, Géraud Krawezik, Ruben Ohana, Nicholas Lourie, Michael McCabe, Rudy Morel, Payel Mukhopadhyay, Mariel Pettee, Kyunghyun Cho, Miles D. Cranmer, Shirley Ho
NeurIPS15
2025 The First Star-by-star $N$-body/Hydrodynamics Simulation of Our Galaxy Coupling with a Surrogate Model
abstract
A major goal of computational astrophysics is to simulate the Milky Way Galaxy with sufficient resolution down to individual stars. However, the scaling fails due to some small-scale, short-timescale phenomena, such as supernova explosions. We have developed a novel integration scheme of N-body/hydrodynamics simulations working with machine learning. This approach bypasses the short timesteps caused by supernova explosions using a surrogate model, thereby improving scalability. With this method, we reached 300 billion particles using 148,900 nodes, equivalent to 7,147,200 CPU cores, breaking through the billion-particle barrier currently faced by state-of-the-art simulations. This resolution allows us to perform the first star-by-star galaxy simulation, which resolves individual stars in the Milky Way Galaxy. The performance scales over 104 CPU cores, an upper limit in the current state-of-the-art simulations using both A64FX and X86-64 processors and NVIDIA CUDA GPUs.
Keiya Hirashima, Michiko S. Fujii, Takayuki R. Saitoh, Naoto Harada, Kentaro Nomura, Kohji Yoshikawa, Yutaka Hirai, Tetsuro Asano, Kana Moriwaki, Masaki Iwasawa, Takashi Okamoto, Junichiro Makino
SC1
2024 The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine Learning
abstract
Machine learning based surrogate models offer researchers powerful tools for accelerating simulation-based workflows. However, as standard datasets in this space often cover small classes of physical behavior, it can be difficult to evaluate the efficacy of new approaches. To address this gap, we introduce the Well: a large-scale collection of datasets containing numerical simulations of a wide variety of spatiotemporal physical systems. The Well draws from domain experts and numerical software developers to provide 15TB of data across 16 datasets covering diverse domains such as biological systems, fluid dynamics, acoustic scattering, as well as magneto-hydrodynamic simulations of extra-galactic fluids or supernova explosions. These datasets can be used individually or as part of a broader benchmark suite. To facilitate usage of the Well, we provide a unified PyTorch interface for training and evaluating models. We demonstrate the function of this library by introducing example baselines that highlight the new challenges posed by the complex dynamics of the Well. The code and data is available at https://github.com/PolymathicAI/the_well.
Ruben Ohana, Michael McCabe, Lucas Meyer, Rudy Morel, Fruzsina Julia Agocs, Miguel Beneitez, Marsha J. Berger, Blakesley Burkhart, Stuart B. Dalziel, Drummond B. Fielding, Daniel Fortunato, Jared A. Goldberg, Keiya Hirashima, Yan-Fei Jiang, Rich R. Kerswell, Suryanarayana Maddu, Jonah Miller, Payel Mukhopadhyay, Stefan S. Nixon, Jeff Shen, Romain Watteaux, Bruno Régaldo-Saint Blancard, François Rozet, Liam Holden Parker, Miles D. Cranmer, Shirley Ho
NeurIPS13