Sungduk Yu

dblp:339/7180 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Deep learning architectures and training · 38% Probabilistic and Bayesian machine learning · 38% Information extraction and text analysis · 19%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Environmental and earth informatics · 100%
Databases, data mining, and information retrieval
2 papers
Data mining · 74% Machine learning and data management · 26%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Environmental and earth informatics
climate modeling
1.522025
ClimSim-Online: A Large Multi-Scale Dataset and Framework for Hybrid Physics-ML Climate Emulation · J. Mach. Learn. Res. 2025
ClimSim: A large multi-scale dataset for hybrid physics-ML climate emulation · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery
0.912025
A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal model
causal world model
0.912025
A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment · ICML 2025
Machine learning › Deep learning architectures and training
scientific machine learning
0.912025
ClimSim-Online: A Large Multi-Scale Dataset and Framework for Hybrid Physics-ML Climate Emulation · J. Mach. Learn. Res. 2025
Natural language and speech › Information extraction and text analysis
temporal information extraction
0.912025
A Semantic Parsing Framework for End-to-End Time Normalization · NeurIPS 2025
Machine learning › Deep learning architectures and training › attention mechanism
transformer attention
0.912025
A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment · ICML 2025
Compilers and program optimization
code generation
0.912025
A Semantic Parsing Framework for End-to-End Time Normalization · NeurIPS 2025
Environmental and earth informatics
climate prediction
0.812024
ChaosBench: A Multi-Channel, Physics-Based Benchmark for Subseasonal-to-Seasonal Climate Prediction · NeurIPS 2024
Environmental and earth informatics › climate prediction
subseasonal-to-seasonal forecasting
0.812024
ChaosBench: A Multi-Channel, Physics-Based Benchmark for Subseasonal-to-Seasonal Climate Prediction · NeurIPS 2024
Environmental and earth informatics
climate science
0.712023
ClimSim: A large multi-scale dataset for hybrid physics-ML climate emulation · NeurIPS 2023
Data mining
dataset construction
0.712023
ClimSim: A large multi-scale dataset for hybrid physics-ML climate emulation · NeurIPS 2023
Natural language and speech › Language models and text generation
large language model
0.312025
A Semantic Parsing Framework for End-to-End Time Normalization · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

machine learning emulators · 1.7large language model · 1.7data augmentation · 1.7containerized pipeline · 1.7vision transformer · 1.5physics-informed modeling · 1.5graph neural network · 1.5ensemble forecasting · 1.5stochastic regression · 1.3regression baseline · 1.3attention mechanism analysis · 0.9
YearPublicationVenuePosition
2025 A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment
abstract
Are generative pre-trained transformer (GPT) models, trained only to predict the next token, implicitly learning a world model from which sequences are generated one token at a time? We address this question by deriving a causal interpretation of the attention mechanism in GPT and presenting a causal world model that arises from this interpretation. Furthermore, we propose that GPT models, at inference time, can be utilized for zero-shot causal structure learning for input sequences, and introduce a corresponding confidence score. Empirical tests were conducted in controlled environments using the setups of the Othello and Chess strategy games. A GPT, pre-trained on real-world games played with the intention of winning, was tested on out-of-distribution synthetic data consisting of sequences of random legal moves. We find that the GPT model is likely to generate legal next moves for out-of-distribution sequences for which a causal structure is encoded in the attention mechanism with high confidence. In cases where it generates illegal moves, it also fails to capture a causal structure.
Raanan Y. Rohekar, Yaniv Gurwicz, Sungduk Yu, Estelle Aflalo, Vasudev Lal
ICML3
2025 A Semantic Parsing Framework for End-to-End Time Normalization
abstract
Time normalization is the task of converting natural language temporal expressions into machine-readable representations. It underpins many downstream applications in information retrieval, question answering, and clinical decision-making. Traditional systems based on the ISO-TimeML schema limit expressivity and struggle with complex constructs such as compositional, event-relative, and multi-span time expressions. In this work, we introduce a novel formulation of time normalization as a code generation task grounded in the SCATE framework, which defines temporal semantics through symbolic and compositional operators. We implement a fully executable SCATE Python library and demonstrate that large language models (LLMs) can generate executable SCATE code. Leveraging this capability, we develop an automatic data augmentation pipeline using LLMs to synthesize large-scale annotated data with code-level validation. Our experiments show that small, locally deployable models trained on this augmented data can achieve strong performance, outperforming even their LLM parents and enabling practical, accurate, and interpretable time normalization.
Xin Su 0008, Sungduk Yu, Phillip Howard, Steven Bethard
NeurIPS2
2025 ClimSim-Online: A Large Multi-Scale Dataset and Framework for Hybrid Physics-ML Climate Emulation
abstract
Modern climate projections lack adequate spatial and temporal resolution due to computational constraints, leading to inaccuracies in representing critical processes like thunderstorms that occur on the sub-resolution scale. Hybrid methods combining physics with machine learning (ML) offer faster, higher fidelity climate simulations by outsourcing compute-hungry, high-resolution simulations to ML emulators. However, these hybrid physics-ML simulations require domain-specific data and workflows that have been inaccessible to many ML experts. This paper is an extended version of our NeurIPS award-winning ClimSim dataset paper. The ClimSim dataset includes 5.7 billion pairs of multivariate input/output vectors spanning ten years at high temporal resolution, capturing the influence of high-resolution, high-fidelity physics on a host climate simulator's macro-scale state. In this extended version, we introduce a significant new contribution in Section 5, which provides a cross-platform, containerized pipeline to integrate ML models into operational climate simulators for hybrid testing. We also implement various baselines of ML models and hybrid simulators to highlight the ML challenges of building stable, skillful emulators. The data (https://huggingface.co/datasets/LEAP/ClimSim_high-res, also in a low-resolution version at https://huggingface.co/datasets/LEAP/ClimSim_low-res and an aquaplanet version at https://huggingface.co/datasets/LEAP/ClimSim_low-res_aqua-planet) and code (https://leap-stc.github.io/ClimSim and https://github.com/leap-stc/climsim-online) are publicly released to support the development of hybrid physics-ML and high-fidelity climate simulations.
Sungduk Yu, Zeyuan Hu 0005, Akshay Subramaniam, Walter M. Hannah, Liran Peng, Zhiyuan Jerry Lin, Mohamed Aziz Bhouri, Ritwik Gupta, Björn Lütjens, Justus C. Will, Gunnar Behrens, Julius Busecke, Nora Loose, Charles Stern, Tom Beucler, Bryce E. Harrop, Helge Heuer, Benjamin R. Hillman, Andrea M. Jenney, Nana Liu, Alistair White, Zhiming Kuang, Fiaz Ahmed, Elizabeth A. Barnes, Noah D. Brenowitz, Christopher S. Bretherton, Veronika Eyring, Savannah L. Ferretti, Nicholas J. Lutsko, Pierre Gentine, Stephan Mandt, J. David Neelin, Rose Yu, Laure Zanna, Nathan M. Urban, Janni Yuval, Ryan Abernathey, Pierre Baldi, Wayne Chuang, Fernando Iglesias-Suarez, Sanket R. Jantre, Po-Lun Ma, Sara Shamekh, Michael S. Pritchard
J. Mach. Learn. Res.1
2024 ChaosBench: A Multi-Channel, Physics-Based Benchmark for Subseasonal-to-Seasonal Climate Prediction
abstract
Accurate prediction of climate in the subseasonal-to-seasonal scale is crucial for disaster preparedness and robust decision making amidst climate change. Yet, forecasting beyond the weather timescale is challenging because it deals with problems other than initial condition, including boundary interaction, butterfly effect, and our inherent lack of physical understanding. At present, existing benchmarks tend to have shorter forecasting range of up-to 15 days, do not include a wide range of operational baselines, and lack physics-based constraints for explainability. Thus, we propose ChaosBench, a challenging benchmark to extend the predictability range of data-driven weather emulators to S2S timescale. First, ChaosBench is comprised of variables beyond the typical surface-atmospheric ERA5 to also include ocean, ice, and land reanalysis products that span over 45 years to allow for full Earth system emulation that respects boundary conditions. We also propose physics-based, in addition to deterministic and probabilistic metrics, to ensure a physically-consistent ensemble that accounts for butterfly effect. Furthermore, we evaluate on a diverse set of physics-based forecasts from four national weather agencies as baselines to our data-driven counterpart such as ViT/ClimaX, PanguWeather, GraphCast, and FourCastNetV2. Overall, we find methods originally developed for weather-scale applications fail on S2S task: their performance simply collapse to an unskilled climatology. Nonetheless, we outline and demonstrate several strategies that can extend the predictability range of existing weather emulators, including the use of ensembles, robust control of error propagation, and the use of physics-informed models. Our benchmark, datasets, and instructions are available at https://leap-stc.github.io/ChaosBench.
Juan Nathaniel, Yongquan Qu, Sungduk Yu, Julius Busecke, Aditya Grover, Pierre Gentine
NeurIPS4
2023 ClimSim: A large multi-scale dataset for hybrid physics-ML climate emulation
abstract
Modern climate projections lack adequate spatial and temporal resolution due to computational constraints. A consequence is inaccurate and imprecise predictions of critical processes such as storms. Hybrid methods that combine physics with machine learning (ML) have introduced a new generation of higher fidelity climate simulators that can sidestep Moore's Law by outsourcing compute-hungry, short, high-resolution simulations to ML emulators. However, this hybrid ML-physics simulation approach requires domain-specific treatment and has been inaccessible to ML experts because of lack of training data and relevant, easy-to-use workflows. We present ClimSim, the largest-ever dataset designed for hybrid ML-physics research. It comprises multi-scale climate simulations, developed by a consortium of climate scientists and ML researchers. It consists of 5.7 billion pairs of multivariate input and output vectors that isolate the influence of locally-nested, high-resolution, high-fidelity physics on a host climate simulator's macro-scale physical state.The dataset is global in coverage, spans multiple years at high sampling frequency, and is designed such that resulting emulators are compatible with downstream coupling into operational climate simulators. We implement a range of deterministic and stochastic regression baselines to highlight the ML challenges and their scoring. The data (https://huggingface.co/datasets/LEAP/ClimSim_high-res) and code (https://leap-stc.github.io/ClimSim) are released openly to support the development of hybrid ML-physics and high-fidelity climate simulations for the benefit of science and society.
Sungduk Yu, Walter M. Hannah, Liran Peng, Zhiyuan Jerry Lin, Mohamed Aziz Bhouri, Ritwik Gupta, Björn Lütjens, Justus C. Will, Gunnar Behrens, Julius Busecke, Nora Loose, Charles Stern, Tom Beucler, Bryce E. Harrop, Benjamin R. Hillman, Andrea M. Jenney, Savannah L. Ferretti, Nana Liu, Anima Anandkumar, Noah D. Brenowitz, Veronika Eyring, Nicholas Geneva, Pierre Gentine, Stephan Mandt, Jaideep Pathak, Akshay Subramaniam, Carl Vondrick, Rose Yu, Laure Zanna, Ryan Abernathey, Fiaz Ahmed, David C. Bader, Pierre Baldi, Elizabeth A. Barnes, Christopher S. Bretherton, Peter M. Caldwell, Wayne Chuang, Yilun Han, Fernando Iglesias-Suarez, Sanket R. Jantre, Karthik Kashinath, Marat Khairoutdinov, Thorsten Kurth, Nicholas J. Lutsko, Po-Lun Ma, Griffin Mooers, J. David Neelin, David A. Randall, Sara Shamekh, Nathan M. Urban, Janni Yuval, Mike Pritchard
NeurIPS1