Duncan Watson-Parris

dblp:254/3021 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0002-5312-4950ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Probabilistic and Bayesian machine learning · 34% Knowledge representation and reasoning · 28% Question answering and dialogue systems · 12%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Environmental and earth informatics · 60% Computing education · 40%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
causal graph discovery
0.912025
Discovering Latent Causal Graphs from Spatiotemporal Data · ICML 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning
0.912025
Discovering Latent Causal Graphs from Spatiotemporal Data · ICML 2025
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.912025
Adapting While Learning: Grounding LLMs for Scientific Problems with Tool Usage Adaptation · ICML 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge management
knowledge internalization
0.912025
Adapting While Learning: Grounding LLMs for Scientific Problems with Tool Usage Adaptation · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
latent variable causal discovery
0.912025
Discovering Latent Causal Graphs from Spatiotemporal Data · ICML 2025
Natural language and speech › Language models and text generation › agentic language model
tool-augmented language models
0.912025
Adapting While Learning: Grounding LLMs for Scientific Problems with Tool Usage Adaptation · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.912025
Discovering Latent Causal Graphs from Spatiotemporal Data · ICML 2025
Computing education
automated assessment
0.912025
ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models · ICLR 2025
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.812024
Multi-Fidelity Residual Neural Processes for Scalable Surrogate Modeling · ICML 2024
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization › surrogate model
multi-fidelity surrogate modeling
0.812024
Multi-Fidelity Residual Neural Processes for Scalable Surrogate Modeling · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
neural processes
0.812024
Multi-Fidelity Residual Neural Processes for Scalable Surrogate Modeling · ICML 2024
Environmental and earth informatics
meteorology
0.512021
RainBench: Towards Data-Driven Global Precipitation Forecasting from Satellite Imagery · AAAI 2021
Environmental and earth informatics › weather forecasting
precipitation forecasting
0.512021
RainBench: Towards Data-Driven Global Precipitation Forecasting from Satellite Imagery · AAAI 2021
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
science question answering
0.312025
Adapting While Learning: Grounding LLMs for Scientific Problems with Tool Usage Adaptation · ICML 2025
Environmental and earth informatics
remote sensing
0.112021
RainBench: Towards Data-Driven Global Precipitation Forecasting from Satellite Imagery · AAAI 2021
Environmental and earth informatics › remote sensing
satellite imagery analysis
0.112021
RainBench: Towards Data-Driven Global Precipitation Forecasting from Satellite Imagery · AAAI 2021

Methods — techniques the papers use, named apart from their topics

expert annotation · 1.7LLM-based question generation · 1.7variational inference · 0.9tool-generated solutions · 0.9supervised fine-tuning · 0.9spatial kernel functions · 0.9identifiability analysis · 0.9accuracy-based problem categorization · 0.9multi-fidelity modeling · 0.8encoder-decoder architecture · 0.8deep learning · 0.5benchmark dataset construction · 0.5
YearPublicationVenuePosition
2025 ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models
abstract
The use of Large Language Models (LLMs) in climate science has recently gained significant attention. However, a critical issue remains: the lack of a comprehensive evaluation framework capable of assessing the quality and scientific validity of model outputs. To address this issue, we develop *ClimaGen* (Climate QA Generator), an adaptive learning framework that generates question-answer pairs from graduate textbooks with climate scientists in the loop. As a result, we present *ClimaQA-Gold*, an expert-annotated benchmark dataset alongside *ClimaQA-Silver*, a large-scale, comprehensive synthetic QA dataset for climate science. Finally, we develop evaluation strategies and compare different LLMs on our benchmarks. Our results offer novel insights into various approaches used to enhance knowledge of climate LLMs. ClimaQA’s source code is publicly available at https://github.com/Rose-STL-Lab/genie-climaqa
Veeramakali Vignesh Manivannan, Yasaman Jafari, Srikar Eranky, Spencer Ho, Rose Yu, Duncan Watson-Parris, Yi-An Ma, Leon Bergen, Taylor Berg-Kirkpatrick
ICLR6
2025 Adapting While Learning: Grounding LLMs for Scientific Problems with Tool Usage Adaptation
abstract
Large Language Models (LLMs) demonstrate promising capabilities in solving scientific problems but often suffer from the issue of hallucination. While integrating LLMs with tools can mitigate this issue, models fine-tuned on tool usage become overreliant on them and incur unnecessary costs. Inspired by how human experts assess problem complexity before selecting solutions, we propose a novel two-component fine-tuning method, Adapting while Learning (AWL). In the first component World Knowledge Learning (WKL), LLMs internalize scientific knowledge by learning from tool-generated solutions. In the second component Tool Usage Adaptation (TUA), we categorize problems as easy or hard based on the model’s accuracy, and train it to maintain direct reasoning for easy problems while switching to tools for hard ones. We validate our method on 6 scientific benchmark datasets across climate science, epidemiology, physics, and other domains. Compared to the original instruct model (8B), models post-trained with AWL achieve 29.11% higher answer accuracy and 12.72% better tool usage accuracy, even surpassing state-of-the-art models including GPT-4o and Claude-3.5 on 4 custom-created datasets. Our code is open-source at https://github.com/Rose-STL-Lab/Adapting-While-Learning.
Bohan Lyu 0001, Yadi Cao, Duncan Watson-Parris, Leon Bergen, Taylor Berg-Kirkpatrick, Rose Yu
ICML3
2025 Discovering Latent Causal Graphs from Spatiotemporal Data
abstract
Many important phenomena in scientific fields like climate, neuroscience, and epidemiology are naturally represented as spatiotemporal gridded data with complex interactions. Inferring causal relationships from these data is a challenging problem compounded by the high dimensionality of such data and the correlations between spatially proximate points. We present SPACY (SPAtiotemporal Causal discoverY), a novel framework based on variational inference, designed to model latent time series and their causal relationships from spatiotemporal data. SPACY alleviates the high-dimensional challenge by discovering causal structures in the latent space. To aggregate spatially proximate, correlated grid points, we use spatial factors, parametrized by spatial kernel functions, to map observational time series to latent representations. Theoretically, we generalize the problem to a continuous spatial domain and establish identifiability when the observations arise from a nonlinear, invertible function of the product of latent series and spatial factors. Using this approach, we avoid assumptions that are often unverifiable, including those about instantaneous effects or sufficient variability. Empirically, SPACY outperforms state-of-the-art baselines on synthetic data, even in challenging settings where existing methods struggle, while remaining scalable for large grids. SPACY also identifies key known phenomena from real-world climate data. An implementation of SPACY is available at \url{https://github.com/Rose-STL-Lab/SPACY/}
Sumanth Varambally, Duncan Watson-Parris, Yi-An Ma, Rose Yu
ICML3
2024 Multi-Fidelity Residual Neural Processes for Scalable Surrogate Modeling
abstract
Multi-fidelity surrogate modeling aims to learn an accurate surrogate at the highest fidelity level by combining data from multiple sources. Traditional methods relying on Gaussian processes can hardly scale to high-dimensional data. Deep learning approaches utilize neural network based encoders and decoders to improve scalability. These approaches share encoded representations across fidelities without including corresponding decoder parameters. This hinders inference performance, especially in out-of-distribution scenarios when the highest fidelity data has limited domain coverage. To address these limitations, we propose Multi-fidelity Residual Neural Processes (MFRNP), a novel multi-fidelity surrogate modeling framework. MFRNP explicitly models the residual between the aggregated output from lower fidelities and ground truth at the highest fidelity. The aggregation introduces decoders into the information sharing step and optimizes lower fidelity decoders to accurately capture both in-fidelity and cross-fidelity information. We show that MFRNP significantly outperforms state-of-the-art in learning partial differential equations and a real-world climate modeling task. Our code is published at: https://github.com/Rose-STL-Lab/MFRNP
Ruijia Niu, Dongxia Wu, Kai Kim, Yi-An Ma, Duncan Watson-Parris, Rose Yu
ICML5
2021 RainBench: Towards Data-Driven Global Precipitation Forecasting from Satellite Imagery
abstract
Extreme precipitation events, such as violent rainfall and hail storms, routinely ravage economies and livelihoods around the developing world. Climate change further aggravates this issue. Data-driven deep learning approaches could widen the access to accurate multi-day forecasts, to mitigate against such events. However, there is currently no benchmark dataset dedicated to the study of global precipitation forecasts. In this paper, we introduce RainBench, a new multi-modal benchmark dataset for data-driven precipitation forecasting. It includes simulated satellite data, a selection of relevant meteorological data from the ERA5 reanalysis product, and IMERG precipitation data. We also release PyRain, a library to process large precipitation datasets efficiently. We present an extensive analysis of our novel dataset and establish baseline results for two benchmark medium-range precipitation forecasting tasks. Finally, we discuss existing data-driven weather forecasting methodologies and suggest future research avenues.
Christian Schröder de Witt, Catherine Tong, Valentina Zantedeschi, Daniele De Martini, Freddie Kalaitzis, Matthew Chantry, Duncan Watson-Parris, Piotr Bilinski
AAAI7