VLDB 2026 Research / reviewers in the wild / expert
Paul C. Hanson
dblp:193/0068
· DBLP profile ↗
8ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0001-8533-6061ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Physics-Guided Foundation Model for Scientific Discovery: An Application to Aquatic ScienceabstractPhysics-guided machine learning (PGML) has become a prevalent approach in studying scientific systems due to its ability to integrate scientific theories for enhancing machine learning (ML) models. However, most PGML approaches are tailored to isolated and relatively simple tasks, which limits their applicability to complex systems involving multiple interacting processes and numerous influencing features. In this paper, we propose a Physics-Guided Foundation Model (PGFM) that combines pre-trained ML models and physics-based models and leverages their complementary strengths to improve the modeling of multiple coupled processes. To effectively conduct pre-training, we construct a simulated environmental system that encompasses a wide range of influencing features and various simulated variables generated by physics-based models. The model is pre-trained in this system to adaptively select important feature interactions guided by multi-task objectives. We then fine-tune the model for each specific task using true observations, while maintaining consistency with established physical theories, such as the principles of mass and energy conservation. We demonstrate the effectiveness of this methodology in modeling water temperature and dissolved oxygen dynamics in real-world lakes. The proposed PGFM is also broadly applicable to a range of scientific fields where physics-based models are being used. Runlong Yu, Chonghao Qiu, Robert Ladwig, Paul C. Hanson, Yiqun Xie, Xiaowei Jia |
AAAI | 4 |
| 2025 | Learning to Retrieve for Environmental Knowledge Discovery: An Augmentation-Adaptive Self-Supervised Learning FrameworkabstractThe discovery of environmental knowledge depends on labeled task-specific data, but is often constrained by the high cost of data collection. Existing machine learning approaches usually struggle to generalize in data-sparse or atypical conditions. To this end, we propose an Augmentation-Adaptive Self-Supervised Learning (A2SL) framework, which retrieves relevant observational samples to enhance modeling of the target ecosys-tem. Specifically, we introduce a multi-level pairwise learning loss to train a scenario encoder that captures varying degrees of similarity among scenarios. These learned similarities drive a retrieval mechanism that supplements a target scenario with relevant data from different locations or time periods. Further-more, to better handle variable scenarios, particularly under atypical or extreme conditions where traditional models struggle, we design an augmentation-adaptive mechanism that selectively enhances these scenarios through targeted data augmentation. Using freshwater ecosystems as a case study, we evaluate A2SL in modeling water temperature and dissolved oxygen dynamics in real-world lakes. Experimental results show that A2SL signif-icantly improves predictive accuracy and enhances robustness in data-scarce and atypical scenarios. Although this study focuses on freshwater ecosystems, the A2SL framework offers a broadly applicable solution in various scientific domains. Code-https://github.com/shiyuanlsy/A2sl Shiyuan Luo, Runlong Yu, Chonghao Qiu, Rahul Ghosh, Robert Ladwig, Paul C. Hanson, Yiqun Xie, Xiaowei Jia |
ICDM | 6 |
| 2024 | Adaptive Process-Guided Learning: An Application in Predicting Lake DO ConcentrationsabstractThis paper introduces a Process-Guided Learning (Pril) framework that integrates physical models with recurrent neural networks (RNNs) to enhance the prediction of dissolved oxygen (DO) concentrations in lakes, which is crucial for sus-taining water quality and ecosystem health. Unlike traditional RNNs, which may deliver high accuracy but often lack physical consistency and broad applicability, the Pril method incorporates differential DO equations for each lake layer, modeling it as a first-order linear solution using a forward Euler scheme with a daily timestep. However, this method is sensitive to numerical instabilities. When drastic fluctuations occur, the numerical integration is neither mass-conservative nor stable. Especially during stratified conditions, exogenous fluxes into each layer cause significant within-day changes in DO concentrations. To address this challenge, we further propose an Adaptive Process-Guided Learning (April) model, which dynamically adjusts timesteps from daily to sub-daily intervals with the aim of mitigating the discrepancies caused by variations in entrainment fluxes. April uses a generator-discriminator architecture to identify days with significant DO fluctuations and employs a multi-step Euler scheme with sub-daily timesteps to effectively manage these variations. We have tested our methods on a wide range of lakes in the Midwestern USA, and demonstrated robust capability in predicting DO concentrations even with limited training data. While primarily focused on aquatic ecosystems, this approach is broadly applicable to diverse scientific and engineering disciplines that utilize process-based models, such as power engineering, climate science, and biomedicine. Runlong Yu, Chonghao Qiu, Robert Ladwig, Paul C. Hanson, Yiqun Xie, Xiaowei Jia |
ICDM | 4 |
| 2024 | Evolution-Based Feature Selection for Predicting Dissolved Oxygen Concentrations in Lakes
Runlong Yu, Robert Ladwig, Peijun Zhu, Paul C. Hanson, Yiqun Xie, Xiaowei Jia |
PPSN (4) | 5 |
| 2023 | Spatiotemporal Classification with limited labels using Constrained Clustering for large datasetsabstractCreating separable representations via representation learning and clustering is critical in analyzing large unstructured datasets with only a few labels. Separable representations can lead to supervised models with better classification capabilities and additionally aid in generating new labeled samples. Most unsupervised and semisupervised methods to analyze large datasets do not leverage the existing small amounts of labels to get better representations. In this paper, we propose a spatiotemporal clustering paradigm that uses spatial and temporal features combined with a constrained loss to produce separable representations. We show the working of this method on the newly published dataset ReaLSAT, a dataset of surface water dynamics for over 680,000 lakes across the world, making it an essential dataset in terms of ecology and sustainability. Using this large un- labelled dataset, we first show how a spatiotemporal representation is better compared to just spatial or temporal representation. We then show how we can learn even better representations using a constrained loss with few labels. We conclude by showing how our method, using few labels, can pick out new labeled samples from the unlabeled data, which can be used to augment supervised methods leading to better classification. Praveen Ravirathinam, Rahul Ghosh, Keyang Xuan, Ankush Khandelwal, Hilary Dugan, Paul C. Hanson, Vipin Kumar 0001 |
SDM | 7 |
| 2017 | GRAPLEr: A distributed collaborative environment for lake ecosystem modeling that integrates overlay networks, high-throughput computing, and WEB servicesabstractSummary The GLEON Research And PRAGMA Lake Expedition—GRAPLE—is a collaborative effort between computer science and lake ecology researchers. It aims to improve our understanding and predictive capacity of the threats to the water quality of our freshwater resources, including climate change. This paper presents GRAPLEr, a distributed computing system used to address the modeling needs of GRAPLE researchers. GRAPLEr integrates and applies overlay virtual network, high‐throughput computing, and WEB service technologies in a novel way. First, its user‐level IP‐over‐P2P overlay network allows compute and storage resources distributed across independently administered institutions (including private and public clouds) to be aggregated into a common virtual network, despite the presence of firewalls and network address translators. Second, resources aggregated by the IP‐over‐P2P virtual network run unmodified high‐throughput‐computing middleware to enable large numbers of model simulations to be executed concurrently across the distributed computing resources. Third, a WEB service interface allows end users to submit job requests to the system using client libraries that integrate with the R statistical computing environment. The paper presents the GRAPLEr architecture, describes its implementation and reports on its performance for batches of general lake model simulations across 3 cloud infrastructures (University of Florida, CloudLab, and Microsoft Azure). Kensworth Subratie, Saumitra Aditya, Srinivas Mahesula, Renato J. O. Figueiredo, Cayelan C. Carey, Paul C. Hanson |
Concurr. Comput. Pract. Exp. | 6 |
| 2015 | A Task-Centered Framework for Computationally-Grounded Science CollaborationsabstractCollaboration is ubiquitous in today's science, yet there is limited support for coordinating scientific work. The general-purpose tools that are typically used (e.g., email, shared document editing, social coding sites), have still not replaced in-person meetings, phone calls, and extensive emails needed to coordinate and track collaborative activities. Scientists with diverse knowledge and skills around the globe could collaborate by opening scientific processes that expose all tasks and activities publicly to achieve a shared scientific question. This paper describes the Organic Data Science framework to support scientific collaborations that revolve around complex science questions that require significant coordination, entice contributors to remain engaged for extended periods of time, and enable continuous growth to accommodate new contributors as the work evolves over time. We discuss how the design of this framework incorporates principles followed by successful on-line communities. We present initial results to date of several communities that are collaborating using this framework. Yolanda Gil, Felix Michel, Varun Ratnakar, Matheus Hauder, Christopher J. Duffy, Hilary Dugan, Paul C. Hanson |
e-Science | 7 |
| 2015 | Supporting Open Collaboration in Science Through Explicit and Linked Semantic Description of Processes
Yolanda Gil, Felix Michel, Varun Ratnakar, Jordan S. Read, Matheus Hauder, Christopher J. Duffy, Paul C. Hanson, Hilary Dugan |
ESWC | 7 |