Christopher J. Duffy

dblp:77/8486 · DBLP profile ↗
← Back
8ranked-venue papers in the field
0as first author
5since 2021 · last 2025
0000-0003-0080-6445ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 5Database Systems & Data Management · 2Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2025 Hierarchically Disentangled Recurrent Network for Factorizing System Dynamics of Multi-scale Systems: An application on Hydrological Systems
abstract
We present a framework for modeling multi-scale processes, and study its performance in the context of stream-flow forecasting in hydrology. Specifically, we propose a novel hierarchical recurrent neural architecture that factorizes the system dynamics at multiple temporal scales and captures their interactions. This framework consists of an inverse and a forward model. The inverse model is used to empirically resolve the system's temporal modes from data (physical model simulations, observed data, or a combination of them from the past), and these states are then used in the forward model to predict streamflow. Experiments on several catchments from the National Weather Service North Central River Forecast Center show that FHNN outperforms standard baselines, including physics-based models and transformer-based approaches. The model demonstrates particular effectiveness in catchments with low runoff ratios and colder climates. We further validate FHNN on the CAMELS (Catchment Attributes and MEteorology for Large-sample Studies), which is a widely used continental-scale hydrology benchmark dataset, confirming consistent performance improvements for 1–7 day streamflow forecasts across diverse hydrological conditions. Additionally, we show that FHNN can maintain accuracy even with limited training data through effective pre-training strategies and training global models.
Rahul Ghosh, Arvind Renganthan, Zac McEachran, Kelly Lindsay, Michael S. Steinbach, John Nieber, Christopher J. Duffy, Vipin Kumar 0001
ICDM7
2024 Towards Entity-Aware Conditional Variational Inference for Heterogeneous Time-Series Prediction: An application to Hydrology
abstract
Many environmental systems (e.g., hydrology basins) can be modeled as entity whose response (e.g., streamflow) depends on drivers (e.g., weather) conditioned on their characteristics (e.g., soil properties). We introduce Entity-aware Conditional Variational Inference (EA-CVI), a novel probabilistic inverse modeling approach, to deduce entity characteristics from observed driver-response data. EA-CVI infers probabilistic latent representations that can accurately predict response for diverse entities, particularly in out-of-sample few-shot settings. EA-CVI's latent embeddings encapsulate diverse entity characteristics within compact, low-dimensional representations. EA-CVI proficiently identifies dominant modes of variation in responses and offers the opportunity to infer a physical interpretation of the underlying attributes that shape these responses. EA-CVI can also generate new data samples by sampling from the learned distribution, making it useful in zero-shot scenarios. EA-CVI addresses the need for uncertainty estimation, particularly during extreme events, rendering it essential for data-driven decision-making in real-world applications. Extensive evaluations on a renowned hydrology benchmark dataset, CAMELS-GB, validate EA-CVI's abilities.
Rahul Ghosh, Arvind Renganathan, Wallace McAliley, Michael S. Steinbach, Christopher J. Duffy, Vipin Kumar 0001
SDM5
2023 Probabilistic Inverse Modeling: An Application in Hydrology
abstract
Rapid advancement in inverse modeling methods have brought into light their susceptibility to imperfect data. This has made it imperative to obtain more explainable and trustworthy estimates from these models. In hydrology, basin characteristics can be noisy or missing, impacting streamflow prediction. We propose a probabilistic inverse model framework that can reconstruct robust hydrology basin characteristics from dynamic input weather driver and streamflow response data. We address two aspects of building more explainable inverse models, uncertainty estimation (uncertainty due to imperfect data and imperfect model) and robustness. This can help improve the trust of water managers, handling of noisy data and reduce costs. We also propose an uncertainty based loss regularization that offers removal of 17% of temporal artifacts in reconstructions, 36% reduction in uncertainty and 4% higher coverage rate for basin characteristics. The forward model performance (streamflow estimation) is also improved by 6% using these uncertainty learning based reconstructions.
Somya Sharma, Rahul Ghosh, Arvind Renganathan, Snigdhansu Chatterjee, John Nieber, Christopher J. Duffy, Vipin Kumar 0001
SDM7
2023 Mini-Batch Learning Strategies for modeling long term temporal dependencies: A study in environmental applications
abstract
In many environmental applications, recurrent neural networks (RNNs) are often used to model physical variables with long temporal dependencies. However, due to minibatch training, temporal relationships between training segments within the batch (intra-batch) as well as between batches (inter-batch) are not considered, which can lead to limited performance. Stateful RNNs aim to address this issue by passing hidden states between batches. Since Stateful RNNs ignore intra-batch temporal dependency, there exists a trade-off between training stability and capturing temporal dependency. In this paper, we provide a quantitative comparison of different Stateful RNN modeling strategies, and propose two strategies to enforce both intra- and inter-batch temporal dependency. First, we extend Stateful RNNs by defining a batch as a temporally ordered set of training segments, which enables intra-batch sharing of temporal information. While this approach significantly improves the performance, it leads to much larger training times due to highly sequential training. To address this issue, we further propose a new strategy which augments a training segment with an initial value of the target variable from the timestep right before the starting of the training segment. In other words, we provide an initial value of the target variable as additional input so that the network can focus on learning changes relative to that initial value. By using this strategy, samples can be passed in any order (mini-batch training) which significantly reduces the training time while maintaining the performance. In demonstrating the utility of our approach in hydrological modeling, we observe that the most significant gains in predictive accuracy occur when these methods are applied to state variables whose values change more slowly, such as soil water and snowpack, rather than continuously moving flux variables such as streamflow.
Shaoming Xu, Ankush Khandelwal, Xiaowei Jia, Licheng Liu, Jared Willard, Rahul Ghosh, Kelly Cutler, Michael S. Steinbach, Christopher J. Duffy, John Nieber, Vipin Kumar 0001
SDM10
2022 Robust Inverse Framework using Knowledge-guided Self-Supervised Learning: An application to Hydrology
abstract
Machine Learning is beginning to provide state-of-the-art performance in a range of environmental applications such as streamflow prediction in a hydrologic basin. However, building accurate broad-scale models for streamflow remains challenging in practice due to the variability in the dominant hydrologic processes, which are best captured by sets of process-related basin characteristics. Existing basin characteristics suffer from noise and uncertainty, among many other things, which adversely impact model performance. To tackle the above challenges, in this paper, we propose a novel Knowledge-guided Self-Supervised Learning (KGSSL) inverse framework to extract system characteristics from driver(input) and response(output) data. This first-of-its-kind framework achieves robust performance even when characteristics are corrupted or missing. We evaluate the KGSSL framework in the context of stream flow modeling using CAMELS (Catchment Attributes and MEteorology for Large-sample Studies) which is a widely used hydrology benchmark dataset. Specifically, KGSSL outperforms baseline by 16% in predicting missing characteristics. Furthermore, in the context of forward modelling, KGSSL inferred characteristics provide a 35% improvement in performance over a standard baseline when the static characteristic are unknown.
Rahul Ghosh, Arvind Renganathan, Kshitij Tayal, Ankush Khandelwal, Xiaowei Jia, Christopher J. Duffy, John Nieber, Vipin Kumar 0001
KDD7
2015 Supporting Open Collaboration in Science Through Explicit and Linked Semantic Description of Processes
Yolanda Gil, Felix Michel, Varun Ratnakar, Jordan S. Read, Matheus Hauder, Christopher J. Duffy, Paul C. Hanson, Hilary Dugan
ESWC6
2010 An object-oriented shared data model for GIS and distributed hydrologic models
abstract
Distributed physical models for the space–time distribution of water, energy, vegetation, and mass flow require new strategies for data representation, model domain decomposition, a priori parameterization, and visualization. The geographic information system (GIS) has been traditionally used to accomplish these data management functionalities in hydrologic applications. However, the interaction between the data management tools and the physical model are often loosely integrated and nondynamic. This is because (a) the data types, semantics, resolutions, and formats for the physical model system and the distributed data or parameters may be different, with significant data preprocessing required before they can be shared; (b) the management tools may not be accessible or shared by the GIS and physical model; and (c) the individual systems may be operating system dependent or are driven by proprietary data structures. The impediment to seamless data flow between the two software components has the effect of increasing the model setup time and analysis time of model output results, and also makes it restrictive to perform sophisticated numerical modeling procedures (real-time forecasting, sensitivity analysis, etc.) that utilize extensive GIS data. These limitations can be offset to a large degree by developing an integrated software component that shares data between the (hydrologic) model and the GIS modules. We contend that the prerequisite for the development of such an integrated software component is a ‘shared data model’, which is designed using an object-oriented strategy. Here we present the design of such a shared data model taking into consideration the data type descriptions, identification of data classes, relationships, and constraints. The developed data model has been used as a method base for developing a coupled GIS interface to Penn State Integrated Hydrologic Model (PIHM), called PIHMgis.
Mukesh Kumar 0002, Gopal Bhatt, Christopher J. Duffy
Int. J. Geogr. Inf. Sci.3
2009 An efficient domain decomposition framework for accurate representation of geodata in distributed hydrologic models
Mukesh Kumar 0002, Gopal Bhatt, Christopher J. Duffy
Int. J. Geogr. Inf. Sci.3