Christopher J. Duffy

dblp:77/8486 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0003-0080-6445ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 Hierarchically Disentangled Recurrent Network for Factorizing System Dynamics of Multi-scale Systems: An application on Hydrological Systems
abstract
We present a framework for modeling multi-scale processes, and study its performance in the context of stream-flow forecasting in hydrology. Specifically, we propose a novel hierarchical recurrent neural architecture that factorizes the system dynamics at multiple temporal scales and captures their interactions. This framework consists of an inverse and a forward model. The inverse model is used to empirically resolve the system's temporal modes from data (physical model simulations, observed data, or a combination of them from the past), and these states are then used in the forward model to predict streamflow. Experiments on several catchments from the National Weather Service North Central River Forecast Center show that FHNN outperforms standard baselines, including physics-based models and transformer-based approaches. The model demonstrates particular effectiveness in catchments with low runoff ratios and colder climates. We further validate FHNN on the CAMELS (Catchment Attributes and MEteorology for Large-sample Studies), which is a widely used continental-scale hydrology benchmark dataset, confirming consistent performance improvements for 1–7 day streamflow forecasts across diverse hydrological conditions. Additionally, we show that FHNN can maintain accuracy even with limited training data through effective pre-training strategies and training global models.
Rahul Ghosh, Arvind Renganthan, Zac McEachran, Kelly Lindsay, Michael S. Steinbach, John Nieber, Christopher J. Duffy, Vipin Kumar 0001
ICDM7
2024 Towards Entity-Aware Conditional Variational Inference for Heterogeneous Time-Series Prediction: An application to Hydrology
abstract
Many environmental systems (e.g., hydrology basins) can be modeled as entity whose response (e.g., streamflow) depends on drivers (e.g., weather) conditioned on their characteristics (e.g., soil properties). We introduce Entity-aware Conditional Variational Inference (EA-CVI), a novel probabilistic inverse modeling approach, to deduce entity characteristics from observed driver-response data. EA-CVI infers probabilistic latent representations that can accurately predict response for diverse entities, particularly in out-of-sample few-shot settings. EA-CVI's latent embeddings encapsulate diverse entity characteristics within compact, low-dimensional representations. EA-CVI proficiently identifies dominant modes of variation in responses and offers the opportunity to infer a physical interpretation of the underlying attributes that shape these responses. EA-CVI can also generate new data samples by sampling from the learned distribution, making it useful in zero-shot scenarios. EA-CVI addresses the need for uncertainty estimation, particularly during extreme events, rendering it essential for data-driven decision-making in real-world applications. Extensive evaluations on a renowned hydrology benchmark dataset, CAMELS-GB, validate EA-CVI's abilities.
Rahul Ghosh, Arvind Renganathan, Wallace McAliley, Michael S. Steinbach, Christopher J. Duffy, Vipin Kumar 0001
SDM5
2023 Probabilistic Inverse Modeling: An Application in Hydrology
abstract
Rapid advancement in inverse modeling methods have brought into light their susceptibility to imperfect data. This has made it imperative to obtain more explainable and trustworthy estimates from these models. In hydrology, basin characteristics can be noisy or missing, impacting streamflow prediction. We propose a probabilistic inverse model framework that can reconstruct robust hydrology basin characteristics from dynamic input weather driver and streamflow response data. We address two aspects of building more explainable inverse models, uncertainty estimation (uncertainty due to imperfect data and imperfect model) and robustness. This can help improve the trust of water managers, handling of noisy data and reduce costs. We also propose an uncertainty based loss regularization that offers removal of 17% of temporal artifacts in reconstructions, 36% reduction in uncertainty and 4% higher coverage rate for basin characteristics. The forward model performance (streamflow estimation) is also improved by 6% using these uncertainty learning based reconstructions.
Somya Sharma, Rahul Ghosh, Arvind Renganathan, Snigdhansu Chatterjee, John Nieber, Christopher J. Duffy, Vipin Kumar 0001
SDM7
2023 Mini-Batch Learning Strategies for modeling long term temporal dependencies: A study in environmental applications
abstract
In many environmental applications, recurrent neural networks (RNNs) are often used to model physical variables with long temporal dependencies. However, due to minibatch training, temporal relationships between training segments within the batch (intra-batch) as well as between batches (inter-batch) are not considered, which can lead to limited performance. Stateful RNNs aim to address this issue by passing hidden states between batches. Since Stateful RNNs ignore intra-batch temporal dependency, there exists a trade-off between training stability and capturing temporal dependency. In this paper, we provide a quantitative comparison of different Stateful RNN modeling strategies, and propose two strategies to enforce both intra- and inter-batch temporal dependency. First, we extend Stateful RNNs by defining a batch as a temporally ordered set of training segments, which enables intra-batch sharing of temporal information. While this approach significantly improves the performance, it leads to much larger training times due to highly sequential training. To address this issue, we further propose a new strategy which augments a training segment with an initial value of the target variable from the timestep right before the starting of the training segment. In other words, we provide an initial value of the target variable as additional input so that the network can focus on learning changes relative to that initial value. By using this strategy, samples can be passed in any order (mini-batch training) which significantly reduces the training time while maintaining the performance. In demonstrating the utility of our approach in hydrological modeling, we observe that the most significant gains in predictive accuracy occur when these methods are applied to state variables whose values change more slowly, such as soil water and snowpack, rather than continuously moving flux variables such as streamflow.
Shaoming Xu, Ankush Khandelwal, Xiaowei Jia, Licheng Liu, Jared Willard, Rahul Ghosh, Kelly Cutler, Michael S. Steinbach, Christopher J. Duffy, John Nieber, Vipin Kumar 0001
SDM10
2022 Robust Inverse Framework using Knowledge-guided Self-Supervised Learning: An application to Hydrology
abstract
Machine Learning is beginning to provide state-of-the-art performance in a range of environmental applications such as streamflow prediction in a hydrologic basin. However, building accurate broad-scale models for streamflow remains challenging in practice due to the variability in the dominant hydrologic processes, which are best captured by sets of process-related basin characteristics. Existing basin characteristics suffer from noise and uncertainty, among many other things, which adversely impact model performance. To tackle the above challenges, in this paper, we propose a novel Knowledge-guided Self-Supervised Learning (KGSSL) inverse framework to extract system characteristics from driver(input) and response(output) data. This first-of-its-kind framework achieves robust performance even when characteristics are corrupted or missing. We evaluate the KGSSL framework in the context of stream flow modeling using CAMELS (Catchment Attributes and MEteorology for Large-sample Studies) which is a widely used hydrology benchmark dataset. Specifically, KGSSL outperforms baseline by 16% in predicting missing characteristics. Furthermore, in the context of forward modelling, KGSSL inferred characteristics provide a 35% improvement in performance over a standard baseline when the static characteristic are unknown.
Rahul Ghosh, Arvind Renganathan, Kshitij Tayal, Ankush Khandelwal, Xiaowei Jia, Christopher J. Duffy, John Nieber, Vipin Kumar 0001
KDD7
2021 Artificial Intelligence for Modeling Complex Systems: Taming the Complexity of Expert Models to Improve Decision Making
abstract
Major societal and environmental challenges involve complex systems that have diverse multi-scale interacting processes. Consider, for example, how droughts and water reserves affect crop production and how agriculture and industrial needs affect water quality and availability. Preventive measures, such as delaying planting dates and adopting new agricultural practices in response to changing weather patterns, can reduce the damage caused by natural processes. Understanding how these natural and human processes affect one another allows forecasting the effects of undesirable situations and study interventions to take preventive measures. For many of these processes, there are expert models that incorporate state-of-the-art theories and knowledge to quantify a system's response to a diversity of conditions. A major challenge for efficient modeling is the diversity of modeling approaches across disciplines and the wide variety of data sources available only in formats that require complex conversions. Using expert models for particular problems requires integration of models with third-party data as well as integration of models across disciplines. Modelers face significant heterogeneity that requires resolving semantic, spatiotemporal, and execution mismatches, which are largely done by hand today and may take more than 2 years of effort. We are developing a modeling framework that uses artificial intelligence (AI) techniques to reduce modeling effort while ensuring utility for decision making. Our work to date makes several innovative contributions: (1) an intelligent user interface that guides analysts to frame their modeling problem and assists them by suggesting relevant choices and automating steps along the way; (2) semantic metadata for models, including their modeling variables and constraints, that ensures model relevance and proper use for a given decision-making problem; and (3) semantic representations of datasets in terms of modeling variables that enable automated data selection and data transformations. This framework is implemented in the MINT (Model INTegration) framework, and currently includes data and models to analyze the interactions between natural and human systems involving climate, water availability, agricultural production, and markets. Our work to date demonstrates the utility of AI techniques to accelerate modeling to support decision-making and uncovers several challenging directions for future work.
Yolanda Gil, Daniel Garijo, Deborah Khider, Craig A. Knoblock, Varun Ratnakar, Maximiliano Osorio, Hernán Vargas, Minh Pham 0004, Jay Pujara, Basel Shbita, Yao-Yi Chiang, Dan Feldman, Yijun Lin 0001, Hayley Song, Vipin Kumar 0001, Ankush Khandelwal, Michael S. Steinbach, Kshitij Tayal, Shaoming Xu, Suzanne A. Pierce, Lissa Pearson, Daniel Hardesty-Lewis, Ewa Deelman, Rafael Ferreira da Silva, Rajiv Mayani, Armen R. Kemanian, Lorne Leonard, Scott D. Peckham, Maria Stoica 0001, Kelly M. Cobourn, Zeya Zhang, Christopher J. Duffy, Lele Shu
ACM Trans. Interact. Intell. Syst.34
2016 Tuning Heterogeneous Computing Platforms for Large-Scale Hydrology Data Management
abstract
HydroTerre is a research prototype platform developed at Penn State for the hydrology community. It provides access to aggregated scientific data sets that are useful for hydrological modeling and research. HydroTerre's frontend is a web service, and a user query can request creation of a data bundle whose size can vary from a few megabytes to 100's of gigabytes. In this article, we present software tuning and optimization strategies for various hardware configurations of the HydroTerre platform. Our goal is to minimize access time to a wide range of data bundle creation queries from users. We use automated schemes to estimate the computational work required for various queries, and identify the best-performing hardware/software configuration. We hope this study is instructive for researchers developing similar data management cyberinfrastructure in other science and engineering fields.
Lorne Leonard, Kamesh Madduri, Christopher J. Duffy
IEEE Trans. Parallel Distributed Syst.3
2015 A Task-Centered Framework for Computationally-Grounded Science Collaborations
abstract
Collaboration is ubiquitous in today's science, yet there is limited support for coordinating scientific work. The general-purpose tools that are typically used (e.g., email, shared document editing, social coding sites), have still not replaced in-person meetings, phone calls, and extensive emails needed to coordinate and track collaborative activities. Scientists with diverse knowledge and skills around the globe could collaborate by opening scientific processes that expose all tasks and activities publicly to achieve a shared scientific question. This paper describes the Organic Data Science framework to support scientific collaborations that revolve around complex science questions that require significant coordination, entice contributors to remain engaged for extended periods of time, and enable continuous growth to accommodate new contributors as the work evolves over time. We discuss how the design of this framework incorporates principles followed by successful on-line communities. We present initial results to date of several communities that are collaborating using this framework.
Yolanda Gil, Felix Michel, Varun Ratnakar, Matheus Hauder, Christopher J. Duffy, Hilary Dugan, Paul C. Hanson
e-Science5
2015 Supporting Open Collaboration in Science Through Explicit and Linked Semantic Description of Processes
Yolanda Gil, Felix Michel, Varun Ratnakar, Jordan S. Read, Matheus Hauder, Christopher J. Duffy, Paul C. Hanson, Hilary Dugan
ESWC6
2010 An object-oriented shared data model for GIS and distributed hydrologic models
abstract
Distributed physical models for the space–time distribution of water, energy, vegetation, and mass flow require new strategies for data representation, model domain decomposition, a priori parameterization, and visualization. The geographic information system (GIS) has been traditionally used to accomplish these data management functionalities in hydrologic applications. However, the interaction between the data management tools and the physical model are often loosely integrated and nondynamic. This is because (a) the data types, semantics, resolutions, and formats for the physical model system and the distributed data or parameters may be different, with significant data preprocessing required before they can be shared; (b) the management tools may not be accessible or shared by the GIS and physical model; and (c) the individual systems may be operating system dependent or are driven by proprietary data structures. The impediment to seamless data flow between the two software components has the effect of increasing the model setup time and analysis time of model output results, and also makes it restrictive to perform sophisticated numerical modeling procedures (real-time forecasting, sensitivity analysis, etc.) that utilize extensive GIS data. These limitations can be offset to a large degree by developing an integrated software component that shares data between the (hydrologic) model and the GIS modules. We contend that the prerequisite for the development of such an integrated software component is a ‘shared data model’, which is designed using an object-oriented strategy. Here we present the design of such a shared data model taking into consideration the data type descriptions, identification of data classes, relationships, and constraints. The developed data model has been used as a method base for developing a coupled GIS interface to Penn State Integrated Hydrologic Model (PIHM), called PIHMgis.
Mukesh Kumar 0002, Gopal Bhatt, Christopher J. Duffy
Int. J. Geogr. Inf. Sci.3
2009 An efficient domain decomposition framework for accurate representation of geodata in distributed hydrologic models
Mukesh Kumar 0002, Gopal Bhatt, Christopher J. Duffy
Int. J. Geogr. Inf. Sci.3
2004 Enhancing the performance of feature selection algorithms for classifying hyperspectral imagery
abstract
A method for enhancing the performance of feature selection algorithms is proposed. The proposed method is a two step process - first a feature subset is selected with optimum mutual information content and then this subset is searched to find a smaller subset, which has the best separability between classes. A subset with "optimum" mutual information content is the one which contains most of the information that is present in the rest of set. An expression has been derived to find such a subset efficiently. The two-step process is shown to reduce the search space drastically. The method is implemented with a simple genetic algorithm (SGA) and tested using hyperspectral remote-sensing images (acquired by AVIRIS sensor) as a data set. Theoretical result shows that the proposed method reduces the computation load by 90%. A computational efficiency to the order /spl sim/20% is obtained on the implementation of proposed method with SGA. The method is sufficiently general to be used to enhance other feature selection algorithms.
Mukesh Kumar 0002, Christopher J. Duffy, Patrick M. Reed
IGARSS2