Rahul Ramachandran

dblp:91/5457 · DBLP profile ↗
← Back
44ranked-venue papers
9as first author
15since 2021 · last 2025
0000-0002-0647-1941ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 36 · 9 first-author · 11 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Multi-Layer Agent-Based Spatiotemporal UoW Recommendation for Workflow Composition
abstract
Software service discovery and recommendation help data scientists build scientific workflows - multi-step data analytics procedures - by automating the manual selection of services. Previous research shows that recommending chainable units of work (UoWs), rather than individual services, improves efficiency and reduces data shimming issues. However, UoW recommendation remains an NP-hard problem. To tackle this challenge, this study introduces a novel framework tailored to recommend UoWs in a goal-driven, context-aware manner, thereby facilitating workflow development. The framework is built around layered structure of software service social networks. At its foundation lies a service dependency network, where each edge represents a dependency between a two-service UoW within a specific context. The next layer abstracts each of these edges into a node, with new edges now representing three-service UoWs. This layering process continues iteratively, with subsequent layers capturing increasingly complex UoWs at higher levels of granularity. At high-order layers, UoW nodes are clustered based on their semantic embeddings, with each cluster represented by an intelligent agent. This approach transforms the workflow recommendation problem into a multi-agent collaboration task, where agents work together to identify high-level UoW groupings before refining selections by navigating down the layered structure for finer-grained recommendations. Experimental results over a real-world dataset confirm the effectiveness of the proposed framework in enhancing workflow composition efficiency.
Xihao Xie, Jia Zhang 0001, Rahul Ramachandran, Tsengdar J. Lee, Seungwon Lee 0005
SSE4
2025 TerraMind: Large-Scale Generative Multimodality for Earth Observation
abstract
We present TerraMind, the first any-to-any generative, multimodal foundation model for Earth observation (EO). Unlike other multimodal models, TerraMind is pretrained on dual-scale representations combining both token-level and pixel-level data across modalities. On a token level, TerraMind encodes high-level contextual information to learn cross-modal relationships, while on a pixel level, TerraMind leverages fine-grained representations to capture critical spatial nuances. We pretrained TerraMind on nine geospatial modalities of a global, large-scale dataset. In this paper, we demonstrate that (i) TerraMind's dual-scale early fusion approach unlocks a range of zero-shot and few-shot applications for Earth observation, (ii) TerraMind introduces "Thinking-in-Modalities" (TiM) -- the capability of generating additional artificial data during finetuning and inference to improve the model output -- and (iii) TerraMind achieves beyond state-of-the-art performance in community-standard benchmarks for EO like PANGAEA. The pretraining dataset, the model weights, and our code are open-sourced under a permissive license.
Johannes Jakubik, Felix Yang, Benedikt Blumenstiel, Erik Scheurer, Rocco Sedona, Stefano Maurogiovanni, Jente Bosmans, Nikolaos Dionelis, Valerio Marsocci, Niklas Kopp, Rahul Ramachandran, Paolo Fraccaro, Thomas Brunschwiler, Gabriele Cavallaro, Juan Bernabé-Moreno, Nicolas Longépé
ICCV11
2024 Enabling Dynamic Data Governance in Science: Design, Implementation, and Future Directions of the Modern Data Governance Framework
abstract
As scientific data volumes exponentially grow, dynamic, flexible and open approaches to data governance are needed. In this paper, we describe our efforts to build an open, scientific Modern Data Governance Framework (mDGF) that streamlines and makes actionable data governance requirements for projects and data providers. We present the goals and design of the mDGF. We also share our envisioned usage for the mDGF and planned future work.
Kaylin M. Bugbee, Rahul Ramachandran, Aaron Kaulfus, Jeanné le Roux, Ge Peng, Deborah K. Smith, Iksha Gurung, Ashish Acharya, Jerika Christman
IGARSS2
2024 Curating AI-Ready Datasets for Equity and Environmental Justice: A Data-Centric AI Case Study
abstract
An equitable and environmentally just community is essential in order to avoid disproportionate burden borne by vulnerable communities. This need becomes pressing in the aftermath of an extreme event such as disaster or hazard when it is difficult for the governing bodies to implement resource allocation as per the need. Artificial Intelligence (AI) algorithms can help surface Equity and Environmental Justice (EEJ) issues when trained on EEJ datasets. However, curating AI-ready EEJ training datasets is challenging due to differences in factors such as heterogeneity, resolution, modality, and level of expertise in labeling. Additionally, EEJ issues involve sensitive information where uncertainties and errors could degrade the performance of AI algorithms. For eg. error in seasonal crop yield information can highly effect the prediction of annual crop yield. To address these challenges, Data-centric AI (DCAI) methods are employed, which enhance AI algorithm performance even with limited training samples. DCAI prioritizes data quality, thereby reducing the adverse effects of uncertainties and errors during the model training process. This research proposes a novel dataset and benchmark for analyzing the effect of the Maui Wildfire of 2023 for Equity and Environmental Justice (EEJ) issues. The proposed dataset aligns with the concepts of DCAI such as annotation quality, data preprocessing, privacy, feature engineering, governance and provenance. The proposed AI-ready dataset is available on HuggingFace at https://huggingface.co/datasets/nasa-impact/ml4ej-maui-wildfire.
Paridhi Parajuli, Rajat Shinde, Iksha Gurung, Manil Maskey, Rahul Ramachandran
IGARSS5
2024 RANGER: Context-Aware Service Unit of Work Recommendation for Incremental Scientific Workflow Composition
Xihao Xie, Jia Zhang 0001, Rahul Ramachandran, Tsengdar J. Lee, Seungwon Lee 0005
WISE (3)4
2023 Blaze: A High-Performance, Scalable, and Efficient Data Transfer Framework with Configurable and Extensible Features : Principles, Implementation, and Evaluation of a Transatlantic Inter-Cloud Data Transfer Case Study
abstract
Blaze is a high-speed data transfer framework that enables efficient and scalable data movement between distributed storage systems. In this paper, we describe the design, implementation, and evaluation of Blaze in the context of a case study involving the transfer of 5.6 petabytes of data from OVH cloud storage to an Amazon S3 bucket. We discuss the technical challenges and design choices that led to creating a single-agent architecture, providing flexibility in agent placement while minimizing operational costs. We also demonstrate the orchestration capabilities of Blaze using Apache Airflow to manage hierarchical workflows for data transfer between storage systems. Our evaluation shows that Blaze achieved the expected 20 Gbps throughput during data transfer and provided significant cost savings compared to other architectures. The results demonstrate that Blaze is a practical solution for high-speed data transfer in large-scale distributed storage environments.
Suresh Marru, Brian Freitag, Dimuthu Wannipurage, Uday Kumar Bommala, Patrick Pradier, Christophe Demange, Nishan Pantha, Tathagata Mukherjee, Betlem Rosich, Eric Monjoux, Rahul Ramachandran
CLOUD11
2023 A Framework for Large Scale Semantic Similarity Search on Satellite Imagery
abstract
Searching for Earth Science phenomena in large archives of Earth Observation Satellite Imagery data requires elaborate processing and spatio-temporal indexing of the images into categories of the said phenomena. Manual tagging is laborious as it needs constant monitoring through vast volumes of satellite data, the volume and velocity of which is ever-increasing. A complete re-indexing is also needed when a new phenomenon of interest is to be searched through the data archive. Previous efforts to automate tagging have leveraged Machine Learning (ML) techniques to classify images into phenomena of interest. In this method, multiple ML algorithms, each specifically trained for detecting a particular phenomenon, are used for spatio-temporal indexing. While doing so negates the need for human indexing, the process of creating ML models for identifying a class of phenomena involves significant time and computation overhead. Moreover, ML algorithms require vast amounts of extremely scarce labeled data. Furthermore, the computation needed for re-indexing the data whenever a new phenomenon is added to be tagged is not negligible. We propose an alternative, data-driven framework to search through vast amounts of satellite data, that eliminates the need for manual indexing, labeling, or creating purpose-built ML classifiers. The proposed method leverages Self-Supervised Learning (SSL) techniques to obtain feature vectors that are used for search and retrieval of satellite images. An Approximate Nearest Neighbors (ANN) algorithm is used to cluster and retrieve images exhibiting similar features, and by extension, similar Earth Science phenomena. Our unique contribution in this work is the orchestration of the methodology with various cloud services that facilitates searching through millions of images within a short span of time. To showcase the framework, we created a web interface to search through 21 years worth of daily satellite imagery with global coverage. In this paper, we discuss the progress we have made in enabling Embedding Based Search within Remote Sensing, and discuss the potential benefits and pitfalls involved in realizing this method. We also aim to provide insights and experiences we documented while developing such a system along with potential limitations of the current stage of the framework.
Muthukumaran Ramasubramanian, Iksha Gurung, Leo Thomas, Kathryn Berger, Soumya Ranjan, Heidi Mok, Sowmya Subramanian, Vitor George, Manil Maskey, Rahul Ramachandran
IGARSS10
2023 Generative Framework Approach to Match Landsat and Sentinel-2 Data
abstract
Linear regression and histogram matching-based techniques have been widely used to minimize the surface reflectance difference between two similar satellite observations such as Landsat-8/9 and Sentinel-2A/B products. However, regionally or globally derived conversion factors may not be suitable for all land cover types and locations, resulting in noticeable residual differences between the sensors. Generative Adversarial Network (GAN) has shown promise in the field of image processing for domain or style transfer. In this work we aim to minimize the surface reflectance difference between Landsat and Sentinel-2 products based on GAN.This work selected 26 pairs of same-day Landsat and Sentinel-2 30-meter spatial resolution surface reflectance images in the green spectral band from NASA’s Harmonized Landsat/Sentinel-2 project (HLS), with each pair having over 90% spatial overlap. Upstream pre-processing steps included atmospheric correction, cloud masking and BRDF normalisation. The generator architecture was based on U-Net and discriminator as PatchGAN. GAN was trained for 43000 iterations. Finally, the Landsat images generated from the Sentinel-2 images by GAN are compared to the original Landsat 8/9 images in terms of SSIM and MSE metrics. Result showed a higher SSIM score for the generated data, which can be interpreted to mean that the generated images were closer to the real Landsat images. And, overall MSE for generated data was lower than that for the original Sentinel-2 images.This study, for the first time, reports a GAN-based spatial matching between Landsat8/9 and Sentinel-2 surface reflectance images. The results indicate that this approach has the potential to map data between the two satellite images with reasonable accuracy. This method may prove to be more robust and could be applied globally, potentially replacing the previous approach of simulated data. The method can be further extended with more data and in other spectral bands and potentially be used in HLS processing.
Sujit Roy, Madhu Sridhar, John Mandel, Brian Freitag, Junchang Ju, Rahul Ramachandran
IGARSS6
2023 Observing Supraglacial Lakes Using Deep Learning and Planetscope Imagery
abstract
Supraglacial lakes (SGL)s result from melt water accumulation in topographic depressions on the surface of glaciers. SGLs primarily affect glacial dynamics through a positive feedback loop in which the albedo-lowering effect of SGLs can escalate surface melt leading to increases in lake extent and depth, amplifying the aforementioned albedo-lowering effect. The implications of accelerated glacial melt include increased sea level rise and modifications to ocean primary productivity. SGLs are critical indicators of surface melt and its downstream impacts and should be monitored efficiently. In situ observations and measurements of SGLs are time consuming, cost-prohibitive and difficult to scale. Earth observation data and machine learning enable scalable monitoring of SGLs through pattern detection and quantification of lake evolution over time [1]. This work presents a model developed by training a convolutional neural network with imagery and labels from NASA Operation IceBridge and predicting SGLs in high temporal and spatial resolution PlanetScope imagery.
Lillianne Thomas, Slesa Adhikari, Iksha Gurung, Aaron Kaulfus, Muthukumaran Ramasubramanian, Manil Maskey, Rahul Ramachandran
IGARSS7
2022 Learning Context-Aware Service Representation for Service Recommendation in Workflow Composition
abstract
As increasingly more software services have been published onto the Internet, it becomes critical yet highly challenging to recommend suitable services to facilitate scientific workflow composition. This paper proposes a novel Natural Language Processing (NLP)-inspired approach to recommending services throughout a workflow development process, based on incrementally learning latent service representation from workflow provenance. A work-flow composition process is formalized as a step-wise, context-aware service selection procedure, which is mapped to next-word prediction in a natural language sentence generation. Historical service dependencies are extracted from workflow provenance to build and enrich a knowledge graph. Each path in the knowledge graph reflects a scenario in a data analytics experiment, which is analogous to a sentence in a conversation. All paths are thus formalized as composable service sequences and are mined, using various patterns, from the established knowledge graph to construct a corpus. Service embeddings are then learned by applying deep learning model from the NLP field. Extensive experiments on the real-world dataset demonstrate the effectiveness and efficiency of the approach.
Xihao Xie, Jia Zhang 0001, Rahul Ramachandran, Tsengdar J. Lee, Seungwon Lee 0005
ICIS3
2022 Selecting Approaches for Enabling Enterprise Data Search: NASA's Science Mission Directorate (SMD) Catalog
abstract
NASA's Science Mission Directorate (SMD) is working to build an open-source science infrastructure to accelerate open, collaborative and interdisciplinary science. One key component in the open-source science infrastructure is the SMD data catalog. In this paper, we present our process for selecting a technical approach to building a NASA SMD enterprise-wide integrated search capability for science users across multiple science disciplines to support discovery and access to complex scientific data.
Kaylin M. Bugbee, Rahul Ramachandran, Ashish Acharya, Dai Hai Ton That, John Hedman, Ahmed Eleish, Charles Driessnack, Wesley Adams, Emily Foshee
IGARSS2
2022 Artificial Intelligence Vis-à-Vis Data Systems
abstract
NASA Earth Science Data Systems (ESDS) program manages a full lifecycle of data collected by all Earth science missions. ESDS also develops capabilities optimized to support rigorous science investigations. As technology landscapes evolve, ESDS has also evolved to transform its internal services and enhance external user centric services. This paper describes how ESDS is (i) adopting artificial intelligence (AI) technology to improve core services and (ii) enabling its users to advance AI driven research and build applications.
Manil Maskey, Rahul Ramachandran, Iksha Gurung, Muthukumaran Ramasubramanian, Anirudh Koul
IGARSS2
2022 Language Model for Earth Science: Exploring Potential Downstream Applications as well as Current Challenges
abstract
The use of deep learning techniques to build transformer language models such as SciBERT and GPT3 have transformed the natural language technology (NLT) landscape. These new NLTs are being used in speech to text and vice versa, automated text classification, sentiment analysis, topic modeling, text summarization, and cognitive assistants. While Earth science has no shortage of unstructured data such as journal and conference papers, little efforts have focused on harnessing NLTs for knowledge extraction and supporting the scientific process. This paper surveys the use of language models in different science. BERT-E, a new Earth science-specific language model, is presented. BERT-E is generated using a transfer learning solution. A language model that has already been trained for general Science (SciBERT) is fine-tuned using abstracts and full text extracted from various Earth science-related articles. A downstream keywords classification application is used for evaluation, and the use of BERT-E shows improved performance. The need to develop a robust set of benchmarks in evaluating the language model such as BERT-E is discussed. Finally, example applications are presented to inspire additional ideas for applications using domain-specific language models.
Rahul Ramachandran, Muthukumaran Ramasubramanian, Prasanna Koirala, Iksha Gurung, Manil Maskey
IGARSS1
2022 Goal-Driven Context-Aware Service Recommendation for Mashup Development
abstract
As service-oriented architecture becoming one prevalent technique to rapidly compose functionalities to customers, increasingly more reusable software components have been published online in the form of web services. To create a mashup, however, it gets not only time-consuming but also error-prone for developers to find suitable services components from such a sea of services. Service discovery and recommendation has thus attracted significant momentum in both academia and industry. This paper proposes a novel incremental recommend-as-you-go approach to recommending next potential service based on the context of a mashup under construction, considering services that have been selected up to the current step as well as the mashup goal. The core technique is an algorithm of learning the embedding of services, which learns their past goal-driven context-aware decision making behaviors in addition to their semantic descriptions and co-occurrence history. A goal exclusionary negative sampling mechanism tailored for mashup development is also developed to improve training performance. Extensive experiments on a real-world dataset demonstrate the effectiveness of this approach.
Xihao Xie, Jia Zhang 0001, Rahul Ramachandran, Tsengdar J. Lee, Seungwon Lee 0005
SNPD3
2021 Augmenting Data Systems with Prediction based Embeddings
abstract
One of the challenges of improving the search and use of complex Earth science data is designing and incorporating semantic components in existing Earth science data systems. Many projects have addressed this by using a knowledge engineering approach. However, using ontologies has inherent limitations as a practical and scalable approach. Data-driven strategies based on natural language processing, coupled with Machine Learning, provide an alternative approach. Data-driven approaches utilize existing corpus available as unstructured text. This paper describes a hybrid strategy that uses a data-driven approach to build an embedding from a large corpus of Earth science journal publications while leveraging existing ontologies to develop validation tests to evaluate the embedding's robustness and correctness. The paper also describes the use of this embedding in two different applications. The first application provides a semantic mapping service to bridge the gap between a science application need and the appropriate instruments or datasets required to address that need. The second application is keyword recommender to make the data set tagging process efficient for the data operators and ensure keyword consistency within a data catalog.
Rahul Ramachandran, Muthukumaran Ramasubramanian, Iksha Gurung, Carson Davis, Derek Koehl, Manil Maskey, Tsengdar J. Lee
IGARSS1
2020 Standardized Algorithm Documentation for Improved Scientific Data Understanding: The Algorithm Publication Tool Prototype
abstract
Algorithm Theoretical Basis Documents (ATBDs) are documents which accompany Earth observation data products generated from algorithms. While ATBDs are essential to scientific reproducibility, these key documents are not standardized and are often difficult to find. In this paper, we present the prototype Algorithm Publication Tool (APT), a cloud-based ATBD authoring and editing tool for NASA's Earth science data systems. A standardized ATBD information model is also described as well as lessons learned from developing the prototype tool.
Kaylin M. Bugbee, Aaron Kaulfus, Alyssa Harris, Sean Bailey, Rahul Ramachandran, Sean Harkins, Aimee Barciauskas, Deborah K. Smith
IGARSS5
2020 Advancing Open Science Through Innovative Data System Solutions: The Joint ESA-NASA Multi-Mission Algorithm and Analysis Platform (MAAP)'s Data Ecosystem
abstract
Collaborative open science practices are changing the way research is conducted. These changes affect how scientists work together on data, code and information. Data systems enhance open science by offering forward thinking technological solutions, such as providing data and computation on the cloud, to enable collaboration, sharing and analysis. In this paper, we present our vision for a conceptual data system on the cloud that enables open science. We also present our work on the Multi-Mission Algorithm and Analysis Platform (MAAP) which has served as a pathfinder data system for this conceptual approach.
Kaylin M. Bugbee, Rahul Ramachandran, Manil Maskey, Aimee Barciauskas, Aaron Kaulfus, Dai Hai Ton That, Katrina Virts, Kel N. Markert, Christopher Lynnes
IGARSS2
2020 Employing Deep Learning to Enable Visual Exploration of Earth Science Events
abstract
Earth science data archives have significantly increased in size due to the number of advanced sensors and science missions. In the meantime, Earth science data systems have not taken advantage of data driven technologies to provide advanced search capabilities. This paper discusses a machine learning-based approach, an enabling data driven technology, to detect Earth science events from image archives. The automated event detection is cataloged in an event database that provides a novel way to explore large archives of data. In addition, a phenomena portal to visually explore events and contextual information is discussed.
Manil Maskey, Rahul Ramachandran, Iksha Gurung, Muthukumaran Ramasubramanian, Brian Freitag, Aaron Kaulfus, Georgios Priftis, Drew Bollinger, Ricardo Mestre
IGARSS2
2020 A Quantitative Analysis on the Use of Supervised Machine Learning in Earth Science
abstract
Recent review papers have discussed the opportunities and challenges of applying machine learning (ML) techniques to Earth science data. A common challenge cited in these papers is the lack of labeled training data. A literature review of Earth science papers over the last 10 years demonstrates that while there is rapid adoption of ML, particularly in biogeoscience and land surface research, the training datasets typically contain only hundreds of samples. This lack of training data limits the use of deep learning algorithms, which require larger volumes of labeled data. In situ training data are most frequently used in almost all domains, followed by model output and satellite data. The atmosphere and solid Earth domains use the largest training datasets, an order of magnitude larger than in biogeoscience papers. Random forest is the most commonly applied ML algorithm in all domains except atmospheric science and biogeoscience, which more frequently use fully connected neural networks.
Katrina Virts, Ashlyn Shirey, George Priftis, Kumar Ankur, Muthukumaran Ramasubramanian, Hassan Muhammad, Ashish Acharya, Rahul Ramachandran
IGARSS8
2019 Alternative Datasets for Identification of Earth Science Events and Data
abstract
Alternative, or non-traditional, data sources can be used to generate datasets which can in turn be analyzed for temporal, spatial and climatological patterns. Events and case studies inferred from the analysis of these patterns can be used by the remote sensing community to more effectively search for Earth observation data. In this paper, we present a new alternative Earth science dataset created from the National Weather Service's Area Forecast Discussion (AFD) documents. We then present an exploratory methodology for identifying interesting climatological patterns within the AFD data and a corresponding motivating example as to how these data and patterns can be used to search for relevant events or case studies.
Kaylin M. Bugbee, Robert Griffin, Brian Freitag, Jeffrey J. Miller, Rahul Ramachandran, Jia Zhang 0001
IGARSS5
2019 Building a Data Ecosystem: A New Data Stewardship Paradigm for the Multi-Mission Algorithm and Analysis Platform (MAAP)
abstract
New adaptive approaches to Earth observation data stewardship need to be adopted in order to allow for higher data volumes, heterogeneous data and constantly evolving technologies. The data ecosystem approach to stewardship offers a viable solution to this need by placing an emphasis on the relationships between data, technologies and people. In this paper, we present the Joint ESA-NASA Multi-Mission Algorithm and Analysis Platform's (MAAP) creation of a data ecosystem to support global aboveground terrestrial carbon dynamics research. We present the components needed to support the MAAP data ecosystem along with two data stewardship workflows used in the MAAP and the development of extended metadata for MAAP.
Kaylin M. Bugbee, Christopher Lynnes, Manil Maskey, Aimee Barciauskas, Rahul Ramachandran, Aaron Kaulfus, Jeanné le Roux, Jeffrey J. Miller, Iksha Gurung, Amanda S. Whitehurst
IGARSS5
2019 Machine Learning Lifecycle for Earth Science Application: A Practical Insight into Production Deployment
abstract
Enterprises are making machine learning for production as an integral part of their future roadmaps and Earth science domain is no exception. However, there is common problem in transitioning machine learning from science to production due to a major difference in constructing a model versus deploying it for people to use to make decisions. Phases of machine learning lifecycle that includes model transition to production using a successful application is discussed.
Manil Maskey, Andrew L. Molthan, Chris Hain, Rahul Ramachandran, Iksha Gurung, Brian Freitag, Jeffrey J. Miller, Muthukumaran Ramasubramanian, Drew Bollinger, Ricardo Mestre, Daniel Cecil
IGARSS4
2019 Applying Deep Learning to Hail Detection: A Case Study
abstract
Deep learning is a subset of machine learning that uses deep neural networks (DNNs) capable of learning representations and extracting valuable information from vast data sets. Similarly, weather phenomena are often identified by patterns in data that serve as precursor signatures. Therefore, deep learning networks can be used to identify signatures of the weather phenomena, or possibly signatures not yet established by forecasters in addition to aiding forecasters in synthesizing the growing amount of meteorological observations. In this article, we demonstrate the value of deep learning for atmospheric science applications by providing a proof of concept, using deep learning for the detection of hail-bearing storms as a test case study. The deep learning network presented in this article obtains a higher precision when presented with multisource data and is able to identify a common feature associated with hail storms-decreased infrared brightness temperatures. This network and case study illustrate the capability of deep networks for the detection of weather phenomena and contribute to the growing awareness of deep learning among atmospheric scientists.
Melinda Pullman, Iksha Gurung, Manil Maskey, Rahul Ramachandran, Sundar A. Christopher
IEEE Trans. Geosci. Remote. Sens.4
2018 ESA-NASA Multi-Mission Analysis Platform for Improving Global Aboveground Terrestrial Carbon Dynamics
abstract
In the context of innovative sensors and a changing ground segments, the concept of ESA-NASA multi-Mission Analysis Platform dedicated to the NISAR, GEDI and Biomass missions is proposed. This analysis platform will be a virtual open and collaborative environment. The goal is to bring together data centre (Earth Observation and non-Earth Observation data), computing resources and hosted processing, collaborative tools (processing tools, data mining tools, user tools, ...), concurrent design and test bench functions, application shops and market place functionalities, accounting tools to manage resource utilisation, communication tools (social network) and documentation.
Clement Albinet, Amanda S. Whitehurst, Henri Laur, Kevin J. Murphy, Bjorn Frommknecht, Klaus Scipal, Andrew E. Mitchell, Benhan Jai, Rahul Ramachandran
IGARSS9
2018 Generalizing a Data Analysis Pipeline in the Cloud to Handle Diverse Use Cases in NASA's EOSDIS
abstract
NASA's Earth Observing System Data and Information System (EOSDIS) is tasked with archiving and distributing Earth Observation data across a range of disciplines, including atmospheric science, oceanography, land processes, natural hazards, solar radiance and even socioeconomic aspects relating to the environment. Driven by rapidly rising data volumes, EOSDIS is migrating to a cloud computing based archive over the next few years. Although this simplifies data management somewhat, the main aim is to provide the data in an environment where end users can bring their analysis to the data rather than attempting to download and manage ever-increasing volumes. To that end, a cloud-based analysis platform is being constructed to enable data transformations, analyses and visualization without egressing the data from the cloud. In this endeavor, we expect a wide variety of users, algorithms and use cases. Consequently, the architecture of this cloud analytics platform is expressly designed to be based on open services, thus fostering an ecosystem that enables the efficient combination of common components with data-specific or analysis-specific components.
Christopher Lynnes, Rahul Ramachandran
IGARSS2
2018 Earth Science Deep Learning: Applications and Lessons Learned
abstract
Deep learning has revolutionized computer vision and natural language processing with various algorithms scaled using high-performance computing. The Data Science and Informatics Group (DSIG) at the NASA Marshall Space Flight Center (MSFC), has been using deep learning for a variety of Earth science applications. This paper provides examples of the applications and also addresses some of the challenges that have been encountered.
Manil Maskey, Rahul Ramachandran, Jeffrey J. Miller, Jia Zhang 0001, Iksha Gurung
IGARSS2
2018 Tropical Cyclone Intensity Estimation Using a Deep Convolutional Neural Network
abstract
Tropical cyclone intensity estimation is a challenging task as it required domain knowledge while extracting features, significant pre-processing, various sets of parameters obtained from satellites, and human intervention for analysis. The inconsistency of results, significant pre-processing of data, complexity of the problem domain, and problems on generalizability are some of the issues related to intensity estimation. In this study, we design a deep convolutional neural network architecture for categorizing hurricanes based on intensity using graphics processing unit. Our model has achieved better accuracy and lower root-mean-square error by just using satellite images than 'state-of-the-art' techniques. Visualizations of learned features at various layers and their deconvolutions are also presented for understanding the learning process.
Ritesh Pradhan, Ramazan Savas Aygün, Manil Maskey, Rahul Ramachandran, Daniel Cecil
IEEE Trans. Image Process.4
2017 A Fine-Grained API Link Prediction Approach Supporting Mashup Recommendation
abstract
Service (API) discovery and recommendation is key to the wide spread of service oriented architecture and service oriented software engineering. Service recommendation typically relies on service linkage prediction calculated by the semantic distances (or similarities) among services based on their collection of inherent attributes. Given a specific context (mashup goal), however, different attributes may contribute differently to a service linkage. In this paper, instead of training a model for all attributes as a whole, a novel approach is presented to simultaneously train separate models for individual attributes. Meanwhile, a latent attribute modeling method is developed to reveal context-aware attribute distribution. Experiments over real-world datasets have demonstrated that this fine-grained method yields higher link prediction accuracy.
Qihao Bao, Jia Zhang 0001, Xiaoyi Duan, Rahul Ramachandran, Tsengdar J. Lee, Yankai Zhang, Seungwon Lee 0005, Patrick Gatlin, Manil Maskey
ICWS4
2017 Linking Design-Time and Run-Time: A Graph-Based Uniform Workflow Provenance Model
abstract
Workflow is an important way to mashup reusable software services to create value-added data analytics services. Workflow provenance is core to understand how services and workflows behaved in the past, which knowledge can be used to provide a better recommendation. Existing workflow provenance management systems handle various types of provenance separately. A typical data science exploration scenario, however, calls for an integrated view of provenance and seamless transition among different types of provenance. In this paper, a graph-based, uniform provenance model is proposed to link together design-time and run-time provenance, by combining retrospective provenance, prospective provenance, and evolution provenance. Such a unified provenance model will not only facilitate workflow mining and exploration, but also facilitate workflow interoperability. The model is formalized into colored Petri nets for verification and monitoring management. A SQL-like query language is developed, which supports basic queries, recursive queries, and cross-provenance queries. To verify the effectiveness of our model, A web-based, collaborative workflow prototyping system is developed as a proof-of-concept. Experiments have been conducted to evaluate the effectiveness of the proposed SQL-like graph query against SQL query.
Xiaoyi Duan, Jia Zhang 0001, Qihao Bao, Rahul Ramachandran, Tsengdar J. Lee, Seungwon Lee 0005
ICWS4
2017 CUMULUS: NASA's cloud based distributed active archive center prototype
abstract
Since 1994, NASA's Earth Science Data System (ESDS) Program, central to facilitating the implementation of NASA's Earth Science strategic plan, has committed to the full and open, public sharing of Earth science data obtained from NASA instruments. A vital responsibility of the ESDS Program is to continuously evolve the entire data and information system (“EOSDIS”) to maximize returns on NASA's collected data. An independent, holistic review of the EOSDIS was conducted to identify gaps and resulted in the following recommendations: investigate (a) whether commercial cloud providers offer potential for storage, processing, and operational efficiencies, and (b) the potential development of new data access and analysis paradigms. In response, ESDS has initiated several prototypes, including the “Cumulus” prototype, to investigate the advantages and risks of leveraging cloud computing. Cumulus is being designed and developed as a “native” cloud-based data ingest, archive and management system that can be used for all future NASA Earth science data streams.
Rahul Ramachandran, Katie Baynes, Kevin J. Murphy, Alireza Jazayeri, Ian Schuler, Dan Pilone
IGARSS1
2016 Snowstorm climatology derived from NASA MERRA reanalysis as an example for event-based virtual collections
abstract
Most of the climatological studies derived from reanalysis datasets to-date have been presence-based rather than event-based. We have gathered event-based climatological statistics from thirty-seven-plus (37+) years of blizzard-like snowstorms, identified and individually tracked, using hourly high-resolution datasets from the NASA's Modern Era Retrospective-analysis for Research and Applications (MERRA) data collection. We have not only extracted summary statistics for all storms, such as cumulative-probability density functions (CDFs, in percentiles) of storm duration and cumulative area coverage, but also per-event statistics for each storm, e.g. beginning/ending times, hourly locations, hourly mean snowfall intensities, and snowfall intensity probability distribution. In addition, we have constructed a web portal where users can view the hourly locations of each snowstorm annotated with conditions of the storm at that hour. Users with accounts on the portal can discover coincident data granules of relevant satellite remote-sensing observations by querying NASA metadata repository, i.e. EOS Clearing House (ECHO) or upcoming Common Metadata Repository (CMR) and create/share individualized virtual collections apposite to their research.
Kwo-Sen Kuo, Kush Shrestha, Amy Lin, Rahul Ramachandran
IGARSS4
2016 Exploiting dark information resources to create new value added services to study Earth science phenomena
abstract
This paper presents two research applications exploiting unused metadata resources in novel ways to aid data discovery and exploration capabilities. The results based on the experiments are encouraging and each application has the potential to serve as a useful standalone component or service in a data system.
Rahul Ramachandran, Manil Maskey, Xiang Li 0043, Kaylin M. Bugbee
IGARSS1
2015 Linking from observations to data to actionable science in the climate data initiative
abstract
A tremendous amount of Earth and Climate related data and information are available from the U.S. Federal government. Among the proposed actions of the President's Climate Action Plan are several efforts to foster the use of existing data to encourage development of additional data products and tools that can be used to improve community resilience and prepare for the impacts of climate change. Building on previous efforts to organize the presentation of this material from Federal web pages and data centers, the Climate Data Initiative is working together with other related efforts to link Earth observation systems through to data resulting from them and on to related web pages, case studies, decision making tools, and other relevant content. Often such information is not located in a single web site, data center, or even a single agency, but distributed across the Federal Government. Linking such information across the breadth of interagency holdings can increase understanding of the complexity of those holdings and their inter-relationships and allow a more cohesive presentation of all of the material.
Curt Tilmes, Ana Pinheiro Privette, Jeffrey Chen, Rahul Ramachandran, Kaylin M. Bugbee, Robert E. Wolfe
IGARSS4
2013 Introducing Provenance Capture into a Legacy Data System
abstract
Accurate provenance information facilitates improved understanding of Earth science data and scientific reproducibility and can serve as an indicator of data quality. Provenance capture is an integral part of many modern workflow systems but may not have been considered in the design of legacy data production systems. Furthermore, in addition to data lineage, it is also important to capture contextual information needed for understanding how a data set was produced. This paper describes our experience in retrofitting a legacy data system to support capture, storage, and dissemination of provenance. Data inputs and transformations are logged automatically, while broader context information describing science algorithms and ancillary files is manually compiled. Provenance and context information are integrated for interactive user access and embedded into data files as XML documents compliant with the “Lineage” specification for geographic metadata defined by the International Organization for Standardization in the ISO 19115-2 standard. Lessons learned from this approach can inform others who need to incorporate provenance into a data system after the fact.
Helen Conover, Rahul Ramachandran, Bruce Beaumont, Ajinkya Kulkarni, Michael McEniry, Kathryn Regner, Sara J. Graves
IEEE Trans. Geosci. Remote. Sens.2
2009 GLIDER: A Comprehensive Software Tool to Visualize, Analyze and Mine Satellite Imagery
abstract
There is a dearth of software tools that allow users to easily visualize, analyze and mine satellite imagery. The few tools that are available are expensive commercial packages that provide limited functionality. As part of a NASA funded project, a software tool named GLIDER is currently being developed to fill this void. GLIDER allows users to visualize and analyze satellite data in its native sensor view. Users can enhance the image by applying different image processing algorithms on the data. GLIDER provides the users with a full suite of pattern recognition and data mining algorithms that can be applied to the satellite imagery to extract thematic information. The suite of algorithms includes both supervised and unsupervised classification algorithms. In addition, users can project satellite imagery and analysis/mining results onto a 3D globe for visualization. GLIDER also allows users to add additional layers to the globe along with the projected image. Users can open multiple views within GLIDER to manage, visualize and analyze many data files all at once. This paper describes the features of GLIDER version 1.0.
Rahul Ramachandran, Sara J. Graves, Todd Berendes, Manil Maskey, C. Chidambaram, Sundar A. Christopher, P. Hogan, Tom Gaskins
IGARSS (3)1
2009 Talkoot Software Appliance for Collaborative Science
abstract
On the emerging ¿Social Web,¿ millions of people offer their knowledge online in a collective knowledge system comprising an active community of motivated members posting problems and solutions in blogs, forums, mailing lists, collaborative portals and other Web 2.0 technologies. These technologies complement formal means of sharing knowledge via conferences and published papers, where it is impossible to share all the research details, and where negative results are rarely included. A small but growing number of scientists and researchers are beginning to harness these Web 2.0 technologies as a transformative way of doing science. With the advent of service oriented architectures, the model of chaining services to create analysis workflows provides the research community unprecedented opportunity to collaborate, sharing their workflows with one another, reproducing and analyzing research results, and leveraging colleagues' expertise to expedite the process of scientific knowledge discovery. A crucial component needed for this unprecedented level of cooperation within the research community is a reusable, extensible and customizable environment for building collaborative ¿open science¿ portals for managing these shared analysis workflows. This paper describes the design and the development of Talkoot, a customizable ¿software appliance¿ to build collaborative portals for earth science services and analysis workflows.
Rahul Ramachandran, Sunil Movva, Helen Conover, Christopher Lynnes
IGARSS (5)1
2008 Mining Scientific Data using the Internet as the Computer
abstract
This paper describes approaches and methodologies facilitating the analysis of large amounts of distributed scientific data. The existence of full-featured analysis tools, such as the Algorithm Development and Mining (ADaM) toolkit and online data repositories now provide easy access and analysis capabilities to large amounts of data. However, there are obstacles to getting the analysis tools and the data together in a workable environment. Does one bring the data to the tools or deploy the tools close to the data? The large size of many current Earth science datasets incurs significant overhead in network transfer for analysis workflows, even with the current advanced networking capabilities. We are developing two solutions for this problem that address different analysis scenarios. The first is a Data Center Deployment of the analysis services for large data selections, orchestrated by a remotely defined analysis workflow. The second is a Data Mining Center approach of providing a cohesive analysis solution for smaller subsets of data. The two approaches can be complementary and thus provide flexibility for researchers to exploit the best solution for their data requirements.
Sara J. Graves, Rahul Ramachandran, Christopher Lynnes, Manil Maskey, Ken Keiser, Long Pham
IGARSS (4)2
2008 Intelligent Data Thinning Algorithms for Satellite Imagery
abstract
This paper presents a study on intelligent data thinning for satellite data. In particular, the focus is on the thinning of the Atmospheric Infrared Sounder (AIRS) profiles. A direct thinning method is first applied to a synthetic data set in order to identify optimal data selection strategies. Experiments on synthetic data suggest that a thinned data set should combine homogeneous samples, and high gradient and variance of gradient samples for optimal performance. This result leads to the modification of our previously developed Density Adjustment Data Thinning algorithm (DADT). The modified DADT (mDADT) algorithm is used to thin the AIRS profiles. Experiments are conducted to compare the thinning performances of mDADT with two simple thinning algorithms. Experiment results show that mDADT algorithm performs better than the two simple thinning algorithms, especially over the regions of significant atmospheric features.
Bradley Zavodsky, Steven Lazarus, Xiang Li 0043, Mike Lueken, Michael Splitt, Rahul Ramachandran, Sunil Movva, Sara J. Graves, William Lapenta
IGARSS (3)6
2006 A Simple and Efficient Feature Extraction Algorithm for Geophysical Phenomena
abstract
A phenomenon is defined as any state or process known through the senses rather than by intuition or reasoning, and thus is an observable event, especially something special or unusual. A geophysical phenomenon in the context of geoscience data can be characterized as a spatial region which is significantly different from the rest of the image; having higher/lower than average background intensity value; and having higher variation in intensity when compared to the remaining data points. This paper will describe two variations of the Phenomena Extraction Algorithm (PEA). The PEA consists of three components: a hierarchical splitting to efficiently decompose geoscience data into smaller regions; a set of statistical tests to determine whether decomposed region meets the definition of a geophysical phenomenon and an optimization algorithm to determine the best thresholds needed by these statistical tests. The two variations of the algorithm were tested on a synthetic dataset in a series of experiments. The results from these experiments will be presented in this paper. The use of PEA in a proof-of-concept effort within Linked Environment for Atmospheric Discovery (LEAD), a large NSF funded Information Technology Research project, will also be described.
Rahul Ramachandran, Xiang Li 0043, Sunil Movva, Sara J. Graves
IGARSS1
2006 Dynamic Filtering and Mining Triggers in Mesoscale Meteorology Forecasting
abstract
Abstract — Mesoscale meteorology forecasting as a data driven application is capable of reacting to events in real-time. We explore a framework for dynamic filtering and mining of data products to generate timely triggers for invoking forecasting applications. In this paper, we present our framework, which couples the Calder stream processing system developed at Indiana University for filter processing and trigger generation, and data mining algorithms developed as part of the ADaM data mining tool kit developed at ITSC, UAH, which detect events for trigger generation.
Nithya N. Vijayakumar, Beth Plale, Rahul Ramachandran, Xiang Li 0043
IGARSS3
2005 Automated detection of frontal systems from numerical model-generated data
abstract
Fronts are significant meteorological phenomena of interest. The extraction of frontal systems from observations and model data can greatly benefit many kinds of research and applications in atmospheric sciences. Due to the huge amount of observational and model data available nowadays, automated extraction of front systems is necessary. This paper presents an automated method to detect frontal systems from numerical model-generated data. In this method, a frontal system is characterized by a vector of features, comprised of parameters derived from the model wind field. K-means clustering is applied to the generated sample set of the feature vectors to partition the feature space and to identify clusters representing the fronts. The probability that a model grid belongs to a front is estimated based on its feature vector. The probability image is generated corresponding to the model grids. A hierarchical thresholding technique is applied to the probability image to identify the frontal systems and a Gaussian Bayes classifier is trained to determine the proper threshold value. This is followed by post processing to filter out false signatures. Experiment results from this method are in good agreement with the ones identified by the domain experts.
Xiang Li 0043, Rahul Ramachandran, Sara J. Graves, Sunil Movva, Bilahari Akkiraju, G. David Emmitt, Steven Greco, Robert Atlas, Joseph Terry, Juan-Carlos Jusem
KDD2
2004 Dayglow removal from FUV auroral images
abstract
Aurora study is a key area in understanding the connection and interaction of the solar-terrestrial system. Auroral events are monitored on the global scale at the Far Ultraviolet (FUV) spectrum by satellite-based sensors. However, the existence of dayglow emission significantly limits scientists is the ability to determine the location and the size of auroral ovals. A dynamic methodology to remove day airglaw emission from the LBHL band UVI images on the Polar satellite is presented in this paper. First, the methodology identifies the geomagnetic latitude bound of the night side auroral oval. Then, the maximum range of geomagnetic latitude bound of the auroral is inferred based on the domain knowledge. Using the non-auroral dayglow pixels, a multi-variable regression fit of the intensity as function of the cosine of the solar zenith angle and cosine of the satellite viewing zenith angle is obtained. The methodology then uses this fitting function to estimate the dayglow intensity and to remove its effect in the LBHL band FUV images. This methodology does not require an external source of input solar flux and the dayglow removal is solely based on individual images. Experiment results show that dayglow effect is removed significantly from the original LBHL FUV images. Using this methodology, the performance of auroral oval detection algorithms is significantly improved
Xiang Li 0043, Rahul Ramachandran, Sunil Movva, Sara J. Graves, Glynn A. Germany, Wladislaw Lyatsky, Arjun Tan
IGARSS2
2004 Agent framework for intelligent data processing
abstract
Data preparation is an important part of the mining process. This work describes MIDAS, an agent framework for intelligent data processing. The objective framework is to provide end users automated data processing services such as subsetting and data format translation by coupling Earth science markup language (ESML) interchange technology and ontologies. These ontology driven agents guide the user through the process of data processing and are able to make decisions for them based on their output requirements. This work describes the design approach used to extend the ESML schema to incorporate semantics using ontologies. It explains the overall architecture including the infrastructure layer, organization and the agent design used in this framework. The subset of performatives derived from the knowledge query and manipulation language (KQML) used by the agents to interact will also be described.
Rahul Ramachandran, Sara J. Graves, Sunil Movva, Xiang Li 0043
IGARSS1
2002 Interchange technology for applications to facilitate generic access to heterogenous data formats
abstract
An Interchange Technology facilitates seamless interactions between applications, tools and services with datasets in heterogeneous formats. The Information Technology and Systems Center at the University of Alabama in Huntsville is currently developing this technology focused on Earth Science data, in general, and particularly on the vast amounts of remotely sensed data. The Earth Science Markup Language (ESML) is one such interchange technology and is comprised of description files and a related library of programming utilities. ESML is based on the eXtensible Markup Language (XML/spl trade/) and allows data "descriptions" to be written in a standard fashion. ESML is unique in that it not only describes the content and structure of the data, but also provides semantic information for the end user application to utilize. Another advantage is that the effort involved to describe legacy data formats in ESML will be small. ESML and the associated library will allow wider interoperability of Earth Science services and tools, enabling Earth Scientists to author, discover, and interpret Earth Science information, and to work more easily with data in a variety of formats and structures. This interchange technology will facilitate the development of dataset-independent searches, visualization, and analysis tools without requiring data to be in any particular format. This paper will describe this interchange technology and provide examples of generic access to heterogeneous remotely sensed data sets.
Rahul Ramachandran, Helen Conover, Sara J. Graves, Ken Keiser
IGARSS1