EDBT 2026 Demo / reviewers in the wild / expert
Giri Prakash
dblp:194/7419
· DBLP profile ↗
7ranked-venue papers in the field
1as first author
2since 2021 · last 2024
0000-0002-2590-5848ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 7 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Data Workbench For Earth and Atmospheric Science Research CommunityabstractWe introduce the Atmospheric Radiation Measurement (ARM) Data Workbench, a web-based interface that combines access to atmospheric datasets and computational resources. It is developed to facilitate a interactive data exploration and data analysis for the research community. The ARM Data Workbench offers a comprehensive solution that combines data locality, user-friendly access, and collaborative capabilities within a single platform. To empower researchers with a familiar analysis and compute environment, the data workbench integrates a JupyterHub which is a popular platform for scientific computing. In this paper, we discuss the features of ARM Data Workbench through which users can access the ARM hosted scientific datasets, order them and get them delivered. We discuss different types of user accounts we provide in order to have access to ARM data and the compute platform to perform analysis. We explain the process of navigating to ARM’s Data Discovery interface to access datasets and JupyterHub platform for scientific users with an example case study. The platform targets new researchers, graduate students and established researchers to help them in their research in a collaborative environment. Sujata Goswami, Kyle Dumas, Wade Darnell, Varsile Tudor Garbulet, Michael Giansiracusa, Giri Prakash |
IEEE Big Data | 6 |
| 2024 | Enhancing Discoverability and Management of Atmospheric Data at Scale: Solutions from the ARM Data CenterabstractThe Atmospheric Radiation Measurement (ARM) is a multi-laboratory and multi-institutional U.S. Department of Energy (DOE) Office of Science National User Facility. The ARM Data Center (ADC), located at Oak Ridge National Laboratory, collects, archives, and shares vast atmospheric data crucial for climate research. The ADC manages over 7 PB of data from 460 instruments worldwide, processing it into more than 11,000 diverse data products using the Network Common Data Form (NetCDF) for machine-independent accessibility. The primary challenge addressed in this paper is the efficient management and distribution of vast and diverse datasets essential for the climate research community, enhancing accessibility through advanced tools like Data Discovery. The ADC has developed advanced infrastructure and software architecture to handle the continuous influx of heterogeneous data to enhance data discoverability, resulting in increased scientific collaboration. In 2023, users from over 34 countries downloaded and utilized ARM data, resulting in 1,455 publications. The ADC’s efforts have significantly improved the discoverability and usability of atmospheric data, fostering extensive scientific research and collaboration. This paper details the solutions implemented by the ADC team for efficient data discovery and distribution, and it demonstrates ARM’s capability of staging processed data for scientific analysis. Chirag Shah 0002, Wade Darnell, Hannah Collier, Harold Shanafield, Michael Giansiracusa, Giri Prakash |
IEEE Big Data | 6 |
| 2020 | Clustering-Based Predictive Analytics to Improve Scientific Data DiscoveryabstractGiven the sheer volume of scientific data archived within the data-intensive projects at the US Department of Energy's Oak Ridge National Laboratory, finding precisely what data we are looking for may not be a trivial task; conversely, we may also miss a more prominent data product. To address such issues, we propose improving the data discovery system and using data analytics methods to comprehend what specific users might be interested in based on their physiological state, search patterns, and past data usage history. This work's primary goal is to prune the complexity, increase the visibility of popular data products, and direct users toward the data that best meet their needs. The proposed algorithm constructs a user profile based on the user's explicit or implicit interactions with the system, such as items they are currently looking at on-site and the key metadata mappings related to the data set. The pattern is then used to build a training data set, which will help find relevant data to recommend to the user. Ranjeet Devarakonda, Jitendra Kumar 0001, Giri Prakash |
IEEE BigData | 3 |
| 2020 | Automated Indexing of Structured Scientific Metadata Using Apache SolrabstractScientific datasets are continuously growing with the amount of raw data being collected worldwide. This amount of data poses the biggest challenge to web search engines on how to retrieve them efficiently. This paper discusses how major scientific data centers are using popular open-source search platforms such as Solr [1] to retrieve structured data stored in data sources such as relational database management systems using its import handler mechanisms [2]. Additionally, we will also focus on how we can configure Solr to serve advanced full-text, faceted search capabilities, along with its key features, which simplify representing and delivering better performance to the scientific search interfaces. Kavya Guntupally, Kyle Dumas, Wade Darnell, Michael C. Crow, Ranjeet Devarakonda, Giri Prakash |
IEEE BigData | 6 |
| 2019 | Big Federal Data Centers Implementing FAIR Data Principles: ARM Data Center ExampleabstractAtmospheric Radiation Measurement (ARM) is a multi-laboratory/multi-institutional, US Department of Energy Office of Science National User Facility. ARM's data is currently hosted at the ARM Data Center (ADC) in Oak Ridge, Tennessee. The ADC holds more than 12,000 data products, with a total holding of more than 1.8 PB of data that dates back to 1992. This includes data from instruments, value-added products, model outputs, field campaigns, and principle investigator contributed data. In this paper, we discuss how big federal scientific data centers, such as ARM, that use modern and scalable architecture apply findable, accessible, interoperable, and reusable (FAIR) data principles to improve overall efficiency. These principles mainly emphasize machine-to-machine interactions that are directly applicable to ARM because of its data volume. Ranjeet Devarakonda, Giri Prakash, Kavya Guntupally, Jitendra Kumar 0001 |
IEEE BigData | 2 |
| 2016 | Next-gen tools for big scientific data: ARM data center exampleabstractThe Atmospheric Radiation Measurement (ARM) Climate Research Facility (www.arm.gov) provides atmospheric observations from diverse climatic regimes around the world. Currently, ARM archives over 22 million user assessable data files, primarily stored in NetCDF file format, with total data volumes close to one Petabyte. In this paper, we will discuss how ARM is currently storing, distributing, cataloging and visualizing such large volumes of multi-dimensional climate observations and model data and also describe their future plan. Ranjeet Devarakonda, Kyle Dumas, Sheman Beus, Everett Neil Rush, Bhargavi Krishna, Robert Records, Giri Prakash |
IEEE BigData | 7 |
| 2016 | HPC infrastructure to support the next-generation ARM facility data operationsabstractThe Department of Energy's (DOE) Atmospheric Radiation Measurement (ARM) Climate Research Facility is establishing an adaptive data services and operations architecture in support of the Next-Generation ARM Facility as explained in its Decadal Vision. In this paper, we describe the capabilities of the ARM Data Center (ADC) and the upcoming high-performance computing infrastructure in support of this Next-Generation ARM Facility. Giri Prakash, Jitendra Kumar 0001, Everett Neil Rush, Robert Records, Anthony Clodfelter, Jimmy W. Voyles |
IEEE BigData | 1 |