Addi Malviya-Thakur

dblp:318/2909 · also Addi Thakur Malviya · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-2681-9992ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Integrating scientific single-page applications with DevSecOps
Lance Drane, Marshall T. McDonnell, Randall Petras, Cody Stiner, Arthur J. Ruckman, Gavin M. Wiggins, Gregory Cage, Seth Hitefield, Jesse McGaha, Andrew Ayres, Michael J. Brim, Rick Archibald, Addi Malviya-Thakur
Future Gener. Comput. Syst.14
2024 Privacy Preserving Federated Learning for Advanced Scientific Ecosystems
abstract
We present a framework to provide privacy preserving (PP) federating learning (FL) across multiple computational and experimental facilities. This work joins the compute capabilities of National Energy Research Scientific Computing Center (NERSC) and Oak Ridge National Laboratory Research Cloud (ORC) with simulated experimental data, such as those produced at the SLAC National Accelerator Laboratory and Spallation Neutron Source (SNS). We describe the software infrastructure developed to provide privacy for computational and experimental networks. We developed algorithmic privacy across the federated system by embedding database security, computation, and communication into the federation architecture, utilizing scientific tools developed by the experimental community.
Rick Archibald, Addi Malviya-Thakur, Marshall T. McDonnell, Gregory Cage, Cody Stiner, Lance Drane, M. Paul Laiu, Michael J. Brim, Mathieu Doucet, William T. Heller, Ryan Coffee
IEEE Big Data2
2023 How R Developers explain their Package Choice: A Survey
abstract
Background: Contemporary software development relies heavily on reusing already implemented functionality, usually in the form of packages. Aims: We aim to shed light on developers' preferences when selecting packages in R language. Method: To do that, we create and administer a survey to over 1000 developers who have added one of two common dataframe enhancement libraries in R to their projects: data.table or tidyr. We design a questionnaire using the Social Contagion Theory (SCT) following prior work on technology adoption and ensure that key dimensions affecting developer choice are considered. Results: Of the 1085 developers we contacted, 803 completed the survey asking them to prioritize various factors known to affect developer perceptions of package quality and to provide their background. Most developers self-identified as data scientists with two to five years of work experience. We found significant differences between the preferences of developers who chose data.table and tidyr. Surprisingly, package reputation based on easy-to-see measures, such as the number of stars on GitHub, was not an important factor for either group. Conclusions: Our findings demonstrate the inherently social nature of package adoption. They can help design future studies on how different populations of developers make decisions on which software packages to use in their projects. Finally, package developers and maintainers can benefit by better understanding the prime concerns of the users of their packages.
Addi Malviya-Thakur, Audris Mockus, Russell Zaretzki, Bogdan C. Bichescu, Randy V. Bradley
ESEM1
2022 Real-World Experiences Adopting Workflows at Exascale on the ExaAM Project
abstract
The purpose of this study is to discuss the experiential lessons associated with adopting scientific workflows in the Exascale Additive Manufacturing project (ExaAM) through the lens of Perceived Characteristic of Innovation (PCI). Besides the implementation, the factors we considered critical to the adoption of the workflow are provenance, sustainable automation, implementation challenges, and integration/compatibility challenges. Through conversations and interviews among the program managers, project leads, and software engineers, we have developed critical insight and strategies to overcome the obstacles and augment the successful adoption and long-term use of these workflows in ExaAM and beyond. We hope our work will pave the way for others in the research community to develop and use workflows in their respective science domains.
Addi Malviya-Thakur, Reed Milewicz, Samuel Grayson, Philip W. Fackler, James F. Belak, John A. Turner
e-Science1
2021 A new methodological framework for hazard detection models in health information technology systems
abstract
The adoption of health information technology (HIT) has facilitated efforts to increase the quality and efficiency of health care services and decrease health care overhead while simultaneously generating massive amounts of digital information stored in electronic health records (EHRs). However, due to patient safety issues resulting from the use of HIT systems, there is an emerging need to develop and implement hazard detection tools to identify and mitigate risks to patients. This paper presents a new methodological framework to develop hazard detection models and to demonstrate its capability by using the US Department of Veterans Affairs' (VA) Corporate Data Warehouse, the data repository for the VA's EHR. The overall purpose of the framework is to provide structure for research and communication about research results. One objective is to decrease the communication barriers between interdisciplinary research stakeholders and to provide structure for detecting hazards and risks to patient safety introduced by HIT systems through errors in the collection, transmission, use, and processing of data in the EHR, as well as potential programming or configuration errors in these HIT systems. A nine-stage framework was created, which comprises programs about feature extraction, detector development, and detector optimization, as well as a support environment for evaluating detector models. The framework forms the foundation for developing hazard detection tools and the foundation for adapting methods to particular HIT systems.
Olufemi A. Omitaomu, Hilda B. Klasky, Mohammed M. Olama, Özgür Özmen, Laura L. Pullum, Addi Malviya-Thakur, P. Teja Kuruganti, Jean M. Scott, Angela Laurio, Frank Drews, Brian C. Sauer, Merry Ward, Jonathan R. Nebeker
J. Biomed. Informatics6
2020 Adaptive Anomaly Detection for Dynamic Clinical Event Sequences
abstract
Over the past decade, health information technology (IT) has enabled the amount of digital information stored in electronic health records (EHRs) to expand greatly. However, according to some studies, hazards in health IT can lead to changes in clinical decisions, care processes, and care outcomes, as well as other issues. Thus, the effects of health IT hazards on patient safety have been at the forefront of recent patient safety research. Nonetheless, hazard detection in health IT remains a challenge. In this paper, the authors assume that safety-related issues in health IT would exhibit anomalous characteristics in EHR data. Although all hazards will exhibit some anomalous characteristics, not all anomalies can be regarded as hazards. The authors hypothesize that errors in health IT could lead to interruptions in the sequence of clinical actions. To this end, the problem of detecting anomalous sequences in big EHR data is considered. This paper focuses on dynamic event sequences, which are a series of clinical actions in motion. The authors propose an adaptive anomaly detection approach that uses higher-order network representation to detect anomalous sequences. Furthermore, the authors propose a contiguous subsequence anomaly detection approach that identifies abnormal subsequences in the detected anomalous sequences. The proposed approaches are tested by using synthetic and real-world EHR data. The proposed methods outperform existing state of the art anomaly detection techniques. To reduce the computational complexity associated with the operational implementation of the proposed approaches, the Apache Spark environment was leveraged, and a much shorter run time together with improved performance were achieved, especially for data with more than 60,000 sequences.
Olufemi A. Omitaomu, Qing Cao 0001, Mohammed M. Olama, Özgür Özmen, Hilda B. Klasky, Laura L. Pullum, Addi Malviya-Thakur, P. Teja Kuruganti, Jean M. Scott, Angela Laurio, Frank Drews, Brian C. Sauer, Merry Ward, Jonathan R. Nebeker
IEEE BigData8