EDBT 2026 Demo / reviewers in the wild / expert
Cristina Garcia-Cardona
dblp:123/4696
· DBLP profile ↗
4ranked-venue papers in the field
2as first author
2since 2021 · last 2024
0000-0002-5641-3491ORCID · reported
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 4 (2 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | An Active Learning-Based Streaming Pipeline for Reduced Data Training of Structure Finding Models in Neutron DiffractometryabstractStructure determination workloads in neutron diffractometry are computationally expensive and routinely require several hours to many days to determine the structure of a material from its neutron diffraction patterns. The potential for machine learning models trained on simulated neutron scattering patterns to significantly speed up these tasks have been reported recently. However, the amount of simulated data needed to train these models grows exponentially with the number of structural parameters to be predicted and poses a significant computational challenge. To overcome this challenge, we introduce a novel batch-mode active learning (AL) policy that uses uncertainty sampling to simulate training data drawn from a probability distribution that prefers labelled examples about which the model is least certain. We confirm its efficacy in training the same models with ∼ 75% less training data while improving the accuracy. We then discuss the design of an efficient stream-based training workflow that uses this AL policy and present a performance study on two heterogeneous platforms to demonstrate that, compared with a conventional training workflow, the streaming workflow delivers ∼ 20% shorter training time without any loss of accuracy. Tianle Wang 0001, Jorge Ramirez, Cristina Garcia-Cardona, Thomas Proffen, Shantenu Jha, Sudip K. Seal |
IEEE Big Data | 3 |
| 2023 | A Deep Learning Pipeline for Optimizing Large-scale Phase Field SimulationsabstractPhase field (PF) simulations are computationally expensive but remain a key analysis tool to understand the complex mechanisms of additive manufacturing (AM) processes. Each PF simulation-aided analysis requires thousands of node hours on leadership-class supercomputers. One of the main goals of these analyses is the study of microstructure evolution during the build process which begins with the onset of nucleation. Nucleation occurs under certain thermomechanical conditions which are not known a priori and many PF simulations are required to identify ranges of input thermo-mechanical parameters that can result in the onset of nucleation. Since many of the simulations do not result in nucleation, an analysis campaign often ends up wasting tremendous amounts of precious computing resources executing nucleation-absent simulations. The goal of this work is to design and train deep learning models to inform a PF simulation about the likelihood of the occurrence of nucleation in a future simulation time-step based on the state summary over a finite number of past time-steps of a running simulation. If the prediction determines that the running simulation is unlikely to reach nucleation in the allotted time, then its execution is stopped immediately ultimately resulting in vast reduction in wasted computations when accrued over all the PF simulations typically performed in a single or multiple analysis campaign(s). The paper presents the performance of a machine learning pipeline that uses a convolutional neural network (CNN) model to learn an embedding which is then used with a self-attention network to build a multi-task deep learning model to predict the likelihood of nucleation. The model also predicts the input parameters used in a simulation. Performance is compared with a baseline pipeline that uses an off-the-shelf LeNet-5 model to learn the initial embedding. Despite their smaller size, performance results indicate significant improvement in accuracy of the proposed models compared to the larger baseline models. Ramakrishnan Kannan, Cristina Garcia-Cardona, Balasubramaniam Radhakrishnan, Sudip K. Seal |
IEEE Big Data | 2 |
| 2020 | Structure Prediction from Neutron Scattering Profiles: A Data Sciences ApproachabstractOne of the main goals of neutron data analysis is to determine the internal structure of materials from their neutron scattering profiles. These structures are defined by a crystallographic class label and a set of real-valued parameters specific to that class. Existing structure analysis approaches use computationally expensive loop refinements methods that routinely take days, and even weeks, to complete. Additionally, the outcomes often rely on the fidelity of physical models that are computed during the refinement process. Here, we evaluate the feasibffity of using trained data-driven machine learning models as fast and accurate substitutes for these expensive methods. We report on the efficacies of a variety of ML models, including convolutional neural networks, auto-encoders, random forests and combinations thereof, in addition to techniques such as transfer learning in predicting these structural parameters. Specifically, we evaluate two categories of models which we call class-conditional and integrated. The first relies on a two-stage inference pipeline in which a crystallographic class label is first predicted followed by regression to predict the length/angle parameters. In the second category, the classification and regression tasks are performed as a single learning task. We train these models on synthetically generated data, validate them against experimental observa-tions and show that integrated models outperform their class-conditional counterparts opening up the possibffity of deep learning models as a viable alternative to existing resource-intensive loop refinement methods in neutron data analysis. Cristina Garcia-Cardona, Ramakrishnan Kannan, J. Travis Johnston, Thomas Proffen, Sudip K. Seal |
IEEE BigData | 1 |
| 2019 | Learning to Predict Material Structure from Neutron Scattering DataabstractUnderstanding structural properties of materials and how they relate to its atomic structure, while extremely challenging, is a key scientific quest that has dominated the landscape of materials research for decades. Neutron and X-ray scattering is a state-of-the-art method to investigate material structure on the atomic scale. Traditional methods of processing neutron scattering data to decipher the structure of target materials have relied on computing scattering patterns using physics-based forward models and comparing them with experimentally gathered scattering profiles within a computationally expensive optimization loop. Here, we report an initial design of a data-driven machine learning pipeline for material structure prediction that is computationally faster (once trained) and potentially more accurate. We describe the architecture of the ML pipeline and a preliminary benchmarking study of shallow machine learning models in terms of their prediction accuracy and limitations. We show that material structure prediction from neutron scattering data using shallow learning models is feasible to within 90% prediction accuracy for certain classes of materials but deeper models are required for more general material structure predictions. Cristina Garcia-Cardona, Ramakrishnan Kannan, J. Travis Johnston, Thomas Proffen, Katharine Page, Sudip K. Seal |
IEEE BigData | 1 |