Kewei Wang 0002

dblp:253/2810-2 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2024 Enhancing Deep Neural Network Classification Performance Through Novel Weight Initialization: t-SNE Supported Walsh Matrix Approach
abstract
Deep Neural Networks, as a subset of AI, outper-form in understanding complex relationships. The key to this success lies in the network's ability to adapt to problem-specific nuances. During model training, the network dynamically optimizes its weights by updating them during backpropagation while trying to minimize the value of the loss function. Throughout this process, the shaping of model weights is crucially linked to how they were initialized. In this study, we introduce the auxiliary network model, called Sup-Walsh (Support Walsh), which reorganizes weights to enhance class boundaries. We tested our approach on three publicly available datasets using popular classification models. For instance, when using AlexNet [1] on the MNIST dataset [2], integrating Sup-Walsh led to a significant increase in accuracy after first epoch from 14.61% to 78.99%. Similarly, GoogleNet [3] on the FashionMNIST dataset [4] showed a notable 31.61% accuracy difference between configurations without and with Sup-Walsh after first epoch. Across nearly all experiments, our proposed method consistently outperformed existing approaches, demonstrating its potential to improve classification accuracy. Code availability: Code is available at Efficient-Weight-Initializer.
Muhammed Nur Talha Kilic, Vishu Gupta, Yuwei Mao, Kewei Wang 0002, Alec Peltekian, Alok N. Choudhary, Wei-keng Liao, Ankit Agrawal 0001
ICMLA4
2024 Deep Learning Based Inverse Modeling for Materials Design: From Microstructure and Property to Processing
abstract
Polycrystalline materials are crucial in various industries, necessitating a comprehensive understanding of the processing-structure-property-performance (PSPP) relationships. Traditional experimental methods are laborious and slow, while computational approaches predominantly address forward problems, deriving structures and properties from processing conditions. Conversely, inferring processing parameters from desired microstructures and properties remains a crucial yet challenging inverse problem due to the complex and nonlinear mappings involved. In this work, we propose a deep learning-based framework exploring non-sequential and sequential models to address two key inverse problems: predicting processing parameters from microstructures and from properties. Focusing on microstructural texture defined by the orientation distribution function (ODF), we apply our framework to copper, generating a dataset of 31,588 unique processing strain rates ($s$-1) in [0, 1] with corresponding ODFs and homogenized properties through simulations. Our inverse prediction results on processing parameters demonstrate high accuracy, with average test RMSEs of 0.0152 from microstructures and 0.0295 from properties. These findings validate the framework's efficacy as a tool for polycrystalline materials process design, enabling the precise determination of processing methods to achieve desired microstructures and properties.
Kewei Wang 0002, Yuwei Mao, Mahmudul Hasan 0016, Md Maruf Billah, Muhammed Nur Talha Kilic, Vishu Gupta, Wei-keng Liao, Alok N. Choudhary, Pinar Acar, Ankit Agrawal 0001
ICMLA1
2022 Using Multi-Resolution Data to Accelerate Neural Network Training in Scientific Applications
abstract
Neural networks are powerful solutions to many scientific applications; however, they usually require long model training time due to large training data sets or large model size. Research has been focused on developing numerical optimization algorithms and parallel processing to reduce the training time. In this work, we propose a multi-resolution strategy that can reduce the training time by training the model with the reduced-resolution data samples at the beginning and later switching to the original resolution data samples. This strategy is motivated by the observation that coarser versions of many applications can be solved faster than their denser counterparts, and the solution to a coarser problem could be used to initialize the solution to the denser problem. When applying the idea to neural network training, coarse data can have a similar effect on the learning curves at the early stage as the dense data but requires less time. Once the curves no longer improve significantly, our strategy switches to using the data in original resolution. The key in this process is the ability to generate multiple resolutions of a problem automatically, which could usually be done with scientific applications with spatial and temporal continuity. We use two real-world scientific applications, CosmoFlow and DeepCAM, to evaluate the proposed mixed-resolution training strategy. Our experiment results demonstrate that the proposed training strategy effectively reduces the end-to-end training time while achieving a comparable accuracy to that of the training only with the original data. While maintaining the same model accuracy, our multi-resolution training strategy reduces the end-to-end training time up to 30% and 23% for CosmoFlow and DeepCAM, respectively.
Kewei Wang 0002, Sunwoo Lee 0001, Jan Balewski, Alex Sim, Peter Nugent, Ankit Agrawal 0001, Alok N. Choudhary, Kesheng Wu, Wei-keng Liao
CCGRID1
2022 A case study on parallel HDF5 dataset concatenation for high energy physics data analysis
Sunwoo Lee 0001, Kaiyuan Hou, Kewei Wang 0002, Saba Sehrish, Marc F. Paterno, Jim Kowalkowski, Quincey Koziol, Robert B. Ross, Ankit Agrawal 0001, Alok N. Choudhary, Wei-keng Liao
Parallel Comput.3
2021 Asynchronous I/O Strategy for Large-Scale Deep Learning Applications
abstract
Many scientific applications have started using deep learning methods for their classification or regression problems. However, for data-intensive scientific applications, I/O performance can be the major performance bottleneck. In order to effectively solve important real-world problems using deep learning methods on High-Performance Computing (HPC) systems, it is essential to address the poor I/O performance issue in large-scale neural network training. In this paper, we propose an asynchronous I/O strategy that can be generally applied to deep learning applications. Our I/O strategy employs an I/O -dedicated thread per process, that performs I/O operations independently of the training progress. The I/O thread reads many training samples at once to reduce the total number of I/O operations per epoch. Given the fixed amount of training data, the fewer the I/O operations per epoch, the shorter the overall I/O time. The I/O operations are also overlapped with the computations using the double-buffering method. We evaluate our I/O strategy using two real-world scientific applications, CosmoFlow and Neuron-Inverter. Our experimental results demonstrate that the proposed I/O strategy significantly improves the scaling performance without affecting the regression performance.
Sunwoo Lee 0001, Qiao Kang, Kewei Wang 0002, Jan Balewski, Alex Sim, Ankit Agrawal 0001, Alok N. Choudhary, Peter Nugent, Kesheng Wu, Wei-keng Liao
HiPC3