Melissa C. Smith

dblp:89/3990 · DBLP profile ↗
← Back
24ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0003-0798-8536ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 8 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 30% High-performance computing · 26% Electronic design automation · 23%
Artificial intelligence
2 papers
Speech recognition and synthesis · 50% Reinforcement learning · 25% Graph learning · 25%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis
speech enhancement
0.512021
Speech Enhancement Using Multi-Stage Self-Attentive Temporal Convolutional Networks · IEEE ACM Trans. Audio Speech Lang. Process. 2021
High-performance computing › numerical linear algebra › linear solver
iterative linear solvers
0.512021
A Parallel Jacobi-Embedded Gauss-Seidel Method · IEEE Trans. Parallel Distributed Syst. 2021
Parallel and multicore computing › parallel algorithms › parallel matrix algorithms
parallel iterative solvers
0.512021
A Parallel Jacobi-Embedded Gauss-Seidel Method · IEEE Trans. Parallel Distributed Syst. 2021
Machine learning › Graph learning › graph neural network
graph convolution
0.312018
Generalized Value Iteration Networks: Life Beyond Lattices · AAAI 2018
Machine learning › Reinforcement learning › value-based reinforcement learning
value iteration network
0.312018
Generalized Value Iteration Networks: Life Beyond Lattices · AAAI 2018
Electronic design automation
design flow
0.212015
Enhancing Hardware Design Flows with MyHDL · FPGA 2015
Electronic design automation
hardware description language
0.212015
Enhancing Hardware Design Flows with MyHDL · FPGA 2015
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.112021
Speech Enhancement Using Multi-Stage Self-Attentive Temporal Convolutional Networks · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Processor architecture and microarchitecture
many-core architecture
0.112021
A Parallel Jacobi-Embedded Gauss-Seidel Method · IEEE Trans. Parallel Distributed Syst. 2021
Performance modeling and evaluation
simulation
0.112015
Enhancing Hardware Design Flows with MyHDL · FPGA 2015
Embedded and real-time systems › real-time scheduling
heuristic scheduling
0.112006
Simulative analysis of dynamic scheduling heuristics for reconfigurable computing of parallel applications · FPGA 2006
Cloud and datacenter computing
job scheduling
0.112006
Simulative analysis of dynamic scheduling heuristics for reconfigurable computing of parallel applications · FPGA 2006
Reconfigurable computing and FPGAs
reconfigurable computing
0.112006
Simulative analysis of dynamic scheduling heuristics for reconfigurable computing of parallel applications · FPGA 2006
Parallel and multicore computing › parallel scheduling
runtime scheduling
0.112006
Simulative analysis of dynamic scheduling heuristics for reconfigurable computing of parallel applications · FPGA 2006

Methods — techniques the papers use, named apart from their topics

temporal convolutional network · 0.5self-attention · 0.5multi-stage learning · 0.5matrix reordering · 0.5jacobi iteration · 0.5domain decomposition · 0.5value iteration · 0.3q-learning · 0.3graph convolution · 0.3verilog · 0.2python · 0.2VHDL · 0.2simulation · 0.1performance prediction model · 0.1
YearPublicationVenuePosition
2025 GEMDiff: a diffusion workflow bridges between normal and tumor gene expression states: a breast cancer case study
abstract
Breast cancer remains a significant global health challenge due to its complexity, which arises from multiple genetic and epigenetic mutations that originate in normal breast tissue. Traditional machine learning models often fall short in addressing the intricate gene interactions that complicate drug design and treatment strategies. In contrast, our study introduces GEMDiff, a novel computational workflow leveraging a diffusion model to bridge the gene expression states between normal and tumor conditions. GEMDiff augments RNAseq data and simulates perturbation transformations between normal and tumor gene states, enhancing biomarker identification. GEMDiff can handle large-scale gene expression data without succumbing to the scalability and stability issues that plague other generative models. By avoiding the need for task-specific hyper-parameter tuning and specific loss functions, GEMDiff can be generalized across various tasks, making it a robust tool for gene expression analysis. The model's ability to augment RNA-seq data and simulate gene perturbations provides a valuable tool for researchers. This capability can be used to generate synthetic data for training other machine learning models, thereby addressing the issue of limited biological data and enhancing the performance of predictive models. The effectiveness of GEMDiff is demonstrated through a case study using breast mRNA gene expression data, identifying 307 core genes involved in the transition from a breast tumor to a normal gene expression state. GEMDiff is open source and available at https://github.com/xai990/GEMDiff.git under the MIT license.
Xusheng Ai, Melissa C. Smith, F. Alex Feltus
Briefings Bioinform.2
2022 Lossy Compression to Reduce Latency of Local Image Transfer for Autonomous Off-Road Perception Systems
abstract
Autonomous vehicles greatly rely on their perception system for navigation. Semantic segmentation provides a much better understanding of a vehicle’s surroundings than object detection. Unfortunately, complete image segmentation comes at a higher computational cost than object detection, which complicates developing a real-time perception system using semantic segmentation. Perception systems contain other bottlenecks too, and are not only limited by their deep learning model. An inherent amount of latency exists in data transfer, specifically through Ethernet. A vehicle’s camera feed must be transferred to an edge device for image processing as part of the autonomous driving decision-making process. This study investigates decreasing image transfer time by using various levels of JPEG compression as well as further understanding how compression affects the accuracy of semantic segmentation. Additionally, as most autonomous driving research focuses on urban environments, we look to explore autonomous unmanned ground vehicles (UGVs) in the off-road space by using the Rellis-3D dataset. We train and evaluate SwiftNet, a state-of-the-art semantic segmentation model, at different JPEG compression ratios and identify the accuracy. The transfer time of these different compression ratios is tested on three images. Results show a continual decrease in accuracy occurs as the compression ratios increase. When training SwiftNet on the train set with no compression, the highest compression ratio of 16.96 achieves a mean intersection over union (mIoU) score of 67.9% compared to the baseline achieving 78.9% mIoU. There is an increase in the accuracy of the higher compression ratios by training SwiftNet on the corresponding compression ratios; the highest compression ratio reaches 74.9% mIoU. Lastly, we notice a positive transfer speedup of these higher compression ratios when inducing JPEG compression in all transfer scenarios: (a) 1870 images (b) 10 images, (c) 1 image. Each scenario has a speedup of 1.18×, 1.14×, and 1.06×, respectively.
Max H. Faykus, Bradley Selee, Jon Calhoun 0001, Melissa C. Smith
IEEE Big Data4
2022 Addressing noise in co-expression network construction
abstract
Gene co-expression networks (GCNs) provide multiple benefits to molecular research including hypothesis generation and biomarker discovery. Transcriptome profiles serve as input for GCN construction and are derived from increasingly larger studies with samples across multiple experimental conditions, treatments, time points, genotypes, etc. Such experiments with larger numbers of variables confound discovery of true network edges, exclude edges and inhibit discovery of context (or condition) specific network edges. To demonstrate this problem, a 475-sample dataset is used to show that up to 97% of GCN edges can be misleading because correlations are false or incorrect. False and incorrect correlations can occur when tests are applied without ensuring assumptions are met, and pairwise gene expression may not meet test assumptions if the expression of at least one gene in the pairwise comparison is a function of multiple confounding variables. The 'one-size-fits-all' approach to GCN construction is therefore problematic for large, multivariable datasets. Recently, the Knowledge Independent Network Construction toolkit has been used in multiple studies to provide a dynamic approach to GCN construction that ensures statistical tests meet assumptions and confounding variables are addressed. Additionally, it can associate experimental context for each edge of the network resulting in context-specific GCNs (csGCNs). To help researchers recognize such challenges in GCN construction, and the creation of csGCNs, we provide a review of the workflow.
Josh J. R. Burns, Benjamin T. Shealy, Mitchell S. Greer, John A. Hadish, Matthew McGowan, Tyler Biggs, Melissa C. Smith, F. Alex Feltus, Stephen P. Ficklin
Briefings Bioinform.7
2022 GEMmaker: process massive RNA-seq datasets on heterogeneous computational infrastructure
abstract
BACKGROUND: Quantification of gene expression from RNA-seq data is a prerequisite for transcriptome analysis such as differential gene expression analysis and gene co-expression network construction. Individual RNA-seq experiments are larger and combining multiple experiments from sequence repositories can result in datasets with thousands of samples. Processing hundreds to thousands of RNA-seq data can result in challenges related to data management, access to sufficient computational resources, navigation of high-performance computing (HPC) systems, installation of required software dependencies, and reproducibility. Processing of larger and deeper RNA-seq experiments will become more common as sequencing technology matures. RESULTS: GEMmaker, is a nf-core compliant, Nextflow workflow, that quantifies gene expression from small to massive RNA-seq datasets. GEMmaker ensures results are highly reproducible through the use of versioned containerized software that can be executed on a single workstation, institutional compute cluster, Kubernetes platform or the cloud. GEMmaker supports popular alignment and quantification tools providing results in raw and normalized formats. GEMmaker is unique in that it can scale to process thousands of local or remote stored samples without exceeding available data storage. CONCLUSIONS: Workflows that quantify gene expression are not new, and many already address issues of portability, reusability, and scale in terms of access to CPUs. GEMmaker provides these benefits and adds the ability to scale despite low data storage infrastructure. This allows users to process hundreds to thousands of RNA-seq samples even when data storage resources are limited. GEMmaker is freely available and fully documented with step-by-step setup and execution instructions.
John A. Hadish, Tyler Biggs, Benjamin T. Shealy, M. Reed Bender, Coleman McKnight, Connor Wytko, Melissa C. Smith, F. Alex Feltus, Loren A. Honaas, Stephen P. Ficklin
BMC Bioinform.7
2021 TIGRA: A Tightly Integrated Generic RISC-V Accelerator Interface
abstract
Field programmable gate array (FPGA) usage in HPC applications is growing with the need for energy efficient and application specific accelerators. Currently, FPGAs are used to accelerate algorithms using OpenCL with communication over PCIe (loosely coupled accelerators) or by modifying existing architectures to incorporate custom logic directly with a CPU (tightly coupled accelerators). However, only the loosely coupled paradigm is feasible to support a variety of acceleration. In this work, we introduce TIGRA, a zero latency interface designed to provide the benefit of tightly coupled accelerators without the developer burden of modifying the underlying architecture, which can enable their usage in HPC. TIGRA is demonstrated on the PicoRV32 processor with AES-128 bit encryption, posit arithmetic, and multiplication.
Brad Green, Dillon Todd, Jon Calhoun 0001, Melissa C. Smith
CLUSTER4
2021 Systematic Evaluation and Enhancement of Speech Recognition in Operational Medical Environments
abstract
Operational medical environments require reliable hands-free solutions to extract data from audio captured under noisy scenarios during rescue missions and provide timely information. However, approaches using automatic speech recognition (ASR) and natural language processing (NLP) techniques are complex as these conversations have a wide range of noise, involve medical terms from multiple speakers, and occur in high-stress environments, among others. These are further complicated by the lack of large training datasets for operational medical scenarios. To address these issues, we developed a platform that enables resilient hands-free data collection, preserves complete documentation through stages of care, and presents the information in near real-time, critical for the medical operation. Our work uniquely focused on systematic evaluation and improvement of a deep neural network-based ASR system by leveraging realistic testing data obtained from medical simulations of battlefield scenarios, which to our knowledge have not been addressed in any prior work. The system performance is shown to improve significantly using multi-style training, language model adaptation for the medical domain, speech enhancement, and NLP techniques.
Snigdhaswin Kar, Prabodh Mishra, Ju Lin, Minjae Woo, Nicholas Deas, Caleb Linduff, Sufeng Niu, Jerome McClendon, D. Hudson Smith, Melissa C. Smith, Ronald W. Gimbel, Kuang-Ching Wang
IJCNN11
2021 Speech Enhancement Using Multi-Stage Self-Attentive Temporal Convolutional Networks
abstract
Multi-stage learning is an effective technique to invoke multiple deep-learning modules sequentially. This paper applies multi-stage learning to speech enhancement by using a multi-stage structure, where each stage comprises a self-attention (SA) block followed by stacks of temporal convolutional network (TCN) blocks with doubling dilation factors. Each stage generates a prediction that is refined in a subsequent stage. A fusion block is inserted at the input of later stages to re-inject original information. The resulting multi-stage speech enhancement system, in short, multi-stage SA-TCN, is compared with state-of-the-art deep-learning speech enhancement methods using the LibriSpeech and VCTK data sets. The multi-stage SA-TCN system's hyper-parameters are fine-tuned, and the impact of the SA block, the fusion block and the number of stages are determined. The use of a multi-stage SA-TCN system as a front-end for automatic speech recognition systems is investigated as well. It is shown that the multi-stage SA-TCN systems perform well relative to other state-of-the-art systems in terms of speech enhancement and speech recognition scores.
Ju Lin, Adriaan J. de Lind van Wijngaarden, Kuang-Ching Wang, Melissa C. Smith
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 A Parallel Jacobi-Embedded Gauss-Seidel Method
abstract
A broad range of scientific simulations involve solving large-scale computationally expensive linear systems of equations. Iterative solvers are typically preferred over direct methods when it comes to large systems due to their lower memory requirements and shorter execution times. However, selecting the appropriate iterative solver is problem-specific and dependent on the type and symmetry of the coefficient matrix. Gauss-Seidel (GS) is an iterative method for solving linear systems that are either strictly diagonally dominant or symmetric positive definite. This technique is an improved version of Jacobi and typically converges in fewer iterations. However, the sequential nature of this algorithm complicates the parallel extraction. In fact, most parallel derivatives of GS rely on the sparsity pattern of the coefficient matrix and require matrix reordering or domain decomposition. In this article, we introduce a new algorithm that exploits the convergence property of GS and adapts the parallel structure of Jacobi. The proposed method works for both dense and sparse systems and is straightforward to implement. We have examined the performance of our method on multicore and many-core architectures. Experimental results demonstrate the superior performance of the proposed algorithm compared with GS and Jacobi. Additionally, performance comparison with built-in Krylov solvers in MATLAB showed that in terms of time per iteration, Krylov methods perform faster on CPUs, but our approach is significantly better when executed on GPUs. Lastly, we apply our method to solve the power flow problem, and the results indicate a significant improvement in runtime, reaching up to 87 times faster speed compared with GS.
Afshin Ahmadi, Felice Manganiello, Amin Khademi, Melissa C. Smith
IEEE Trans. Parallel Distributed Syst.4
2020 Improved Speech Enhancement Using a Time-Domain GAN with Mask Learning
abstract
Speech enhancement is an essential component in robust automatic speech recognition (ASR) systems. Most speech enhancement methods are nowadays based on neural networks that use feature-mapping or mask-learning. This paper proposes a novel speech enhancement method that integrates time-domain feature mapping and mask learning into a unified framework using a Generative Adversarial Network (GAN). The proposed framework processes the received waveform and decouples speech and noise signals, which are fed into two short-time Fourier transform (STFT) convolution 1-D layers that map the waveforms to spectrograms in the complex domain. These speech and noise spectrograms are then used to compute the speech mask loss. The proposed method is evaluated using the TIMIT data set for seen and unseen signal-to-noise ratio conditions. It is shown that the proposed method outperforms the speech enhancement methods that use Deep Neural Network (DNN) based speech enhancement or a Speech Enhancement Generative Adversarial Network (SEGAN).
Ju Lin, Sufeng Niu, Adriaan J. de Lind van Wijngaarden, Jerome McClendon, Melissa C. Smith, Kuang-Ching Wang
INTERSPEECH5
2019 Speech Enhancement Using Forked Generative Adversarial Networks with Spectral Subtraction
Ju Lin, Sufeng Niu, Zice Wei, Adriaan J. de Lind van Wijngaarden, Melissa C. Smith, Kuang-Ching Wang
INTERSPEECH6
2018 Generalized Value Iteration Networks: Life Beyond Lattices
abstract
In this paper, we introduce a generalized value iteration network (GVIN), which is an end-to-end neural network planning module. GVIN emulates the value iteration algorithm by using a novel graph convolution operator, which enables GVIN to learn and plan on irregular spatial graphs. We propose three novel differentiable kernels as graph convolution operators and show that the embedding-based kernel achieves the best performance. Furthermore, we present episodic Q-learning, an improvement upon traditional n-step Q-learning that stabilizes training for VIN and GVIN. Lastly, we evaluate GVIN on planning problems in 2D mazes, irregular graphs, and real-world street networks, showing that GVIN generalizes well for both arbitrary graphs and unseen graphs of larger scaleand outperforms a naive generalization of VIN (discretizing a spatial graph into a 2D image).
Sufeng Niu, Siheng Chen, Hanyu Guo, Colin Targonski, Melissa C. Smith, Jelena Kovacevic
AAAI5
2018 Artificial Intelligence and Deep Learning Applications for Automotive Manufacturing
abstract
Artificial Intelligence (AI) and Deep Learning has been steadily gaining importance due it’s potential for a broad set of science and industry applications. The success of deep learning techniques has found many applications, e.g. in the domain of computer vision and natural language understanding. Developing AI applications is a complex task with many challenges related to data collection, model training, and deployment.In this paper, we evaluate architectures, models and deployment issues related to the usage of deep learning techniques in the automotive manufacturing domain. Particularly, we focus on different computer vision problems in automotive manufacturing processes, e.g., in logistics processes. We developed several deep learning models that help to improve the quality and efficiency of these processes. Finally, we provide an analysis of the architecture, datasets and models used, and provide performance metrics for each of the different models.
André Luckow, Ken Kennedy, Marcin Ziolkowski, Emil Djerekarov, Matthew Cook 0004, Edward B. Duffy, Michael Schleiss, Bennie Vorster, Edwin Weill, Ankit Kulshrestha, Melissa C. Smith
IEEE BigData11
2015 Enhancing Hardware Design Flows with MyHDL
abstract
MyHDL is a Python based HDL that harnesses the power and versatility of Python for hardware development. MyHDL has excellent simulation capabilities and also allows for conversion to Verilog and VHDL, so developers can enter a conventional design flow as desired. Verilog and VHDL are used extensively, particularly because most synthesis tools only support these two languages. However, they are simply outdated; poor parameterization limits high level design and modern abstraction features such as classes are missing.
Keerthan Jaic, Melissa C. Smith
FPGA2
2015 Subjective versus objective: classifying analytical models for productive heterogeneous performance prediction
Vivek K. Pallipuram, Melissa C. Smith, Nilim Sarma, Ranajeet Anand, Edwin Weill, Karan Sapra
J. Supercomput.2
2014 A practical network intrusion detection system for inline FPGAs on 10GbE network adapters
abstract
A network intrusion detection system (NIDS), such as SNORT, analyzes incoming packets to identify potential security threats. Pattern matching is arguably the most important and most computationally intensive component of a NIDS. Software-based NIDS implementations drop up to 90% of packets during increased network load even at lower network bandwidth. We propose an alternative hybrid-NIDS that couples an FPGA with a network adapter to provide hardware support for pattern matching and software support for post processing. The proposed system, SFAOENIDS, offers an extensible open-source NIDS for Solarflare AOE devices. The pattern matching engine-the primary component of the hardware architecture was designed based on the requirements of typical NIDS implementations. In testing on a real network environment, the SFAOENIDS hardware implementation, operating at 200 MHz, handles a 10Gbps data rate without dropping packets while simultaneously minimizing the server CPU load.
Keerthan Jaic, Melissa C. Smith, Nilim Sarma
ASAP2
2014 Combining Hadoop and GPU to preprocess large Affymetrix microarray data
abstract
High density oligonucleotide array (microarray) from Affymetrix has been widely used for the measurements of gene expressions. Currently, public data repositories, such as Gene Expression Omnibus (GEO) of the National Center for Biotechnology Information (NCBI), have accumulated large amounts of microarray data. Efficient integrative analysis of those microarray data will provide significant knowledge about biological systems. None of the existing microarray preprocessing and quality assessment tools can handle very large microarray datasets with tens of thousands of experiments. The preprocessing and quality assessment of microarray datasets contain both data-intensive and compute-intensive tasks. In this paper, we develop a new set of tools using a mix of the Hadoop (for data intensive tasks) and the General-Purpose Graphics Processing Units (GPGPUs) (for compute intensive tasks) to efficiently process large microarray data. Evaluation of our new tools on large microarray datasets with ten thousands of experiments showed promising superior performance. We demonstrate that the combination of Hadoop and GPGPU computation is effective for complex scientific applications that contain both data-intensive and compute-intensive tasks. Our new tool set will make it possible to utilize valuable large microarray data in the public repositories.
Sufeng Niu, Nilim Sarma, Pengfei Xuan, Melissa C. Smith, Pradip K. Srimani, Feng Luo 0001
IEEE BigData5
2014 A regression-based performance prediction framework for synchronous iterative algorithms on general purpose graphical processing unit clusters
abstract
SUMMARY Heterogeneous performance prediction models are valuable tools to accurately predict application runtime, allowing for efficient design space exploration and application mapping. The existing performance models require intricate system architecture knowledge, making the modeling task difficult. In this research, we propose a regression‐based performance prediction framework for general purpose graphical processing unit (GPGPU) clusters that statistically abstracts the system architecture characteristics, enabling performance prediction without detailed system architecture knowledge. The regression‐based framework targets deterministic synchronous iterative algorithms using our synchronous iterative GPGPU execution model and is broken into two components: the computation component that models the GPGPU device and host computations and the communication component that models the network‐level communications. The computation component regression models use algorithm characteristics such as the number of floating‐point operations and total bytes as predictor variables and are trained using several small, instrumented executions of synchronous iterative algorithms that include a range of floating‐point operations‐to‐byte requirements. The regression models for network‐level communications are developed using micro‐benchmarks and employ data transfer size and processor count as predictor variables. Our performance prediction framework achieves prediction accuracy over 90% compared with the actual implementations for several tested GPGPU cluster configurations. The end goal of this research is to offer the scientific computing community, an accurate and easy‐to‐use performance prediction framework that empowers users to optimally utilize the heterogeneous resources. Copyright © 2013 John Wiley & Sons, Ltd.
Vivek K. Pallipuram, Melissa C. Smith, Nimisha Raut
Concurr. Comput. Pract. Exp.2
2012 Optimization of Shared High-Performance Reconfigurable Computing Resources
abstract
In the field of high-performance computing, systems harboring reconfigurable devices, such as field-programmable gate arrays (FPGAs), are gaining more widespread interest. Such systems range from supercomputers with tightly coupled reconfigurable hardware to clusters with reconfigurable devices at each node. The use of these architectures for scientific computing provides an alternative for computationally demanding problems and has advantages in metrics, such as operating cost/performance and power/performance. However, performance optimization of these systems can be challenging even with knowledge of the system’s characteristics. Our analytic performance model includes parameters representing the reconfigurable hardware, application load imbalance across the nodes, background user load, basic message-passing communication, and processor heterogeneity. In this article, we provide an overview of the analytical model and demonstrate its application for optimization and scheduling of high-performance reconfigurable computing (HPRC) resources. We examine cost functions for minimum runtime and other optimization problems commonly found in shared computing resources. Finally, we discuss additional scheduling issues and other potential applications of the model.
Melissa C. Smith, Gregory D. Peterson
ACM Trans. Embed. Comput. Syst.1
2012 A comparative study of GPU programming models and architectures using neural networks
Vivek K. Pallipuram, Mohammad Ashraf Bhuiyan, Melissa C. Smith
J. Supercomput.3
2011 Performance, optimization, and fitness: Connecting applications to architectures
abstract
Abstract Recent trends involving multicore processors and graphical processing units (GPUs) focus on exploiting task‐ and thread‐level parallelism. In this paper, we have analyzed various aspects of the performance of these architectures including NVIDIA GPUs, and multicore processors such as Intel Xeon, AMD Opteron, IBM's Cell Broadband Engine. The case study used in this paper is a biological spiking neural network (SNN), implemented with the Izhikevich, Wilson, Morris–Lecar, and Hodgkin–Huxley neuron models. The four SNN models have varying requirements for communication and computation making them useful for performance analysis of the hardware platforms. We report and analyze the variation of performance with network (problem size) scaling, available optimization techniques and execution configuration. A Fitness performance model, that predicts the suitability of the architecture for accelerating an application, is proposed and verified with the SNN implementation results. The Roofline model, another existing performance model, has also been utilized to determine the hardware bottleneck(s) and attainable peak performance of the architectures. Significant speedups for the four SNN neuron models utilizing these architectures are reported; the maximum speedup of 574x was observed in our GPU implementation. Our results and analysis show that a proper match of architecture with algorithm complexity provides the best performance. Copyright © 2010 John Wiley & Sons, Ltd.
Mohammad Ashraf Bhuiyan, Melissa C. Smith, Vivek K. Pallipuram
Concurr. Comput. Pract. Exp.2
2011 An analytical model for multilevel performance prediction of Multi-FPGA systems
abstract
Power limitations in semiconductors have made explicitly parallel device architectures such as Field-Programmable Gate Arrays (FPGAs) increasingly attractive for use in scalable systems. However, mitigating the significant cost of FPGA development requires efficient design-space exploration to plan and evaluate a range of potential algorithm and platform choices prior to implementation. The authors propose the RC Amenability Test for Scalable Systems (RATSS), an analytical model which enables straightforward, fast, and reasonably accurate performance prediction prior to implementation by extending current modeling concepts to multi-FPGA designs. RATSS provides a comprehensive strategic model to evaluate applications based on the computation and communication requirements of the algorithm and capabilities of the FPGA platform. The RATSS model targets data-parallel applications on current scalable FPGA systems. Three case studies with RATSS demonstrate nearly 90% prediction accuracy as compared to corresponding implementations.
Brian Holland, Alan D. George, Herman Lam, Melissa C. Smith
ACM Trans. Reconfigurable Technol. Syst.4
2007 An Application Specific Memory Characterization Technique for Co-processor Accelerators
abstract
Commodity accelerator technologies including reconfigurable devices provide an order of magnitude performance improvement compared to mainstream microprocessor systems. A number of compute-intensive scientific applications, therefore, can potentially benefit from commodity computing devices available in the form of co-processor accelerators. However, there has been little progress in accelerating production-level scientific applications using these technologies due to several programming and performance challenges. One of the key perfomance challenges is performance sustainability. While computation is often accelerated substantially by accelerator devices, the achievable performance is significantly lower once the data transfer costs and overheads are incorporated. We present an application-specific memory characterization technique for an FPGA-accelerated system that enabled us to reduce data transfer overhead by a factor of five for a production-scale scientific application. Our proposed technique extends to applications that exhibit similar memory behavior and to co-processor accelerator systems that support data streaming, pipelining, and overlapped execution.
Sadaf R. Alam, Jeffrey S. Vetter, Melissa C. Smith
ASAP3
2006 Simulative analysis of dynamic scheduling heuristics for reconfigurable computing of parallel applications
abstract
The demand for processing power has been ever increasing with the growth of high-performance computing (HPC) applications and so have the constraints restricting the solutions to such requirements. High-performance distributed and parallel computing and custom-built, hardware-based computing have attempted to address this problem with some success but at a substantial cost. Recently, systems augmented with Field-Programmable Gate Arrays (FPGAs) offering a fusion of traditional parallel and distributed machines with customizable and dynamically reconfigurable hardware have emerged as a cost-effective alternative to traditional systems. However, providing a robust runtime environment for such systems to which HPC users have become accustomed has been fraught with numerous challenges. Dynamic scheduling of large-scale HPC applications in such parallel reconfigurable computing (RC) environments is one such challenge and has not been sufficiently studied to our knowledge. In this paper, we simulatively analyze the performance of several common HPC scheduling heuristics that can be used by an automated job management service to schedule application tasks on a parallel RC system. We also present a performance prediction model which the scheduling heuristics employ to schedule several common HPC applications on a collection of typical FPGA processing platforms.
Rajagopal Subramaniyan, Ian A. Troxel, Alan D. George, Melissa C. Smith
FPGA4
2005 Parallel application performance on shared high performance reconfigurable computing resources
Melissa C. Smith, Gregory D. Peterson
Perform. Evaluation1