EDBT 2026 Demo / reviewers in the wild / expert
Melissa C. Smith
dblp:89/3990
· DBLP profile ↗
24ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0003-0798-8536ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 8 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Parallel and multicore computing · 30% High-performance computing · 26% Electronic design automation · 23% | |
| Artificial intelligence
2 papers |
Speech recognition and synthesis · 50% Reinforcement learning · 25% Graph learning · 25% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Speech recognition and synthesis
speech enhancement |
0.5 | 1 | 2021 | Speech Enhancement Using Multi-Stage Self-Attentive Temporal Convolutional Networks · IEEE ACM Trans. Audio Speech Lang. Process. 2021 |
High-performance computing › numerical linear algebra › linear solver
iterative linear solvers |
0.5 | 1 | 2021 | A Parallel Jacobi-Embedded Gauss-Seidel Method · IEEE Trans. Parallel Distributed Syst. 2021 |
Parallel and multicore computing › parallel algorithms › parallel matrix algorithms
parallel iterative solvers |
0.5 | 1 | 2021 | A Parallel Jacobi-Embedded Gauss-Seidel Method · IEEE Trans. Parallel Distributed Syst. 2021 |
Machine learning › Graph learning › graph neural network
graph convolution |
0.3 | 1 | 2018 | Generalized Value Iteration Networks: Life Beyond Lattices · AAAI 2018 |
Machine learning › Reinforcement learning › value-based reinforcement learning
value iteration network |
0.3 | 1 | 2018 | Generalized Value Iteration Networks: Life Beyond Lattices · AAAI 2018 |
Electronic design automation
design flow |
0.2 | 1 | 2015 | Enhancing Hardware Design Flows with MyHDL · FPGA 2015 |
Electronic design automation
hardware description language |
0.2 | 1 | 2015 | Enhancing Hardware Design Flows with MyHDL · FPGA 2015 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.1 | 1 | 2021 | Speech Enhancement Using Multi-Stage Self-Attentive Temporal Convolutional Networks · IEEE ACM Trans. Audio Speech Lang. Process. 2021 |
Processor architecture and microarchitecture
many-core architecture |
0.1 | 1 | 2021 | A Parallel Jacobi-Embedded Gauss-Seidel Method · IEEE Trans. Parallel Distributed Syst. 2021 |
Performance modeling and evaluation
simulation |
0.1 | 1 | 2015 | Enhancing Hardware Design Flows with MyHDL · FPGA 2015 |
Embedded and real-time systems › real-time scheduling
heuristic scheduling |
0.1 | 1 | 2006 | Simulative analysis of dynamic scheduling heuristics for reconfigurable computing of parallel applications · FPGA 2006 |
Cloud and datacenter computing
job scheduling |
0.1 | 1 | 2006 | Simulative analysis of dynamic scheduling heuristics for reconfigurable computing of parallel applications · FPGA 2006 |
Reconfigurable computing and FPGAs
reconfigurable computing |
0.1 | 1 | 2006 | Simulative analysis of dynamic scheduling heuristics for reconfigurable computing of parallel applications · FPGA 2006 |
Parallel and multicore computing › parallel scheduling
runtime scheduling |
0.1 | 1 | 2006 | Simulative analysis of dynamic scheduling heuristics for reconfigurable computing of parallel applications · FPGA 2006 |
Methods — techniques the papers use, named apart from their topics
temporal convolutional network · 0.5self-attention · 0.5multi-stage learning · 0.5matrix reordering · 0.5jacobi iteration · 0.5domain decomposition · 0.5value iteration · 0.3q-learning · 0.3graph convolution · 0.3verilog · 0.2python · 0.2VHDL · 0.2simulation · 0.1performance prediction model · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GEMDiff: a diffusion workflow bridges between normal and tumor gene expression states: a breast cancer case studyabstractBreast cancer remains a significant global health challenge due to its complexity, which arises from multiple genetic and epigenetic mutations that originate in normal breast tissue. Traditional machine learning models often fall short in addressing the intricate gene interactions that complicate drug design and treatment strategies. In contrast, our study introduces GEMDiff, a novel computational workflow leveraging a diffusion model to bridge the gene expression states between normal and tumor conditions. GEMDiff augments RNAseq data and simulates perturbation transformations between normal and tumor gene states, enhancing biomarker identification. GEMDiff can handle large-scale gene expression data without succumbing to the scalability and stability issues that plague other generative models. By avoiding the need for task-specific hyper-parameter tuning and specific loss functions, GEMDiff can be generalized across various tasks, making it a robust tool for gene expression analysis. The model's ability to augment RNA-seq data and simulate gene perturbations provides a valuable tool for researchers. This capability can be used to generate synthetic data for training other machine learning models, thereby addressing the issue of limited biological data and enhancing the performance of predictive models. The effectiveness of GEMDiff is demonstrated through a case study using breast mRNA gene expression data, identifying 307 core genes involved in the transition from a breast tumor to a normal gene expression state. GEMDiff is open source and available at https://github.com/xai990/GEMDiff.git under the MIT license. Xusheng Ai, Melissa C. Smith, F. Alex Feltus |
Briefings Bioinform. | 2 |
| 2022 | Lossy Compression to Reduce Latency of Local Image Transfer for Autonomous Off-Road Perception SystemsabstractAutonomous vehicles greatly rely on their perception system for navigation. Semantic segmentation provides a much better understanding of a vehicle’s surroundings than object detection. Unfortunately, complete image segmentation comes at a higher computational cost than object detection, which complicates developing a real-time perception system using semantic segmentation. Perception systems contain other bottlenecks too, and are not only limited by their deep learning model. An inherent amount of latency exists in data transfer, specifically through Ethernet. A vehicle’s camera feed must be transferred to an edge device for image processing as part of the autonomous driving decision-making process. This study investigates decreasing image transfer time by using various levels of JPEG compression as well as further understanding how compression affects the accuracy of semantic segmentation. Additionally, as most autonomous driving research focuses on urban environments, we look to explore autonomous unmanned ground vehicles (UGVs) in the off-road space by using the Rellis-3D dataset. We train and evaluate SwiftNet, a state-of-the-art semantic segmentation model, at different JPEG compression ratios and identify the accuracy. The transfer time of these different compression ratios is tested on three images. Results show a continual decrease in accuracy occurs as the compression ratios increase. When training SwiftNet on the train set with no compression, the highest compression ratio of 16.96 achieves a mean intersection over union (mIoU) score of 67.9% compared to the baseline achieving 78.9% mIoU. There is an increase in the accuracy of the higher compression ratios by training SwiftNet on the corresponding compression ratios; the highest compression ratio reaches 74.9% mIoU. Lastly, we notice a positive transfer speedup of these higher compression ratios when inducing JPEG compression in all transfer scenarios: (a) 1870 images (b) 10 images, (c) 1 image. Each scenario has a speedup of 1.18×, 1.14×, and 1.06×, respectively. Max H. Faykus, Bradley Selee, Jon Calhoun 0001, Melissa C. Smith |
IEEE Big Data | 4 |
| 2022 | Addressing noise in co-expression network constructionabstractGene co-expression networks (GCNs) provide multiple benefits to molecular research including hypothesis generation and biomarker discovery. Transcriptome profiles serve as input for GCN construction and are derived from increasingly larger studies with samples across multiple experimental conditions, treatments, time points, genotypes, etc. Such experiments with larger numbers of variables confound discovery of true network edges, exclude edges and inhibit discovery of context (or condition) specific network edges. To demonstrate this problem, a 475-sample dataset is used to show that up to 97% of GCN edges can be misleading because correlations are false or incorrect. False and incorrect correlations can occur when tests are applied without ensuring assumptions are met, and pairwise gene expression may not meet test assumptions if the expression of at least one gene in the pairwise comparison is a function of multiple confounding variables. The 'one-size-fits-all' approach to GCN construction is therefore problematic for large, multivariable datasets. Recently, the Knowledge Independent Network Construction toolkit has been used in multiple studies to provide a dynamic approach to GCN construction that ensures statistical tests meet assumptions and confounding variables are addressed. Additionally, it can associate experimental context for each edge of the network resulting in context-specific GCNs (csGCNs). To help researchers recognize such challenges in GCN construction, and the creation of csGCNs, we provide a review of the workflow. Josh J. R. Burns, Benjamin T. Shealy, Mitchell S. Greer, John A. Hadish, Matthew McGowan, Tyler Biggs, Melissa C. Smith, F. Alex Feltus, Stephen P. Ficklin |
Briefings Bioinform. | 7 |
| 2022 | GEMmaker: process massive RNA-seq datasets on heterogeneous computational infrastructureabstractBACKGROUND: Quantification of gene expression from RNA-seq data is a prerequisite for transcriptome analysis such as differential gene expression analysis and gene co-expression network construction. Individual RNA-seq experiments are larger and combining multiple experiments from sequence repositories can result in datasets with thousands of samples. Processing hundreds to thousands of RNA-seq data can result in challenges related to data management, access to sufficient computational resources, navigation of high-performance computing (HPC) systems, installation of required software dependencies, and reproducibility. Processing of larger and deeper RNA-seq experiments will become more common as sequencing technology matures. RESULTS: GEMmaker, is a nf-core compliant, Nextflow workflow, that quantifies gene expression from small to massive RNA-seq datasets. GEMmaker ensures results are highly reproducible through the use of versioned containerized software that can be executed on a single workstation, institutional compute cluster, Kubernetes platform or the cloud. GEMmaker supports popular alignment and quantification tools providing results in raw and normalized formats. GEMmaker is unique in that it can scale to process thousands of local or remote stored samples without exceeding available data storage. CONCLUSIONS: Workflows that quantify gene expression are not new, and many already address issues of portability, reusability, and scale in terms of access to CPUs. GEMmaker provides these benefits and adds the ability to scale despite low data storage infrastructure. This allows users to process hundreds to thousands of RNA-seq samples even when data storage resources are limited. GEMmaker is freely available and fully documented with step-by-step setup and execution instructions. John A. Hadish, Tyler Biggs, Benjamin T. Shealy, M. Reed Bender, Coleman McKnight, Connor Wytko, Melissa C. Smith, F. Alex Feltus, Loren A. Honaas, Stephen P. Ficklin |
BMC Bioinform. | 7 |
| 2021 | TIGRA: A Tightly Integrated Generic RISC-V Accelerator InterfaceabstractField programmable gate array (FPGA) usage in HPC applications is growing with the need for energy efficient and application specific accelerators. Currently, FPGAs are used to accelerate algorithms using OpenCL with communication over PCIe (loosely coupled accelerators) or by modifying existing architectures to incorporate custom logic directly with a CPU (tightly coupled accelerators). However, only the loosely coupled paradigm is feasible to support a variety of acceleration. In this work, we introduce TIGRA, a zero latency interface designed to provide the benefit of tightly coupled accelerators without the developer burden of modifying the underlying architecture, which can enable their usage in HPC. TIGRA is demonstrated on the PicoRV32 processor with AES-128 bit encryption, posit arithmetic, and multiplication. Brad Green, Dillon Todd, Jon Calhoun 0001, Melissa C. Smith |
CLUSTER | 4 |
| 2021 | Systematic Evaluation and Enhancement of Speech Recognition in Operational Medical EnvironmentsabstractOperational medical environments require reliable hands-free solutions to extract data from audio captured under noisy scenarios during rescue missions and provide timely information. However, approaches using automatic speech recognition (ASR) and natural language processing (NLP) techniques are complex as these conversations have a wide range of noise, involve medical terms from multiple speakers, and occur in high-stress environments, among others. These are further complicated by the lack of large training datasets for operational medical scenarios. To address these issues, we developed a platform that enables resilient hands-free data collection, preserves complete documentation through stages of care, and presents the information in near real-time, critical for the medical operation. Our work uniquely focused on systematic evaluation and improvement of a deep neural network-based ASR system by leveraging realistic testing data obtained from medical simulations of battlefield scenarios, which to our knowledge have not been addressed in any prior work. The system performance is shown to improve significantly using multi-style training, language model adaptation for the medical domain, speech enhancement, and NLP techniques. Snigdhaswin Kar, Prabodh Mishra, Ju Lin, Minjae Woo, Nicholas Deas, Caleb Linduff, Sufeng Niu, Jerome McClendon, D. Hudson Smith, Melissa C. Smith, Ronald W. Gimbel, Kuang-Ching Wang |
IJCNN | 11 |
| 2021 | Speech Enhancement Using Multi-Stage Self-Attentive Temporal Convolutional NetworksabstractMulti-stage learning is an effective technique to invoke multiple deep-learning modules sequentially. This paper applies multi-stage learning to speech enhancement by using a multi-stage structure, where each stage comprises a self-attention (SA) block followed by stacks of temporal convolutional network (TCN) blocks with doubling dilation factors. Each stage generates a prediction that is refined in a subsequent stage. A fusion block is inserted at the input of later stages to re-inject original information. The resulting multi-stage speech enhancement system, in short, multi-stage SA-TCN, is compared with state-of-the-art deep-learning speech enhancement methods using the LibriSpeech and VCTK data sets. The multi-stage SA-TCN system's hyper-parameters are fine-tuned, and the impact of the SA block, the fusion block and the number of stages are determined. The use of a multi-stage SA-TCN system as a front-end for automatic speech recognition systems is investigated as well. It is shown that the multi-stage SA-TCN systems perform well relative to other state-of-the-art systems in terms of speech enhancement and speech recognition scores. Ju Lin, Adriaan J. de Lind van Wijngaarden, Kuang-Ching Wang, Melissa C. Smith |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | A Parallel Jacobi-Embedded Gauss-Seidel MethodabstractA broad range of scientific simulations involve solving large-scale computationally expensive linear systems of equations. Iterative solvers are typically preferred over direct methods when it comes to large systems due to their lower memory requirements and shorter execution times. However, selecting the appropriate iterative solver is problem-specific and dependent on the type and symmetry of the coefficient matrix. Gauss-Seidel (GS) is an iterative method for solving linear systems that are either strictly diagonally dominant or symmetric positive definite. This technique is an improved version of Jacobi and typically converges in fewer iterations. However, the sequential nature of this algorithm complicates the parallel extraction. In fact, most parallel derivatives of GS rely on the sparsity pattern of the coefficient matrix and require matrix reordering or domain decomposition. In this article, we introduce a new algorithm that exploits the convergence property of GS and adapts the parallel structure of Jacobi. The proposed method works for both dense and sparse systems and is straightforward to implement. We have examined the performance of our method on multicore and many-core architectures. Experimental results demonstrate the superior performance of the proposed algorithm compared with GS and Jacobi. Additionally, performance comparison with built-in Krylov solvers in MATLAB showed that in terms of time per iteration, Krylov methods perform faster on CPUs, but our approach is significantly better when executed on GPUs. Lastly, we apply our method to solve the power flow problem, and the results indicate a significant improvement in runtime, reaching up to 87 times faster speed compared with GS. Afshin Ahmadi, Felice Manganiello, Amin Khademi, Melissa C. Smith |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2020 | Improved Speech Enhancement Using a Time-Domain GAN with Mask LearningabstractSpeech enhancement is an essential component in robust automatic speech recognition (ASR) systems. Most speech enhancement methods are nowadays based on neural networks that use feature-mapping or mask-learning. This paper proposes a novel speech enhancement method that integrates time-domain feature mapping and mask learning into a unified framework using a Generative Adversarial Network (GAN). The proposed framework processes the received waveform and decouples speech and noise signals, which are fed into two short-time Fourier transform (STFT) convolution 1-D layers that map the waveforms to spectrograms in the complex domain. These speech and noise spectrograms are then used to compute the speech mask loss. The proposed method is evaluated using the TIMIT data set for seen and unseen signal-to-noise ratio conditions. It is shown that the proposed method outperforms the speech enhancement methods that use Deep Neural Network (DNN) based speech enhancement or a Speech Enhancement Generative Adversarial Network (SEGAN). Ju Lin, Sufeng Niu, Adriaan J. de Lind van Wijngaarden, Jerome McClendon, Melissa C. Smith, Kuang-Ching Wang |
INTERSPEECH | 5 |
| 2019 | Speech Enhancement Using Forked Generative Adversarial Networks with Spectral Subtraction
Ju Lin, Sufeng Niu, Zice Wei, Adriaan J. de Lind van Wijngaarden, Melissa C. Smith, Kuang-Ching Wang |
INTERSPEECH | 6 |
| 2018 | Generalized Value Iteration Networks: Life Beyond LatticesabstractIn this paper, we introduce a generalized value iteration network (GVIN), which is an end-to-end neural network planning module. GVIN emulates the value iteration algorithm by using a novel graph convolution operator, which enables GVIN to learn and plan on irregular spatial graphs. We propose three novel differentiable kernels as graph convolution operators and show that the embedding-based kernel achieves the best performance. Furthermore, we present episodic Q-learning, an improvement upon traditional n-step Q-learning that stabilizes training for VIN and GVIN. Lastly, we evaluate GVIN on planning problems in 2D mazes, irregular graphs, and real-world street networks, showing that GVIN generalizes well for both arbitrary graphs and unseen graphs of larger scaleand outperforms a naive generalization of VIN (discretizing a spatial graph into a 2D image). Sufeng Niu, Siheng Chen, Hanyu Guo, Colin Targonski, Melissa C. Smith, Jelena Kovacevic |
AAAI | 5 |
| 2018 | Artificial Intelligence and Deep Learning Applications for Automotive ManufacturingabstractArtificial Intelligence (AI) and Deep Learning has been steadily gaining importance due it’s potential for a broad set of science and industry applications. The success of deep learning techniques has found many applications, e.g. in the domain of computer vision and natural language understanding. Developing AI applications is a complex task with many challenges related to data collection, model training, and deployment.In this paper, we evaluate architectures, models and deployment issues related to the usage of deep learning techniques in the automotive manufacturing domain. Particularly, we focus on different computer vision problems in automotive manufacturing processes, e.g., in logistics processes. We developed several deep learning models that help to improve the quality and efficiency of these processes. Finally, we provide an analysis of the architecture, datasets and models used, and provide performance metrics for each of the different models. André Luckow, Ken Kennedy, Marcin Ziolkowski, Emil Djerekarov, Matthew Cook 0004, Edward B. Duffy, Michael Schleiss, Bennie Vorster, Edwin Weill, Ankit Kulshrestha, Melissa C. Smith |
IEEE BigData | 11 |
| 2015 | Enhancing Hardware Design Flows with MyHDLabstractMyHDL is a Python based HDL that harnesses the power and versatility of Python for hardware development. MyHDL has excellent simulation capabilities and also allows for conversion to Verilog and VHDL, so developers can enter a conventional design flow as desired. Verilog and VHDL are used extensively, particularly because most synthesis tools only support these two languages. However, they are simply outdated; poor parameterization limits high level design and modern abstraction features such as classes are missing. Keerthan Jaic, Melissa C. Smith |
FPGA | 2 |
| 2015 | Subjective versus objective: classifying analytical models for productive heterogeneous performance prediction
Vivek K. Pallipuram, Melissa C. Smith, Nilim Sarma, Ranajeet Anand, Edwin Weill, Karan Sapra |
J. Supercomput. | 2 |
| 2014 | A practical network intrusion detection system for inline FPGAs on 10GbE network adaptersabstractA network intrusion detection system (NIDS), such as SNORT, analyzes incoming packets to identify potential security threats. Pattern matching is arguably the most important and most computationally intensive component of a NIDS. Software-based NIDS implementations drop up to 90% of packets during increased network load even at lower network bandwidth. We propose an alternative hybrid-NIDS that couples an FPGA with a network adapter to provide hardware support for pattern matching and software support for post processing. The proposed system, SFAOENIDS, offers an extensible open-source NIDS for Solarflare AOE devices. The pattern matching engine-the primary component of the hardware architecture was designed based on the requirements of typical NIDS implementations. In testing on a real network environment, the SFAOENIDS hardware implementation, operating at 200 MHz, handles a 10Gbps data rate without dropping packets while simultaneously minimizing the server CPU load. Keerthan Jaic, Melissa C. Smith, Nilim Sarma |
ASAP | 2 |
| 2014 | Combining Hadoop and GPU to preprocess large Affymetrix microarray dataabstractHigh density oligonucleotide array (microarray) from Affymetrix has been widely used for the measurements of gene expressions. Currently, public data repositories, such as Gene Expression Omnibus (GEO) of the National Center for Biotechnology Information (NCBI), have accumulated large amounts of microarray data. Efficient integrative analysis of those microarray data will provide significant knowledge about biological systems. None of the existing microarray preprocessing and quality assessment tools can handle very large microarray datasets with tens of thousands of experiments. The preprocessing and quality assessment of microarray datasets contain both data-intensive and compute-intensive tasks. In this paper, we develop a new set of tools using a mix of the Hadoop (for data intensive tasks) and the General-Purpose Graphics Processing Units (GPGPUs) (for compute intensive tasks) to efficiently process large microarray data. Evaluation of our new tools on large microarray datasets with ten thousands of experiments showed promising superior performance. We demonstrate that the combination of Hadoop and GPGPU computation is effective for complex scientific applications that contain both data-intensive and compute-intensive tasks. Our new tool set will make it possible to utilize valuable large microarray data in the public repositories. Sufeng Niu, Nilim Sarma, Pengfei Xuan, Melissa C. Smith, Pradip K. Srimani, Feng Luo 0001 |
IEEE BigData | 5 |
| 2014 | A regression-based performance prediction framework for synchronous iterative algorithms on general purpose graphical processing unit clustersabstractSUMMARY Heterogeneous performance prediction models are valuable tools to accurately predict application runtime, allowing for efficient design space exploration and application mapping. The existing performance models require intricate system architecture knowledge, making the modeling task difficult. In this research, we propose a regression‐based performance prediction framework for general purpose graphical processing unit (GPGPU) clusters that statistically abstracts the system architecture characteristics, enabling performance prediction without detailed system architecture knowledge. The regression‐based framework targets deterministic synchronous iterative algorithms using our synchronous iterative GPGPU execution model and is broken into two components: the computation component that models the GPGPU device and host computations and the communication component that models the network‐level communications. The computation component regression models use algorithm characteristics such as the number of floating‐point operations and total bytes as predictor variables and are trained using several small, instrumented executions of synchronous iterative algorithms that include a range of floating‐point operations‐to‐byte requirements. The regression models for network‐level communications are developed using micro‐benchmarks and employ data transfer size and processor count as predictor variables. Our performance prediction framework achieves prediction accuracy over 90% compared with the actual implementations for several tested GPGPU cluster configurations. The end goal of this research is to offer the scientific computing community, an accurate and easy‐to‐use performance prediction framework that empowers users to optimally utilize the heterogeneous resources. Copyright © 2013 John Wiley & Sons, Ltd. Vivek K. Pallipuram, Melissa C. Smith, Nimisha Raut |
Concurr. Comput. Pract. Exp. | 2 |
| 2012 | Optimization of Shared High-Performance Reconfigurable Computing ResourcesabstractIn the field of high-performance computing, systems harboring reconfigurable devices, such as field-programmable gate arrays (FPGAs), are gaining more widespread interest. Such systems range from supercomputers with tightly coupled reconfigurable hardware to clusters with reconfigurable devices at each node. The use of these architectures for scientific computing provides an alternative for computationally demanding problems and has advantages in metrics, such as operating cost/performance and power/performance. However, performance optimization of these systems can be challenging even with knowledge of the system’s characteristics. Our analytic performance model includes parameters representing the reconfigurable hardware, application load imbalance across the nodes, background user load, basic message-passing communication, and processor heterogeneity. In this article, we provide an overview of the analytical model and demonstrate its application for optimization and scheduling of high-performance reconfigurable computing (HPRC) resources. We examine cost functions for minimum runtime and other optimization problems commonly found in shared computing resources. Finally, we discuss additional scheduling issues and other potential applications of the model. Melissa C. Smith, Gregory D. Peterson |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2012 | A comparative study of GPU programming models and architectures using neural networks
Vivek K. Pallipuram, Mohammad Ashraf Bhuiyan, Melissa C. Smith |
J. Supercomput. | 3 |
| 2011 | Performance, optimization, and fitness: Connecting applications to architecturesabstractAbstract Recent trends involving multicore processors and graphical processing units (GPUs) focus on exploiting task‐ and thread‐level parallelism. In this paper, we have analyzed various aspects of the performance of these architectures including NVIDIA GPUs, and multicore processors such as Intel Xeon, AMD Opteron, IBM's Cell Broadband Engine. The case study used in this paper is a biological spiking neural network (SNN), implemented with the Izhikevich, Wilson, Morris–Lecar, and Hodgkin–Huxley neuron models. The four SNN models have varying requirements for communication and computation making them useful for performance analysis of the hardware platforms. We report and analyze the variation of performance with network (problem size) scaling, available optimization techniques and execution configuration. A Fitness performance model, that predicts the suitability of the architecture for accelerating an application, is proposed and verified with the SNN implementation results. The Roofline model, another existing performance model, has also been utilized to determine the hardware bottleneck(s) and attainable peak performance of the architectures. Significant speedups for the four SNN neuron models utilizing these architectures are reported; the maximum speedup of 574x was observed in our GPU implementation. Our results and analysis show that a proper match of architecture with algorithm complexity provides the best performance. Copyright © 2010 John Wiley & Sons, Ltd. Mohammad Ashraf Bhuiyan, Melissa C. Smith, Vivek K. Pallipuram |
Concurr. Comput. Pract. Exp. | 2 |
| 2011 | An analytical model for multilevel performance prediction of Multi-FPGA systemsabstractPower limitations in semiconductors have made explicitly parallel device architectures such as Field-Programmable Gate Arrays (FPGAs) increasingly attractive for use in scalable systems. However, mitigating the significant cost of FPGA development requires efficient design-space exploration to plan and evaluate a range of potential algorithm and platform choices prior to implementation. The authors propose the RC Amenability Test for Scalable Systems (RATSS), an analytical model which enables straightforward, fast, and reasonably accurate performance prediction prior to implementation by extending current modeling concepts to multi-FPGA designs. RATSS provides a comprehensive strategic model to evaluate applications based on the computation and communication requirements of the algorithm and capabilities of the FPGA platform. The RATSS model targets data-parallel applications on current scalable FPGA systems. Three case studies with RATSS demonstrate nearly 90% prediction accuracy as compared to corresponding implementations. Brian Holland, Alan D. George, Herman Lam, Melissa C. Smith |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2007 | An Application Specific Memory Characterization Technique for Co-processor AcceleratorsabstractCommodity accelerator technologies including reconfigurable devices provide an order of magnitude performance improvement compared to mainstream microprocessor systems. A number of compute-intensive scientific applications, therefore, can potentially benefit from commodity computing devices available in the form of co-processor accelerators. However, there has been little progress in accelerating production-level scientific applications using these technologies due to several programming and performance challenges. One of the key perfomance challenges is performance sustainability. While computation is often accelerated substantially by accelerator devices, the achievable performance is significantly lower once the data transfer costs and overheads are incorporated. We present an application-specific memory characterization technique for an FPGA-accelerated system that enabled us to reduce data transfer overhead by a factor of five for a production-scale scientific application. Our proposed technique extends to applications that exhibit similar memory behavior and to co-processor accelerator systems that support data streaming, pipelining, and overlapped execution. Sadaf R. Alam, Jeffrey S. Vetter, Melissa C. Smith |
ASAP | 3 |
| 2006 | Simulative analysis of dynamic scheduling heuristics for reconfigurable computing of parallel applicationsabstractThe demand for processing power has been ever increasing with the growth of high-performance computing (HPC) applications and so have the constraints restricting the solutions to such requirements. High-performance distributed and parallel computing and custom-built, hardware-based computing have attempted to address this problem with some success but at a substantial cost. Recently, systems augmented with Field-Programmable Gate Arrays (FPGAs) offering a fusion of traditional parallel and distributed machines with customizable and dynamically reconfigurable hardware have emerged as a cost-effective alternative to traditional systems. However, providing a robust runtime environment for such systems to which HPC users have become accustomed has been fraught with numerous challenges. Dynamic scheduling of large-scale HPC applications in such parallel reconfigurable computing (RC) environments is one such challenge and has not been sufficiently studied to our knowledge. In this paper, we simulatively analyze the performance of several common HPC scheduling heuristics that can be used by an automated job management service to schedule application tasks on a parallel RC system. We also present a performance prediction model which the scheduling heuristics employ to schedule several common HPC applications on a collection of typical FPGA processing platforms. Rajagopal Subramaniyan, Ian A. Troxel, Alan D. George, Melissa C. Smith |
FPGA | 4 |
| 2005 | Parallel application performance on shared high performance reconfigurable computing resources
Melissa C. Smith, Gregory D. Peterson |
Perform. Evaluation | 1 |