Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Gregory D. Peterson

dblp:26/4323 · DBLP profile ↗
← Back
26ranked-venue papers
1as first author
1since 2021 · last 2022
0000-0002-0875-5278ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 1 first-authorArtificial intelligence and machine learning · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
High-performance computing · 35% Performance modeling and evaluation · 24% Reconfigurable computing and FPGAs · 22%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 24 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA accelerator
0.442012
A High Performance and Memory Efficient LU Decomposer on FPGAs · IEEE Trans. Computers 2012
High-Performance Mixed-Precision Linear Solver for FPGAs · IEEE Trans. Computers 2008
Poster reception - Reconfigurable accelerator for quantum Monte Carlo simulations in N-body systems · SC 2006
High-performance computing › numerical linear algebra
dense matrix decomposition
0.112012
A High Performance and Memory Efficient LU Decomposer on FPGAs · IEEE Trans. Computers 2012
High-performance computing › numerical linear algebra
linear algebra kernel
0.112012
A High Performance and Memory Efficient LU Decomposer on FPGAs · IEEE Trans. Computers 2012
High-performance computing › numerical linear algebra › matrix factorization
LU decomposition
0.112012
A High Performance and Memory Efficient LU Decomposer on FPGAs · IEEE Trans. Computers 2012
Performance modeling and evaluation › parallel system performance
parallel performance modeling
0.112012
An Effective Execution Time Approximation Method for Parallel Computing · IEEE Trans. Parallel Distributed Syst. 2012
Performance modeling and evaluation › system-level analysis › architecture evaluation
accelerator comparison
0.112011
Comparing Hardware Accelerators in Scientific Applications: A Case Study · IEEE Trans. Parallel Distributed Syst. 2011
GPUs and heterogeneous computing
GPU computing
0.112011
Comparing Hardware Accelerators in Scientific Applications: A Case Study · IEEE Trans. Parallel Distributed Syst. 2011
GPUs and heterogeneous computing › heterogeneous programming models
OpenCL
0.112011
Comparing Hardware Accelerators in Scientific Applications: A Case Study · IEEE Trans. Parallel Distributed Syst. 2011
High-performance computing › numerical linear algebra
linear solver
0.112008
High-Performance Mixed-Precision Linear Solver for FPGAs · IEEE Trans. Computers 2008
High-performance computing
scientific computing systems
0.112008
High-Performance Mixed-Precision Linear Solver for FPGAs · IEEE Trans. Computers 2008
Performance modeling and evaluation
pseudorandom number generation
0.112006
Poster reception - A reconfigurable supercomputing library for accelerated parallel lagged-Fibonacci pseudorandom number generation · SC 2006
Emerging computing paradigms › quantum computing › quantum simulation
quantum monte carlo simulation
0.112006
Poster reception - Reconfigurable accelerator for quantum Monte Carlo simulations in N-body systems · SC 2006
Bioinformatics and computational biology › synthetic biology
genetic circuits
0.012004
Engineering in the biological substrate: information processing in genetic circuits · Proc. IEEE 2004
Bioinformatics and computational biology
synthetic biology
0.012004
Engineering in the biological substrate: information processing in genetic circuits · Proc. IEEE 2004
Reconfigurable computing and FPGAs
FPGA architecture
0.012012
A High Performance and Memory Efficient LU Decomposer on FPGAs · IEEE Trans. Computers 2012
Reconfigurable computing and FPGAs › coarse-grained reconfigurable architecture
processing element array
0.012012
A High Performance and Memory Efficient LU Decomposer on FPGAs · IEEE Trans. Computers 2012
High-performance computing › scientific computing
scientific computing application
0.012011
Comparing Hardware Accelerators in Scientific Applications: A Case Study · IEEE Trans. Parallel Distributed Syst. 2011
Processor architecture and microarchitecture › computer arithmetic
floating-point arithmetic
0.012008
High-Performance Mixed-Precision Linear Solver for FPGAs · IEEE Trans. Computers 2008
Performance modeling and evaluation › simulation
monte carlo simulation
0.012006
Poster reception - A reconfigurable supercomputing library for accelerated parallel lagged-Fibonacci pseudorandom number generation · SC 2006
High-performance computing
n-body simulation
0.012006
Poster reception - Reconfigurable accelerator for quantum Monte Carlo simulations in N-body systems · SC 2006
Parallel and multicore computing › parallel computing
parallel scientific computing
0.012006
Poster reception - A reconfigurable supercomputing library for accelerated parallel lagged-Fibonacci pseudorandom number generation · SC 2006
High-performance computing
scientific computing
0.012006
Poster reception - Reconfigurable accelerator for quantum Monte Carlo simulations in N-body systems · SC 2006
Integrated circuit design › digital logic
logic gate
0.012004
Engineering in the biological substrate: information processing in genetic circuits · Proc. IEEE 2004
Integrated circuit design › analog and mixed-signal circuits
oscillator design
0.012004
Engineering in the biological substrate: information processing in genetic circuits · Proc. IEEE 2004

Methods — techniques the papers use, named apart from their topics

neural network · 1.1domain adaptation · 1.1adversarial learning · 1.1space-time mapping · 0.1numerical approximation · 0.1loop blocking · 0.1extreme value theory · 0.1block-cyclic distribution · 0.1block LU decomposition · 0.1performance evaluation · 0.1brook+ · 0.1VHDL · 0.1CUDA · 0.1simulation · 0.0silicon mimetic approach · 0.0modeling · 0.0
YearPublicationVenuePosition
2022 BioADAPT-MRC: adversarial learning-based domain adaptation improves biomedical machine reading comprehension task
abstract
MOTIVATION: Biomedical machine reading comprehension (biomedical-MRC) aims to comprehend complex biomedical narratives and assist healthcare professionals in retrieving information from them. The high performance of modern neural network-based MRC systems depends on high-quality, large-scale, human-annotated training datasets. In the biomedical domain, a crucial challenge in creating such datasets is the requirement for domain knowledge, inducing the scarcity of labeled data and the need for transfer learning from the labeled general-purpose (source) domain to the biomedical (target) domain. However, there is a discrepancy in marginal distributions between the general-purpose and biomedical domains due to the variances in topics. Therefore, direct-transferring of learned representations from a model trained on a general-purpose domain to the biomedical domain can hurt the model's performance. RESULTS: We present an adversarial learning-based domain adaptation framework for the biomedical machine reading comprehension task (BioADAPT-MRC), a neural network-based method to address the discrepancies in the marginal distributions between the general and biomedical domain datasets. BioADAPT-MRC relaxes the need for generating pseudo labels for training a well-performing biomedical-MRC model. We extensively evaluate the performance of BioADAPT-MRC by comparing it with the best existing methods on three widely used benchmark biomedical-MRC datasets-BioASQ-7b, BioASQ-8b and BioASQ-9b. Our results suggest that without using any synthetic or human-annotated data from the biomedical domain, BioADAPT-MRC can achieve state-of-the-art performance on these datasets. AVAILABILITY AND IMPLEMENTATION: BioADAPT-MRC is freely available as an open-source project at https://github.com/mmahbub/BioADAPT-MRC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Maria Mahbub, Sudarshan Srinivasan, Edmon Begoli, Gregory D. Peterson
Bioinform.4
2020 AIR: Iterative refinement acceleration using arbitrary dynamic precision
Gregory D. Peterson, Dimitrios S. Nikolopoulos, Hans Vandierendonck
Parallel Comput.2
2014 Specmaster: an OpenCL-based peptide search engine for tandem mass spectrometry
abstract
SUMMARY Graphics processing units and multicore processors are now pervasive in computational sciences and high‐performance computing. Their high arithmetic throughput and memory bandwidth combined with their ever increasing programmability make them suitable for a widening variety of applications. We give a high level overview of Specmaster, a Myrimatch port that can use every open computing language device available in a machine to identify peptides in tandem mass spectrometry data. We then highlight device‐specific optimizations for multi‐core CPUs and graphics processing units as well as describe our framework for implementing these optimizations while still using a single code base. We also provide performance results of Specmaster running on four different architectures and compare these numbers with Myrimatch. Finally, we improve on our existing work by showing Specmaster dynamically load‐balancing on 3 Radeon 7970s and 32 Interlagos 6272 cores as well as comparing the quality of our search results with Myrimatch. Copyright © 2013 John Wiley & Sons, Ltd.
Rick Weber, David D. Jenkins, Gregory D. Peterson
Concurr. Comput. Pract. Exp.3
2012 From CUDA to OpenCL: Towards a performance-portable solution for multi-platform GPU programming
Rick Weber, Piotr Luszczek, Stanimire Tomov, Gregory D. Peterson, Jack J. Dongarra
Parallel Comput.5
2012 Application accelerators in HPC - Editorial introduction
Volodymyr V. Kindratenko, Gregory D. Peterson
Parallel Comput.2
2012 Compressed sensing and Cholesky decomposition on FPGAs and GPUs
Depeng Yang, Gregory D. Peterson, Husheng Li
Parallel Comput.2
2012 A High Performance and Memory Efficient LU Decomposer on FPGAs
abstract
LU decomposition for dense matrices is an important linear algebra kernel that is widely used in both scientific and engineering applications. To efficiently perform large matrix LU decomposition on FPGAs with limited local memory, a block LU decomposition algorithm on FPGAs applicable to arbitrary matrix size is proposed. Our algorithm applies a series of transformations, including loop blocking and space-time mapping, onto sequential nonblocking LU decomposition. We also introduce a high performance and memory efficient hardware architecture, which mainly consists of a linear array of processing elements (PEs), to implement our block LU decomposition algorithm. Our design can achieve optimum performance under various hardware resource constraints. Furthermore, our algorithm and design can be easily extended to the multi-FPGA platform by using a block-cyclic data distribution and inter-FPGA communication scheme. A total of 36 PEs can be integrated into a Xilinx Virtex-5 XC5VLX330 FPGA on our self-designed PCI-Express card, reaching a sustained performance of 8.50 GFLOPS at 133 MHz for a matrix size of 16,384, which outperforms several general-purpose processors. For a Xilinx Virtex-6 XC6VLX760, a newer FPGA, we predict that a total of 180 PEs can be integrated, reaching 70.66 GFLOPS at 200 MHz. Compared to the previous work, our design can integrate twice the number of PEs into the same FPGA and has significantly higher performance.
Guiming Wu, Yong Dou, Junqing Sun, Gregory D. Peterson
IEEE Trans. Computers4
2012 Optimization of Shared High-Performance Reconfigurable Computing Resources
abstract
In the field of high-performance computing, systems harboring reconfigurable devices, such as field-programmable gate arrays (FPGAs), are gaining more widespread interest. Such systems range from supercomputers with tightly coupled reconfigurable hardware to clusters with reconfigurable devices at each node. The use of these architectures for scientific computing provides an alternative for computationally demanding problems and has advantages in metrics, such as operating cost/performance and power/performance. However, performance optimization of these systems can be challenging even with knowledge of the system’s characteristics. Our analytic performance model includes parameters representing the reconfigurable hardware, application load imbalance across the nodes, background user load, basic message-passing communication, and processor heterogeneity. In this article, we provide an overview of the analytical model and demonstrate its application for optimization and scheduling of high-performance reconfigurable computing (HPRC) resources. We examine cost functions for minimum runtime and other optimization problems commonly found in shared computing resources. Finally, we discuss additional scheduling issues and other potential applications of the model.
Melissa C. Smith, Gregory D. Peterson
ACM Trans. Embed. Comput. Syst.2
2012 An Effective Execution Time Approximation Method for Parallel Computing
abstract
In performance modeling of parallel synchronous iterative applications, the longest individual execution time among parallel processors determines the iteration time and often must be estimated for performance analysis. This involves the mean maximum calculation which has been a challenge in computer modeling for a long time. For large systems, numerical methods are not suitable because of heavy computation requirements and inaccuracy caused by rounding. On the other hand, previous approximation methods face challenges of accuracy and generality, especially for heterogeneous computing environments. This paper presents an interesting property of extreme values to enable Effective Mean Maximum Approximation (EMMA). Compared to previous mean maximum execution time approximation methods, this method is more accurate and general to different computational environments.
Junqing Sun, Gregory D. Peterson
IEEE Trans. Parallel Distributed Syst.2
2011 Comparing Hardware Accelerators in Scientific Applications: A Case Study
abstract
Multicore processors and a variety of accelerators have allowed scientific applications to scale to larger problem sizes. We present a performance, design methodology, platform, and architectural comparison of several application accelerators executing a Quantum Monte Carlo application. We compare the application's performance and programmability on a variety of platforms including CUDA with Nvidia GPUs, Brook+ with ATI graphics accelerators, OpenCL running on both multicore and graphics processors, C++ running on multicore processors, and a VHDL implementation running on a Xilinx FPGA. We show that OpenCL provides application portability between multicore processors and GPUs, but may incur a performance cost. Furthermore, we illustrate that graphics accelerators can make simulations involving large numbers of particles feasible.
Rick Weber, Akila Gothandaraman, Robert J. Hinde, Gregory D. Peterson
IEEE Trans. Parallel Distributed Syst.4
2010 Blocking LU Decomposition for FPGAs
abstract
To efficiently perform large matrix LU decomposition on FPGAs with limited local memory, the original algorithm needs to be blocked. In this paper, we propose a block LU decomposition algorithm for FPGAs, which is applicable for matrices of arbitrary size. We introduce a high performance hardware design, which mainly consists of a linear array of processing elements (PEs), to implement our block LU decomposition algorithm. A total of 36 PEs can be integrated into a Xilinx Virtex-5 xc5vlx330 FPGA on our self-designed PCI-Express card, reaching a sustained performance of 8.50 GFLOPS at 133 MHz, which outperforms previous work.
Guiming Wu, Yong Dou, Gregory D. Peterson
FCCM3
2010 Space-Time Turbo Bayesian Compressed Sensing for UWB Systems
abstract
A Space-Time Turbo Bayesian Compressed Sensing (STTBCS) algorithm is proposed for Ultra-Wideband (UWB) systems in this paper. Based on the sparsity of UWB signals, the STTBCS algorithm provides an efficient approach for integrating spatial and temporal redundancies. A space-time structure is also designed for exploiting and transferring the spatial and temporal \emph{a priori} information for signal reconstructions in the framework of Bayesian Compressed Sensing (BCS). Simulation results using experimental UWB echo signals demonstrate that our STTBCS algorithm achieves good performance for UWB systems, compared with the traditional BCS and multitask BCS algorithms.
Depeng Yang, Husheng Li, Gregory D. Peterson
ICC3
2009 An FPGA Implementation for Solving Least Square Problem
abstract
This paper proposes a high performance least square solver on FPGAs using the Cholesky decomposition method. Our design can be realized by iteratively adopting a single triangular linear equation solver for modified Cholesky decomposition and forward/backward substitutions. Good performance is achieved by optimizing the Cholesky factorization algorithms, reordering the computation and thus alleviating the data dependency. Dedicated hardware architecture for solving triangular linear equations is designed and implemented for different precision requirements. Compared to software on a Pentium 4, our design achieves a significant speedup.
Depeng Yang, Gregory D. Peterson, Husheng Li, Junqing Sun
FCCM2
2009 Feedback orthogonal pruning pursuit for pulse acquisition in UWB communications
abstract
A feedback structure and signal recovery algorithm, orthogonal pruning pursuit, are proposed to apply sequential compressed sensing (CS) on ultra-wideband (UWB) communications. The prior knowledge is exploited by the feedback structure in time sequence UWB signals. The orthogonal pruning pursuit algorithm is utilized to reconstruct UWB echo signals by deleting zero entries to pursuit true non-zero elements. Numerical simulations show that using the proposed scheme, far fewer measurements are needed to achieve good performance compared with other CS algorithms for UWB communications. This saves substantial hardware resources in practical systems.
Depeng Yang, Husheng Li, Gregory D. Peterson
PIMRC3
2008 FPGA acceleration of a quantum Monte Carlo application
Akila Gothandaraman, Gregory D. Peterson, G. Lee Warren, Robert J. Hinde, Robert J. Harrison
Parallel Comput.2
2008 High-Performance Mixed-Precision Linear Solver for FPGAs
abstract
Compared to higher-precision data formats, lower-precision data formats result in higher performance for computational intensive applications on FPGAs because of their lower resource cost, reduced memory bandwidth requirements, and higher circuit frequency. On the other hand, scientific computations usually demand highly accurate solutions. This paper seeks to utilize lower-precision data formats whenever possible for higher performance without losing the accuracy of higher-precision data formats by using mixed-precision algorithms and architectures. First, we analyze the floating-point performance of different data formats on FPGAs. Second, we introduce mixed-precision iterative refinement algorithms for linear solvers and give error analysis. Finally, we propose an innovative architecture for a mixed-precision direct solver for reconfigurable computing. Our results show that our mixed-precision algorithm and architecture significantly improve the performance of linear solvers on FPGAs.
Junqing Sun, Gregory D. Peterson, Olaf O. Storaasli
IEEE Trans. Computers2
2007 Sparse Matrix-Vector Multiplication Design on FPGAs
abstract
Creating a high throughput sparse matrix vector multiplication (SpMxV) implementation depends on a balanced system design. In this paper, we introduce the innovative SpMxV solver designed for FPGAs (SSF). Besides high computational throughput, system performance is optimized by reducing initialization time and overheads, minimizing and overlapping I/O operations, and increasing scalability. SSF accepts any matrix size and can be easily adapted to different data formats. SSF minimizes the control logic by taking advantage of the data flow via an innovative accumulation circuit which uses pipelined floating point adders. Compared to optimized software codes on a Pentium 4 microprocessor, our design achieves up to 20x speedup.
Junqing Sun, Gregory D. Peterson, Olaf O. Storaasli
FCCM2
2006 FPGA Implementation of Evolvable Block-based Neural Networks
abstract
This paper presents a hardware implementation approach for Block-based Neural Networks (BbNNs) on a Programmable System-On-Chip. This is an intrinsic online evolution system that can be genetically evolved and adapted to changes in input data patterns dynamically without any need for multiple FPGA reconfigurations to accommodate various network structure/parameter changes. This removes a considerable bottleneck for performance. The research presented here is a first step towards an evolvable system that can be implemented as an embedded system.
Saumil G. Merchant, Gregory D. Peterson, Sang Ki Park, Seong G. Kong
IEEE Congress on Evolutionary Computation2
2006 Continuous Heartbeat Monitoring Using Evolvable Block-based Neural Networks
abstract
This paper presents continuous heartbeat monitoring using evolvable block-based neural networks (BbNNs). An evolutionary algorithm is used to optimize the structure and weights of BbNN simultaneously. A BbNN, trained with the Hermite transform coefficients and a time interval between the two neighboring R peaks of ECG signal, promises a patient-specific heartbeat monitoring system. BbNNs reconfigure the structure and internal weights to cope with individual differences and the changes in physical conditions. Simulation results using the MIT-BIH Arrhythmia database demonstrate a high accuracy of 98.7% on average for the classification of ventricular ectopic beats (VEBs), being a substantial improvement over conventional techniques.
Seong G. Kong, Gregory D. Peterson
IJCNN3
2006 Poster reception - A reconfigurable supercomputing library for accelerated parallel lagged-Fibonacci pseudorandom number generation
abstract
To help promote more widespread adoption of hardware acceleration in parallel scientific computing we present portable, flexible design components for pseudorandom number generation. Due to the success of the Scalable Parallel Random Number Generators (SPRNG) software library in stochastic computations (e.g., Monte Carlo simulations), we developed an efficient and portable hardware architecture fully compatible with SPRNG's Parallel Additive Lagged Fibonacci Generator (PALFG). Our general design produces identical results for all the parameter sets that SPRNG supports and yields high performance parallel random number generators, which can each generate 162 million 31 bit uniform random integers per second on Xilinx Virtex II Pro FPGAs. The friendly design interface makes it easy for users to integrate into their applications, particularly computational scientists unfamiliar with reconfigurable hardware. Due to its fast generation speed and friendly interface, this uniform random number generator is being targeted as an open core for parallel scientific computing.
Yu Bi, Gregory D. Peterson, G. Lee Warren, Robert J. Harrison
SC2
2006 Poster reception - Reconfigurable accelerator for quantum Monte Carlo simulations in N-body systems
abstract
Recent advances in FPGA technology make them an attractive platform for accelerating scientific computing applications. We present a novel hardware accelerator for Quantum Monte Carlo simulations in N-body systems. The design is deeply pipelined and exploits the inherent fine-grained parallelism available using an FPGA for all calculations. The design is implemented on a Xilinx Virtex II Pro XC2VP30 device and preliminary results indicate a maximum operating frequency of 100MHz. A single instance of our design offers an estimated speedup of 20x and accuracy comparable to the serial code running on a 2.8GHz Intel Pentium 4 processor. This architecture performs all computations with fixed-point representation and delivers accuracy on the order of or better than double-precision floating point. After deploying a single instance on the present FPGA platform, targeting our design on the Cray XD1 platform with a high gate-density FPGA will allow us to operate multiple cores in parallel.
Akila Gothandaraman, G. Lee Warren, Gregory D. Peterson, Robert J. Harrison
SC3
2005 ECG signal classification using block-based neural networks
abstract
This paper investigates the application of evolvable block-based neural networks (BbNNs) to ECG signal classification. A BbNN consists of a two-dimensional (2-D) array of modular basic blocks that can be easily implemented using reconfigurable digital hardware. BbNNs are evolved for each patient in order to provide personalized health monitoring. A genetic algorithm evolves the internal structure and associated weights of a BbNN using training patterns that consist of morphological and temporal features extracted from the ECG signal of a patient. The remaining part of the ECG record serves as the test signal. The BbNN was tested for ten records collected from different patients provided by the MIT-BIH Arrhythmia database. The evolved BbNNs produced higher than 90% classification accuracies.
Seong G. Kong, Gregory D. Peterson
IJCNN3
2005 Parallel application performance on shared high performance reconfigurable computing resources
Melissa C. Smith, Gregory D. Peterson
Perform. Evaluation2
2004 Engineering in the biological substrate: information processing in genetic circuits
abstract
We review the rapidly evolving efforts to analyze, model, simulate, and engineer genetic and biochemical information processing systems within living cells. We begin by showing that the fundamental elements of information processing in electronic and genetic systems are strikingly similar, and follow this theme through a review of efforts to create synthetic genetic circuits. In particular, we describe and review the "silicon mimetic" approach, where genetic circuits are engineered to mimic the functionality of semiconductor devices such as logic gates, latched circuits, and oscillators. This is followed with a review of the analysis, modeling, and simulation of natural and synthetic genetic circuits, which often proceed in a manner similar to that used for electronic systems. We conclude by presenting examples of naturally occurring genetic and biochemical systems that recently have been conceptualized in terms familiar to systems engineers. Our review of these newly forming fields of research demonstrates that the expertise and skills contained within electrical and computer engineering disciplines apply not only to design within biological systems, but also to the development of a deeper understanding of biological functionality. This review of these efforts points to the emergence of both engineering and basic science disciplines following parallel paths.
Michael L. Simpson, Chris D. Cox, Gregory D. Peterson, Gary S. Sayler
Proc. IEEE3
2003 Modeling mobile-agent-based collaborative processing in sensor networks using generalized stochastic Petri nets
abstract
In mobile-agent-based distributed sensor networks (MADSNs), instead of moving data from an individual sensor node to a processing center as a typical scenario in the client/server-based computing, mobile agents on that carry the executable code are dispatched from the processing center to the sensor nodes and process data locally. Because of the complicated behavior of the mobile agents, there has not been much work done in modeling and simulation. In this paper, a Generalized Stochastic Petri Net (GSPN) is used to model the mobile agents in DSNs. GSPN is a very popular modeling tool for systems that feature concurrency, synchronization and randomness. One of the most challenging problems in GSPN modeling is to design a mechanism for breaking the transition conflicts. This paper presents a random transition selector based on joint entropy and a so-called "rolling rocks random (R3) selector" to break conflicts among immediate transitions. The GSPN model based on this transition selector is synthesized on a Xilinx Virtex 1000E Field Programmable Gate Array (FPGA) using reconfigurable components. Simulation results show that the proposed transition selector performs better than the commonly used random selector.
Hongtao Du, Hairong Qi 0001, Gregory D. Peterson
SMC3
1993 Performance of a Globally-Clocked Parallel Simulator
abstract
A performance model for a globally-clocked, discrete-event queueing network simulator is developed and validated against measured results. The use of architectural enhancements for improving the performance of the algorithm is investigated. Both scaled and fixed problem sizes are investigated, with a maximum measured scaled speedup of 47 and 64 processors. The model is very accurate, predicting runtime within 5% of measured results.
Gregory D. Peterson, Roger D. Chamberlain
ICPP (3)1