Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Robert H. Klenke

dblp:74/3544 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
0since 2021 · last 2011
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 5 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Performance modeling and evaluation · 39% Reconfigurable computing and FPGAs · 21% Electronic design automation · 14%

Topics — the 20 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation › tracing
instruction tracing
0.112011
Microblaze: an application-independent fpga-based profiler (abstract only) · FPGA 2011
Performance modeling and evaluation
profiling
0.112011
Microblaze: an application-independent fpga-based profiler (abstract only) · FPGA 2011
Memory systems
DRAM
0.122000
Dynamic Access Ordering for Streamed Computations · IEEE Trans. Computers 2000
Access Order and Effective Bandwidth for Streams on a Direct Rambus Memory · HPCA 1999
Reconfigurable computing and FPGAs › FPGA-based processor implementation
soft-core processor
0.012011
Microblaze: an application-independent fpga-based profiler (abstract only) · FPGA 2011
Interconnection networks and networks-on-chip › switch architecture
crossbar switch
0.012001
Performance Modeling of Hierarchical Crossbar-Based Multicomputer Systems · IEEE Trans. Computers 2001
Performance modeling and evaluation › network performance analysis
interconnection network performance
0.012001
Performance Modeling of Hierarchical Crossbar-Based Multicomputer Systems · IEEE Trans. Computers 2001
Processor architecture and microarchitecture › memory system microarchitecture
memory ordering
0.012000
Dynamic Access Ordering for Streamed Computations · IEEE Trans. Computers 2000
Memory systems
memory bandwidth
0.011999
Access Order and Effective Bandwidth for Streams on a Direct Rambus Memory · HPCA 1999
Electronic design automation › hardware simulation
multilevel simulation
0.011999
Resolving unknown inputs in mixed-level simulation with sequential elements · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1999
High-performance computing › data-intensive computing
streaming computation
0.011999
Access Order and Effective Bandwidth for Streams on a Direct Rambus Memory · HPCA 1999
Electronic design automation
hardware verification and test
0.021999
An analysis of fault partitioned parallel test generation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1996
Resolving unknown inputs in mixed-level simulation with sequential elements · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1999
Electronic design automation
high-level synthesis
0.011998
A Top-Down Design Environment for Developing Pipelined Datapaths · DAC 1998
Integrated circuit design
digital system design
0.011997
An Integrated Design Environment for Performance and Dependability Analysis · DAC 1997
Parallel and multicore computing
parallel algorithms
0.011996
An analysis of fault partitioned parallel test generation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1996
Electronic design automation › hardware verification and test › test generation
parallel test generation
0.011996
An analysis of fault partitioned parallel test generation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1996
Electronic design automation › hardware verification and test
test generation
0.011996
An analysis of fault partitioned parallel test generation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1996
Memory systems › memory bandwidth
memory bandwidth optimization
0.012000
Dynamic Access Ordering for Streamed Computations · IEEE Trans. Computers 2000
Electronic design automation › hardware/software co-design
co-simulation
0.011999
Resolving unknown inputs in mixed-level simulation with sequential elements · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1999
Performance modeling and evaluation › performance model construction › memory system performance modeling
memory bottleneck analysis
0.011999
Access Order and Effective Bandwidth for Streams on a Direct Rambus Memory · HPCA 1999
Processor architecture and microarchitecture › microprocessor design › processor core design
datapath design
0.011998
A Top-Down Design Environment for Developing Pipelined Datapaths · DAC 1998

Methods — techniques the papers use, named apart from their topics

instruction flow tracing · 0.1analytical modeling · 0.0simulation · 0.0stream memory controller · 0.0hardware prefetching · 0.0compile-time stream detection · 0.0streaming hardware · 0.0mixed-level modeling · 0.0cacheline burst analysis · 0.0cycle-based modeling · 0.0
YearPublicationVenuePosition
2011 Microblaze: an application-independent fpga-based profiler (abstract only)
abstract
Monitoring the functional behavior of an application is an important capability that assists in exploring the performance of a target application against different SW/HW implementations. Recently, there have been efforts to exploit the ability to trace the internal signals of the soft-core processors for developing FPGA-based profiling tools to monitor programs running on these processors. However, these previously developed techniques are application-dependent, i.e., they require the designer either to edit the HDL code or the application code to obtain the desired trace information when targeting new applications. In this research, we propose an application-independent profiling technique using the MicroBlaze/FPGA platform where profiling library or user-defined functions can be achieved by tracing the unique instruction flow that distinguishes functions from each other rather than monitoring the program counter value (the addresses of the functions). Hence, modifying the application code or targeting new application does not require reconfiguring the FPGA or modifying the application code for targeting the same functions. This technique can be used to analyze the target application at the source code level, observing the dominant operations and demanded resources that characterize the system behavior. In addition, this technique can assist in selecting the appropriate processor architecture for a given application by considering MicroBlaze as a reference architecture from which the functional behavior of the target application can be mapped to the performance of other architectures.
Fadi Obeidat, Robert H. Klenke
FPGA2
2001 Interfaces for mixed-level simulation with sequential elements
Robert H. Klenke, James H. Aylor, Moshe Meyassed, William W. Dungan
J. Syst. Archit.1
2001 Performance Modeling of Hierarchical Crossbar-Based Multicomputer Systems
abstract
Crossbar networks have been widely used as the interconnection mechanism in many multiprocessor/multicomputer systems. Crossbar networks can be categorized into three major topological classes: full-crossbar networks, multistage interconnection networks (MINs), and networks consisting of multiple levels of full crossbar connections, called hierarchical crossbar interconnection networks (HCINs). A significant amount of previous work exists in the area of performance modeling of systems with full-crossbar networks or multistage interconnection networks. However, performance modeling of multicomputer systems with HCINs has not been widely studied. This paper presents both analytical and simulation models for performance evaluation of an HCIN based on the commercial Mercury RACEway crossbar switch. The effective data transfer rate for message passing is taken as the primary performance metric and the models predict how this metric varies with the traffic load on the system. The analytical results are compared to the simulation results for different standard configurations of Mercury RACE Multicomputer Systems.
Robert H. Klenke, James H. Aylor
IEEE Trans. Computers2
2000 Dynamic Access Ordering for Streamed Computations
abstract
Memory bandwidth is rapidly becoming the limiting performance factor for many applications, particularly for streaming computations such as scientific vector processing or multimedia (de)compression. Although these computations lack the temporal locality of reference that makes traditional caching schemes effective, they have predictable access patterns. Since most modern DRAM components support modes that make it possible to perform some access sequences faster than others, the predictability of the stream accesses makes it possible to reorder them to get better memory performance. We describe a Stream Memory Controller (SMC) system that combines compile-time detection of streams with execution-time selection of the access order and issue. The SMC effectively prefetches read-streams, buffers write-streams, and reorders the accesses to exploit the existing memory bandwidth as much as possible. Unlike most other hardware prefetching or stream buffer designs, this system does not increase bandwidth requirements. The SMC is practical to implement, using existing compiler technology and requiring only a modest amount of special purpose hardware. We present simulation results for fast-page mode and Rambus DRAM memory systems and we describe a prototype system with which we have observed performance improvements for inner loops by factors of 13 over traditional access methods.
Sally A. McKee, William A. Wulf, James H. Aylor, Robert H. Klenke, Maximo H. Salinas, Sung I. Hong, Dee A. B. Weikle
IEEE Trans. Computers4
1999 Access Order and Effective Bandwidth for Streams on a Direct Rambus Memory
abstract
Processor speeds are increasing rapidly and memory speeds are not keeping up. Streaming computations (such as multimedia or scientific applications) are among those whose performance is most limited by the memory bottleneck. Rambus hopes to bridge the processor/memory performance gap with a recently introduced DRAM that can deliver up to 1.6 Gbytes/sec. We analyze the performance of these interesting new memory devices on the inner loops of streaming computations, both for traditional memory controllers that treat all DRAM transactions as random cacheline accesses, and for controllers augmented with streaming hardware. For our benchmarks, we find that accessing unit-stride streams in cacheline bursts in the natural order of the computation exploits from 44-76% of the peak bandwidth of a memory system composed of a single Direct RDRAM device, and that accessing streams via a streaming mechanism with a simple access ordering scheme can improve performance by factors of 1.18 to 2.25.
Sung I. Hong, Sally A. McKee, Maximo H. Salinas, Robert H. Klenke, James H. Aylor, William A. Wulf
HPCA4
1999 Resolving unknown inputs in mixed-level simulation with sequential elements
abstract
It is well known that techniques such as performance modeling that can effectively evaluate design alternatives early in the design process can greatly increase the quality of the ultimate implementation, while at the same time, decrease the design time. In order to gain the maximum benefit from performance modeling, it must be integrated into the design process such that the performance model can be directly refined into an implementation. This paper presents techniques for developing interfaces between abstract performance models and detailed behavioral models to enable this refinement process. These mixed-level modeling interfaces, as they are called, allow abstract performance models to be cosimulated with detailed behavioral models. Because of the differences in the level of detail between abstract performance models and detailed behavioral models, not all of the inputs to the behavioral model can be derived from information in the performance model. Techniques for determining values for these "unknown" inputs are presented, These techniques have been designed to generate bounds on the system performance that converge as the model is refined.
Moshe Meyassed, Robert H. Klenke, James H. Aylor
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1998 A Top-Down Design Environment for Developing Pipelined Datapaths
abstract
This paper presents a design environment for cycle-based systems, such as microprocessors, that permits modeling of these systems at various levels, from the abstract system level, through the detailed RTL level, to an actual implementation. The environment allows the models to be refined to lower levels in a step-wise manner. The environment provides the ability to obtain meaningful metrics from abstract models of a processor's architecture. This capability allows design alternatives to be evaluated earlier in the design cycle, thus eliminating costly redesign and reducing the processor time to market.
Robert M. McGraw, James H. Aylor, Robert H. Klenke
DAC3
1997 An Integrated Design Environment for Performance and Dependability Analysis
abstract
This paper presents an integrated design environment thatsupports the design and analysis of digital systems from initialconcept to the final implementation. The environment supports bothsystem level performance and dependability analysis from acommon modeling representation. A tool called ADEPT (AdvancedDesign Environment Prototype Tool) has been developed toimplement the environment. ADEPT is based on IEEE 1076 VHDLand uses commercial schematic capture systems as a front end viaan EDIF interface. Several examples are presented whichdemonstrate various aspects of the environment.
Robert H. Klenke, Moshe Meyassed, James H. Aylor, Barry W. Johnson, Ramesh Rao, Anup Ghosh
DAC1
1996 Design and Evaluation of Dynamic Access Ordering Hardware
abstract
Memory bandwidth is rapidly becoming the limiting performance factor for many applications, particularly for streaming computations such as scientific vector processing or multimedia (de)compression. Although these computations lack the temporal locality of reference that makes caches effective, they have predictable access patterns. Since most modern DRAM components support modes that make it possible to perform some access sequences faster than others, the predictability of the stream accesses makes it possible to reorder them to get better memory performance. We describe and evaluate a Stream Memory Controller system that combines compile-time detection of streams with execution-time selection of the access order and issue. The technique is practical to implement, using existing compiler technology and requiring only a modest amount of special-purpose hardware. With our prototype system, we have observed performance improvements by factors of 13 over normal caching. 1. INTRODUCTION...
Sally A. McKee, Assaji Aluwihare, Benjamin H. Clark, Robert H. Klenke, Trevor C. Landon, Christopher W. Oliver, Maximo H. Salinas, Adam E. Szymkowiak, Kenneth L. Wright, William A. Wulf, James H. Aylor
International Conference on Supercomputing4
1996 An analysis of fault partitioning algorithms for fault partitioned ATPG
abstract
Generation of test vectors for the VLSI devices used in contemporary digital systems is becoming much more difficult as these devices increase in size and complexity. Automatic Test Pattern Generation (ATPG) techniques are commonly used to generate these tests. Since ATPG is an NP complete problem with complexity exponential to circuit size, the application of parallel processing techniques to accelerate the process of generating test vectors is an promising area of research. The simplest approach to parallelization of the test generation process is to simply divide the processing of the fault list across multiple processors. Each individual processor then performs the normal rest generation process on its own portion of the fault list, typically without interaction with the other processors. The major drawback of this technique, called fault partitioning, is that the processors perform redundant work generating test vectors for faults covered by vectors generated on another processor. This problem has been solved with the introduction of dynamic load balancing and detected fault broadcasting. Previous research has indicated that algorithmic fault partitioning moderately improves the performance of fault partitioned ATPG without detected fault broadcasting by reducing redundant work. However algorithmic fault partitioning can add significant preprocessing time to the ATPG process. This paper presents results that show that algorithmic partitioning is unnecessary prior to fault partitioned parallel ATPG using detected fault broadcasting and dynamic load balancing. Considering preprocessing time, random fault partitioning is shown to be the most efficient technique for partitioning faults prior to fault partitioned ATPG.
Robert H. Klenke, James H. Aylor, Joseph M. Wolf
VTS1
1996 An analysis of fault partitioned parallel test generation
abstract
Generation of test vectors for the VLSI devices used in contemporary digital systems is becoming much more difficult as these devices increase in size and complexity. Automatic Test Pattern Generation (ATPG) techniques are commonly used to generate these tests. Since ATPG is an NP complete problem with complexity exponential to circuit size, the application of parallel processing techniques to accelerate the process of generating test vectors is an active area of research. The simplest approach to parallelization of the test generation process is to simply divide the processing of the fault list across multiple processors, Each individual processor then performs the normal test generation process on its own portion of the fault list, typically without interaction with the other processors. The major drawback of this technique, called fault partitioning, is that the processors perform redundant work generating test vectors for faults covered by vectors generated on another processor. An earlier approach to reducing this redundant work involved transmitting generated test vectors among the processors and fault simulating them on each processor. This paper presents a comparison of the vector broadcasting approach with the simpler and more effective approach of fault broadcasting. In fault broadcasting, fault simulation is performed on the entire fault list on each processor. The resulting list of detected faults is then transmitted to all the other processors. The results show that this technique produces greater speedups and smaller test sets than the test vector broadcasting technique. Analytical models are developed which can be used to determine the cost of the various parts of the parallel ATPG algorithm. These models are validated using data from benchmark circuits.
Joseph M. Wolf, Lori M. Kaufman, Robert H. Klenke, James H. Aylor, Ronald Waxman
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1995 Refinement of system-level designs using hybrid modeling
abstract
The design of complex systems requires an enormous amount of modeling and simulation at the various levels of detail to assure high quality results. Most design methodologies require the use of separate modeling languages and simulation environments for each refinement step and simulation technique. Since one of the major goals of design is to reduce the time to market for a given system, the current approaches are becoming unacceptable. A new unified modeling (UM) methodology is under development at the University of Virginia which provides a top-down/bottom-up design methodology based on a single modeling environment. This single environment methodology is made possible through a technique known as hybrid modeling, which enable the simulation and analysis of behavioral components within a system-level model. By simulating behavioral components along with components at the system modeling level, high risk portions of the design can be examined at earlier stages in a given product's development; thus reducing the time to market. The paper examines the hybrid modeling techniques which make this multi-level modeling possible.
Robert M. McGraw, Moshe Meyassed, Robert H. Klenke, James H. Aylor, Ronald D. Williams
ICECCS3
1993 Workstation Based Parallel Test Generation
abstract
Generation of test vectors for the VLSI devices used in contemporary digital system is becoming much more difficult as these devices increase in size. Automatic Test Pattern Generation (ATPG) techniques are commonly used to generate these tests. Since ATPG is an NP complete problem with complexity exponential to circuit size, the application of parallel processing techniques to accelerate the process of finding test patterns is an active area of research. This paper presents an approach to parallelization of the test generation problem that is targeted to a network-of-workstations environment. The system is based upon partitioning of the fault list across multiple processors and includes enhancements designed to address the main drawbacks of this technique, namely unequal load balancing and generation of redundant vectors. The technique is generalized enough that it can be applied to any test generation system regardless of the ATPG or fault simulation algorithm employed. Results were gathered to determine the impact of workstation processing load and network communications load on the performance of the system.>
Robert H. Klenke, Lori M. Kaufman, James H. Aylor, Ronald Waxman, Padmini Narayan
ITC1
1993 Parallelization methods for circuit partitioning based parallel automatic test pattern generation
abstract
Generation of test vectors for the VLSI devices used in contemporary digital systems is becoming much more difficult as these devices increase in size. Automatic Test Pattern Generation (ATPG) techniques are commonly used to generate these tests. Parallel processing techniques can be applied to accelerate the process of finding test patterns. One problem with this approach is that most currently available distributed memory multicomputers have a limited amount of memory on each processor which limits the size of the circuit database that can be contained on a single node. Topological partitioning of the circuit database across several processors can increase the size of VLSI circuits that can be processed on a given parallel machine. This paper presents the architecture of a topologically partitioned ATPG system and several partitioning algorithms that can be used to partition the circuit-under-test. This paper also presents several parallelization methods that may be applied to topologically partitioned ATPG on a distributed memory multicomputer. Results of using these parallelization techniques along with topological partitioning are presented.>
Robert H. Klenke, Ronald D. Williams, James H. Aylor
VTS1