VLDB 2026 Research / reviewers in the wild / expert
Robert H. Klenke
dblp:74/3544
· DBLP profile ↗
14ranked-venue papers
5as first author
0since 2021 · last 2011
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 5 first-authorSoftware engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Performance modeling and evaluation · 39% Reconfigurable computing and FPGAs · 21% Electronic design automation · 14% |
Topics — the 20 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation › tracing
instruction tracing |
0.1 | 1 | 2011 | Microblaze: an application-independent fpga-based profiler (abstract only) · FPGA 2011 |
Performance modeling and evaluation
profiling |
0.1 | 1 | 2011 | Microblaze: an application-independent fpga-based profiler (abstract only) · FPGA 2011 |
Memory systems
DRAM |
0.1 | 2 | 2000 | Dynamic Access Ordering for Streamed Computations · IEEE Trans. Computers 2000 Access Order and Effective Bandwidth for Streams on a Direct Rambus Memory · HPCA 1999 |
Reconfigurable computing and FPGAs › FPGA-based processor implementation
soft-core processor |
0.0 | 1 | 2011 | Microblaze: an application-independent fpga-based profiler (abstract only) · FPGA 2011 |
Interconnection networks and networks-on-chip › switch architecture
crossbar switch |
0.0 | 1 | 2001 | Performance Modeling of Hierarchical Crossbar-Based Multicomputer Systems · IEEE Trans. Computers 2001 |
Performance modeling and evaluation › network performance analysis
interconnection network performance |
0.0 | 1 | 2001 | Performance Modeling of Hierarchical Crossbar-Based Multicomputer Systems · IEEE Trans. Computers 2001 |
Processor architecture and microarchitecture › memory system microarchitecture
memory ordering |
0.0 | 1 | 2000 | Dynamic Access Ordering for Streamed Computations · IEEE Trans. Computers 2000 |
Memory systems
memory bandwidth |
0.0 | 1 | 1999 | Access Order and Effective Bandwidth for Streams on a Direct Rambus Memory · HPCA 1999 |
Electronic design automation › hardware simulation
multilevel simulation |
0.0 | 1 | 1999 | Resolving unknown inputs in mixed-level simulation with sequential elements · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1999 |
High-performance computing › data-intensive computing
streaming computation |
0.0 | 1 | 1999 | Access Order and Effective Bandwidth for Streams on a Direct Rambus Memory · HPCA 1999 |
Electronic design automation
hardware verification and test |
0.0 | 2 | 1999 | An analysis of fault partitioned parallel test generation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1996 Resolving unknown inputs in mixed-level simulation with sequential elements · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1999 |
Electronic design automation
high-level synthesis |
0.0 | 1 | 1998 | A Top-Down Design Environment for Developing Pipelined Datapaths · DAC 1998 |
Integrated circuit design
digital system design |
0.0 | 1 | 1997 | An Integrated Design Environment for Performance and Dependability Analysis · DAC 1997 |
Parallel and multicore computing
parallel algorithms |
0.0 | 1 | 1996 | An analysis of fault partitioned parallel test generation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1996 |
Electronic design automation › hardware verification and test › test generation
parallel test generation |
0.0 | 1 | 1996 | An analysis of fault partitioned parallel test generation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1996 |
Electronic design automation › hardware verification and test
test generation |
0.0 | 1 | 1996 | An analysis of fault partitioned parallel test generation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1996 |
Memory systems › memory bandwidth
memory bandwidth optimization |
0.0 | 1 | 2000 | Dynamic Access Ordering for Streamed Computations · IEEE Trans. Computers 2000 |
Electronic design automation › hardware/software co-design
co-simulation |
0.0 | 1 | 1999 | Resolving unknown inputs in mixed-level simulation with sequential elements · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1999 |
Performance modeling and evaluation › performance model construction › memory system performance modeling
memory bottleneck analysis |
0.0 | 1 | 1999 | Access Order and Effective Bandwidth for Streams on a Direct Rambus Memory · HPCA 1999 |
Processor architecture and microarchitecture › microprocessor design › processor core design
datapath design |
0.0 | 1 | 1998 | A Top-Down Design Environment for Developing Pipelined Datapaths · DAC 1998 |
Methods — techniques the papers use, named apart from their topics
instruction flow tracing · 0.1analytical modeling · 0.0simulation · 0.0stream memory controller · 0.0hardware prefetching · 0.0compile-time stream detection · 0.0streaming hardware · 0.0mixed-level modeling · 0.0cacheline burst analysis · 0.0cycle-based modeling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2011 | Microblaze: an application-independent fpga-based profiler (abstract only)abstractMonitoring the functional behavior of an application is an important capability that assists in exploring the performance of a target application against different SW/HW implementations. Recently, there have been efforts to exploit the ability to trace the internal signals of the soft-core processors for developing FPGA-based profiling tools to monitor programs running on these processors. However, these previously developed techniques are application-dependent, i.e., they require the designer either to edit the HDL code or the application code to obtain the desired trace information when targeting new applications. In this research, we propose an application-independent profiling technique using the MicroBlaze/FPGA platform where profiling library or user-defined functions can be achieved by tracing the unique instruction flow that distinguishes functions from each other rather than monitoring the program counter value (the addresses of the functions). Hence, modifying the application code or targeting new application does not require reconfiguring the FPGA or modifying the application code for targeting the same functions. This technique can be used to analyze the target application at the source code level, observing the dominant operations and demanded resources that characterize the system behavior. In addition, this technique can assist in selecting the appropriate processor architecture for a given application by considering MicroBlaze as a reference architecture from which the functional behavior of the target application can be mapped to the performance of other architectures. Fadi Obeidat, Robert H. Klenke |
FPGA | 2 |
| 2001 | Interfaces for mixed-level simulation with sequential elements
Robert H. Klenke, James H. Aylor, Moshe Meyassed, William W. Dungan |
J. Syst. Archit. | 1 |
| 2001 | Performance Modeling of Hierarchical Crossbar-Based Multicomputer SystemsabstractCrossbar networks have been widely used as the interconnection mechanism in many multiprocessor/multicomputer systems. Crossbar networks can be categorized into three major topological classes: full-crossbar networks, multistage interconnection networks (MINs), and networks consisting of multiple levels of full crossbar connections, called hierarchical crossbar interconnection networks (HCINs). A significant amount of previous work exists in the area of performance modeling of systems with full-crossbar networks or multistage interconnection networks. However, performance modeling of multicomputer systems with HCINs has not been widely studied. This paper presents both analytical and simulation models for performance evaluation of an HCIN based on the commercial Mercury RACEway crossbar switch. The effective data transfer rate for message passing is taken as the primary performance metric and the models predict how this metric varies with the traffic load on the system. The analytical results are compared to the simulation results for different standard configurations of Mercury RACE Multicomputer Systems. Robert H. Klenke, James H. Aylor |
IEEE Trans. Computers | 2 |
| 2000 | Dynamic Access Ordering for Streamed ComputationsabstractMemory bandwidth is rapidly becoming the limiting performance factor for many applications, particularly for streaming computations such as scientific vector processing or multimedia (de)compression. Although these computations lack the temporal locality of reference that makes traditional caching schemes effective, they have predictable access patterns. Since most modern DRAM components support modes that make it possible to perform some access sequences faster than others, the predictability of the stream accesses makes it possible to reorder them to get better memory performance. We describe a Stream Memory Controller (SMC) system that combines compile-time detection of streams with execution-time selection of the access order and issue. The SMC effectively prefetches read-streams, buffers write-streams, and reorders the accesses to exploit the existing memory bandwidth as much as possible. Unlike most other hardware prefetching or stream buffer designs, this system does not increase bandwidth requirements. The SMC is practical to implement, using existing compiler technology and requiring only a modest amount of special purpose hardware. We present simulation results for fast-page mode and Rambus DRAM memory systems and we describe a prototype system with which we have observed performance improvements for inner loops by factors of 13 over traditional access methods. Sally A. McKee, William A. Wulf, James H. Aylor, Robert H. Klenke, Maximo H. Salinas, Sung I. Hong, Dee A. B. Weikle |
IEEE Trans. Computers | 4 |
| 1999 | Access Order and Effective Bandwidth for Streams on a Direct Rambus MemoryabstractProcessor speeds are increasing rapidly and memory speeds are not keeping up. Streaming computations (such as multimedia or scientific applications) are among those whose performance is most limited by the memory bottleneck. Rambus hopes to bridge the processor/memory performance gap with a recently introduced DRAM that can deliver up to 1.6 Gbytes/sec. We analyze the performance of these interesting new memory devices on the inner loops of streaming computations, both for traditional memory controllers that treat all DRAM transactions as random cacheline accesses, and for controllers augmented with streaming hardware. For our benchmarks, we find that accessing unit-stride streams in cacheline bursts in the natural order of the computation exploits from 44-76% of the peak bandwidth of a memory system composed of a single Direct RDRAM device, and that accessing streams via a streaming mechanism with a simple access ordering scheme can improve performance by factors of 1.18 to 2.25. Sung I. Hong, Sally A. McKee, Maximo H. Salinas, Robert H. Klenke, James H. Aylor, William A. Wulf |
HPCA | 4 |
| 1999 | Resolving unknown inputs in mixed-level simulation with sequential elementsabstractIt is well known that techniques such as performance modeling that can effectively evaluate design alternatives early in the design process can greatly increase the quality of the ultimate implementation, while at the same time, decrease the design time. In order to gain the maximum benefit from performance modeling, it must be integrated into the design process such that the performance model can be directly refined into an implementation. This paper presents techniques for developing interfaces between abstract performance models and detailed behavioral models to enable this refinement process. These mixed-level modeling interfaces, as they are called, allow abstract performance models to be cosimulated with detailed behavioral models. Because of the differences in the level of detail between abstract performance models and detailed behavioral models, not all of the inputs to the behavioral model can be derived from information in the performance model. Techniques for determining values for these "unknown" inputs are presented, These techniques have been designed to generate bounds on the system performance that converge as the model is refined. Moshe Meyassed, Robert H. Klenke, James H. Aylor |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1998 | A Top-Down Design Environment for Developing Pipelined DatapathsabstractThis paper presents a design environment for cycle-based systems, such as microprocessors, that permits modeling of these systems at various levels, from the abstract system level, through the detailed RTL level, to an actual implementation. The environment allows the models to be refined to lower levels in a step-wise manner. The environment provides the ability to obtain meaningful metrics from abstract models of a processor's architecture. This capability allows design alternatives to be evaluated earlier in the design cycle, thus eliminating costly redesign and reducing the processor time to market. Robert M. McGraw, James H. Aylor, Robert H. Klenke |
DAC | 3 |
| 1997 | An Integrated Design Environment for Performance and Dependability AnalysisabstractThis paper presents an integrated design environment thatsupports the design and analysis of digital systems from initialconcept to the final implementation. The environment supports bothsystem level performance and dependability analysis from acommon modeling representation. A tool called ADEPT (AdvancedDesign Environment Prototype Tool) has been developed toimplement the environment. ADEPT is based on IEEE 1076 VHDLand uses commercial schematic capture systems as a front end viaan EDIF interface. Several examples are presented whichdemonstrate various aspects of the environment. Robert H. Klenke, Moshe Meyassed, James H. Aylor, Barry W. Johnson, Ramesh Rao, Anup Ghosh |
DAC | 1 |
| 1996 | Design and Evaluation of Dynamic Access Ordering HardwareabstractMemory bandwidth is rapidly becoming the limiting performance factor for many applications, particularly for streaming computations such as scientific vector processing or multimedia (de)compression. Although these computations lack the temporal locality of reference that makes caches effective, they have predictable access patterns. Since most modern DRAM components support modes that make it possible to perform some access sequences faster than others, the predictability of the stream accesses makes it possible to reorder them to get better memory performance. We describe and evaluate a Stream Memory Controller system that combines compile-time detection of streams with execution-time selection of the access order and issue. The technique is practical to implement, using existing compiler technology and requiring only a modest amount of special-purpose hardware. With our prototype system, we have observed performance improvements by factors of 13 over normal caching. 1. INTRODUCTION... Sally A. McKee, Assaji Aluwihare, Benjamin H. Clark, Robert H. Klenke, Trevor C. Landon, Christopher W. Oliver, Maximo H. Salinas, Adam E. Szymkowiak, Kenneth L. Wright, William A. Wulf, James H. Aylor |
International Conference on Supercomputing | 4 |
| 1996 | An analysis of fault partitioning algorithms for fault partitioned ATPGabstractGeneration of test vectors for the VLSI devices used in contemporary digital systems is becoming much more difficult as these devices increase in size and complexity. Automatic Test Pattern Generation (ATPG) techniques are commonly used to generate these tests. Since ATPG is an NP complete problem with complexity exponential to circuit size, the application of parallel processing techniques to accelerate the process of generating test vectors is an promising area of research. The simplest approach to parallelization of the test generation process is to simply divide the processing of the fault list across multiple processors. Each individual processor then performs the normal rest generation process on its own portion of the fault list, typically without interaction with the other processors. The major drawback of this technique, called fault partitioning, is that the processors perform redundant work generating test vectors for faults covered by vectors generated on another processor. This problem has been solved with the introduction of dynamic load balancing and detected fault broadcasting. Previous research has indicated that algorithmic fault partitioning moderately improves the performance of fault partitioned ATPG without detected fault broadcasting by reducing redundant work. However algorithmic fault partitioning can add significant preprocessing time to the ATPG process. This paper presents results that show that algorithmic partitioning is unnecessary prior to fault partitioned parallel ATPG using detected fault broadcasting and dynamic load balancing. Considering preprocessing time, random fault partitioning is shown to be the most efficient technique for partitioning faults prior to fault partitioned ATPG. Robert H. Klenke, James H. Aylor, Joseph M. Wolf |
VTS | 1 |
| 1996 | An analysis of fault partitioned parallel test generationabstractGeneration of test vectors for the VLSI devices used in contemporary digital systems is becoming much more difficult as these devices increase in size and complexity. Automatic Test Pattern Generation (ATPG) techniques are commonly used to generate these tests. Since ATPG is an NP complete problem with complexity exponential to circuit size, the application of parallel processing techniques to accelerate the process of generating test vectors is an active area of research. The simplest approach to parallelization of the test generation process is to simply divide the processing of the fault list across multiple processors, Each individual processor then performs the normal test generation process on its own portion of the fault list, typically without interaction with the other processors. The major drawback of this technique, called fault partitioning, is that the processors perform redundant work generating test vectors for faults covered by vectors generated on another processor. An earlier approach to reducing this redundant work involved transmitting generated test vectors among the processors and fault simulating them on each processor. This paper presents a comparison of the vector broadcasting approach with the simpler and more effective approach of fault broadcasting. In fault broadcasting, fault simulation is performed on the entire fault list on each processor. The resulting list of detected faults is then transmitted to all the other processors. The results show that this technique produces greater speedups and smaller test sets than the test vector broadcasting technique. Analytical models are developed which can be used to determine the cost of the various parts of the parallel ATPG algorithm. These models are validated using data from benchmark circuits. Joseph M. Wolf, Lori M. Kaufman, Robert H. Klenke, James H. Aylor, Ronald Waxman |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1995 | Refinement of system-level designs using hybrid modelingabstractThe design of complex systems requires an enormous amount of modeling and simulation at the various levels of detail to assure high quality results. Most design methodologies require the use of separate modeling languages and simulation environments for each refinement step and simulation technique. Since one of the major goals of design is to reduce the time to market for a given system, the current approaches are becoming unacceptable. A new unified modeling (UM) methodology is under development at the University of Virginia which provides a top-down/bottom-up design methodology based on a single modeling environment. This single environment methodology is made possible through a technique known as hybrid modeling, which enable the simulation and analysis of behavioral components within a system-level model. By simulating behavioral components along with components at the system modeling level, high risk portions of the design can be examined at earlier stages in a given product's development; thus reducing the time to market. The paper examines the hybrid modeling techniques which make this multi-level modeling possible. Robert M. McGraw, Moshe Meyassed, Robert H. Klenke, James H. Aylor, Ronald D. Williams |
ICECCS | 3 |
| 1993 | Workstation Based Parallel Test GenerationabstractGeneration of test vectors for the VLSI devices used in contemporary digital system is becoming much more difficult as these devices increase in size. Automatic Test Pattern Generation (ATPG) techniques are commonly used to generate these tests. Since ATPG is an NP complete problem with complexity exponential to circuit size, the application of parallel processing techniques to accelerate the process of finding test patterns is an active area of research. This paper presents an approach to parallelization of the test generation problem that is targeted to a network-of-workstations environment. The system is based upon partitioning of the fault list across multiple processors and includes enhancements designed to address the main drawbacks of this technique, namely unequal load balancing and generation of redundant vectors. The technique is generalized enough that it can be applied to any test generation system regardless of the ATPG or fault simulation algorithm employed. Results were gathered to determine the impact of workstation processing load and network communications load on the performance of the system.> Robert H. Klenke, Lori M. Kaufman, James H. Aylor, Ronald Waxman, Padmini Narayan |
ITC | 1 |
| 1993 | Parallelization methods for circuit partitioning based parallel automatic test pattern generationabstractGeneration of test vectors for the VLSI devices used in contemporary digital systems is becoming much more difficult as these devices increase in size. Automatic Test Pattern Generation (ATPG) techniques are commonly used to generate these tests. Parallel processing techniques can be applied to accelerate the process of finding test patterns. One problem with this approach is that most currently available distributed memory multicomputers have a limited amount of memory on each processor which limits the size of the circuit database that can be contained on a single node. Topological partitioning of the circuit database across several processors can increase the size of VLSI circuits that can be processed on a given parallel machine. This paper presents the architecture of a topologically partitioned ATPG system and several partitioning algorithms that can be used to partition the circuit-under-test. This paper also presents several parallelization methods that may be applied to topologically partitioned ATPG on a distributed memory multicomputer. Results of using these parallelization techniques along with topological partitioning are presented.> Robert H. Klenke, Ronald D. Williams, James H. Aylor |
VTS | 1 |