Lucas Roh

dblp:29/2401 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
0since 2021 · last 2001
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-authorSoftware engineering, systems software and programming languages · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 63% Processor architecture and microarchitecture · 25% Performance modeling and evaluation · 12%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory hierarchy
cache hierarchy
0.011995
Design of storage hierarchy in multithreaded architectures · MICRO 1995
Processor architecture and microarchitecture
multithreading
0.011995
Design of storage hierarchy in multithreaded architectures · MICRO 1995
Memory systems
cache
0.011993
An evaluation of bottom-up and top-down thread generation techniques · MICRO 1993
Memory systems › cache
cache miss
0.011993
An evaluation of bottom-up and top-down thread generation techniques · MICRO 1993
Memory systems
memory referencing behavior
0.011993
An evaluation of bottom-up and top-down thread generation techniques · MICRO 1993
Memory systems › cache
prefetching
0.011993
An evaluation of bottom-up and top-down thread generation techniques · MICRO 1993
Performance modeling and evaluation
workload characterization
0.011993
An evaluation of bottom-up and top-down thread generation techniques · MICRO 1993
Processor architecture and microarchitecture
latency hiding
0.011995
Design of storage hierarchy in multithreaded architectures · MICRO 1995
Processor architecture and microarchitecture › multithreading
multithreaded execution
0.011995
Design of storage hierarchy in multithreaded architectures · MICRO 1995

Methods — techniques the papers use, named apart from their topics

simulation-based evaluation · 0.0profiling · 0.0cache simulation · 0.0
YearPublicationVenuePosition
2001 Resource Management in Dataflow-Based Multithreaded Execution
Lucas Roh, Bhanu Shankar, A. P. Wim Böhm, Walid A. Najjar
J. Parallel Distributed Comput.1
1997 Algorithms and Design for a Second-Order Automatic Differentiation Module
abstract
This article describes approaches to computing second-order derivatives with automatic differentiation (AD) based on the forward mode and the propagation of univariate Taylor series. Performance results are given that show the speedup possible with these techniques relative to existing approaches. We also describe a new source transformation AD module for computing second-order derivatives of C and Fortran codes and the underlying infrastructure used to create a language-independent translation tool. 1 Introduction Automatic differentiation (AD) provides an efficient and accurate method to obtain derivatives for use in sensitivity analysis, parameter identification and optimization. Current tools are targeted primarily at computing first-order derivatives, namely gradients and Jacobians. Prior to AD, derivative values were obtained through divided difference methods, symbolic manipulation or hand-coding, all of which have drawbacks when compared with AD (see [4] for a dis- This wo...
Jason Abate, Christian H. Bischof, Lucas Roh, Alan Carle
ISSAC3
1997 ADIC: An Extensible Automatic Differentiation Tool for ANSI-C
abstract
In scientific computing, we often require the derivatives ∂f/∂x of a function f expressed as a program with respect to some input parameter(s) x, say. Automatic Differentiation (AD) techniques augment the program with derivative computation by applying the chain rule of calculus to elementary operations in an automated fashion. This article introduces ADIC (Automatic Differentiation of C), a new AD tool for ANSI-C programs. ADIC is currently the only tool for ANSI-C that employs a source-to-source program transformation approach; that is, it takes a C code and produces a new C code that computes the original results as well as the derivatives. We first present ADIC ‘by example’ to illustrate the functionality and ease of use of ADIC and then describe in detail the architecture of ADIC. ADIC incorporates a modular design that provides a foundation for both rapid prototyping of better AD algorithms and their sharing across AD tools for different languages. A component architecture called AIF (Automatic Differentiation Intermediate Form) separates core AD concepts from their language-specific implementation and allows the development of generic AD modules that can be reused directly in other AIF-based AD tools. The language-specific ADIC front-end and back-end canonicalize C programs to make them fit for semantic augmentation and manage, for example, the association of a program variable with its derivative object. We also report on applications of ADIC to a semiconductor device simulator, 3-D CFD grid generator, vehicle simulator, and neural network code. © 1997 John Wiley & Sons, Ltd.
Christian H. Bischof, Lucas Roh, A. J. Mauer-Oats
Softw. Pract. Exp.2
1996 Generation, Optimization, and Evaluation of Multithreaded Code
Lucas Roh, Walid A. Najjar, Bhanu Shankar, A. P. Wim Böhm
J. Parallel Distributed Comput.1
1995 Analysis of communications and overhead reduction in multithreaded execution
Lucas Roh, Walid A. Najjar
PACT1
1995 Control of loop parallelism in multithreaded code
Bhanu Shankar, Lucas Roh, A. P. Wim Böhm, Walid A. Najjar
PACT2
1995 Design of storage hierarchy in multithreaded architectures
abstract
Multithreaded execution models attempt to combine some aspects of dataflow-like execution with von Neumann model execution. Their main objective is to mask the latency of inter-processor communications and remote memory accesses in large scale multiprocessors. An important issue in the analysis and evaluation of multithreaded execution is the design and performance of the storage hierarchy. Because of the sequential execution of threads, the locality of access within an executing thread can be exploited using registers and cache. At the inter-thread level, however, the locality of accesses to memory and its effect on the cache is not yet well understood. A storage model which can exploit this locality is developed and evaluated. The results indicate there is a large amount of inter-thread locality that can be exploited and that we can get an efficient storage system by exploiting the characteristics of nonblocking threads.
Lucas Roh, Walid A. Najjar
MICRO1
1993 An evaluation of bottom-up and top-down thread generation techniques
abstract
Due to increasing cache-miss latencies, cache control instructions are being implemented for future systems. The authors study the memory referencing behavior of individual machine-level instructions using simulations of fully-associative caches under MIN replacement. Their objective is to obtain a deeper understanding of useful program behavior that can be eventually employed at optimizing programs and to motivate architectural features aimed at improving the efficacy of memory hierarchies. The simulation results show that a very small number of load/store instructions account for a majority of data cache misses. Specifically, fewer than 10 instructions account for half the misses for six out of nine SPEC89 benchmarks. Selectively prefetching data referenced by a small number of instructions identified through profiling can reduce overall miss ratio significantly while only incurring a small number of unnecessary prefetches.>
A. P. Wim Böhm, Walid A. Najjar, Bhanu Shankar, Lucas Roh
MICRO4