VLDB 2026 Research / reviewers in the wild / expert
Derek B. Gottlieb
dblp:70/480
· DBLP profile ↗
6ranked-venue papers
1as first author
0since 2021 · last 2005
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Processor architecture and microarchitecture · 56% Reconfigurable computing and FPGAs · 44% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
clustered architecture |
0.0 | 1 | 2004 | A reconfigurable unit for a clustered programmable-reconfigurable processor · FPGA 2004 |
Reconfigurable computing and FPGAs › reconfigurable architecture
reconfigurable arrays |
0.0 | 1 | 2004 | A reconfigurable unit for a clustered programmable-reconfigurable processor · FPGA 2004 |
Processor architecture and microarchitecture
pipelining |
0.0 | 1 | 2004 | A reconfigurable unit for a clustered programmable-reconfigurable processor · FPGA 2004 |
Methods — techniques the papers use, named apart from their topics
pipelining scheme · 0.0interleaved reconfigurable array · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2005 | Exploiting Pipelining to Tolerate Wire Delays in a Programmable-Reconfigurable ProcessorabstractAs fabrication technologies advance, increasing wire delays in semiconductor systems are leading to larger and larger gaps between the clock rates of circuits implemented in reconfigurable logic and those of conventional microprocessors. In this paper, we present a pipelining scheme for the Amalgam programmable-reconfigurable processor that divides long wire delays into multi-cycle operations and supports overlapping of independent computations. On streaming benchmark programs, this pipelining scheme increases the clock rates of Amalgam's reconfigurable clusters by up to 72%, allowing the pipelined Amalgam to maintain a 2.6/spl times/ performance advantage over a purely-programmable processor in a wide range of fabrication processes. Chi-Wei Wang, Nicholas P. Carter, Richard B. Kujoth, Jeffrey J. Cook, Derek B. Gottlieb |
FPL | 5 |
| 2004 | A reconfigurable unit for a clustered programmable-reconfigurable processorabstractIn a clustered programmable-reconfigurable processor, multiple programmable processors and blocks of reconfigurable logic communicate through a register-based communication mechanism, which reduces the impact of wire delay on clock cycle time. In this paper, we present a circuit-level design for the reconfigurable clusters used on the Amalgam programmable-reconfigurable processor. We outline our interleaved reconfigurable array design, which provides high bandwidth to and from the register file without requiring large amounts of register control logic. We characterize the latency of operations in our array, and present results that show the impact that this latency has on overall system performance in a range of fabrication processes. Finally, we present a pipelining scheme that enables the array to operate at clock rates closer to those of programmable processors and allows for better scaling in future technologies. Richard B. Kujoth, Chi-Wei Wang, Derek B. Gottlieb, Jeffrey J. Cook, Nicholas P. Carter |
FPGA | 3 |
| 2003 | Mapping computation kernels to clustered programmable-reconfigurable processorsabstractReconfigurable computing systems have shown the potential to surpass conventional processor architectures in performance for a growing range of applications. That performance, however, must be attained without significantly changing the design effort on the programmer's part, and without drastically increasing compilation time. In this paper, we present our compiler framework for mapping computation kernels to the reconfigurable clusters of Amalgam, a clustered programmable-reconfigurable processor. We first promote the use of the gated singular-assignment program dependence graph, a parallel intermediate program representation, to represent computation kernels. We then present an algorithm for mapping a computation kernel into the control FSM and datapath for a reconfigurable cluster. Finally, we describe our fast datapath synthesis tool-flow which preserves regularity and reduces the problem size by not flattening the datapath to gates. Jeffrey J. Cook, Lee Baugh, Derek B. Gottlieb, Nicholas P. Carter |
FPT | 3 |
| 2002 | Mapping Algorithms to the Amalgam Programmable-Reconfigurable ProcessorabstractThe Amalgam programmable-reconfigurable processor is designed to provide the computational power required by upcoming embedded applications without requiring the design of application-specific hardware. It integrates multiple programmable processors and blocks of reconfigurable logic onto a single chip, using a clustered architecture, similar to the one used on the M-Machine to reduce wire length and delay and allow implementation at high clock rates. The clustered architecture provides tremendous flexibility, allowing applications to exploit parallelism at whatever granularity is best-suited to the application, while the combination of reconfigurable logic and programmable processors delivers much higher performance than could be achieved through programmable processors alone. This abstract presents the results of our initial experiments in hand-mapping applications onto Amalgam. Five applications (IDCT, Rijndael encryption, nQueens, DNA sequence comparison, and image dithering) have been implemented, achieving speedups ranging from 8.7/spl times/ to 23.2/spl times/ over the performance of a single programmable cluster by using the complete resources of an Amalgam chip. Jeffrey J. Cook, Derek B. Gottlieb, Joshua D. Walstrom, Steve Ferrera, Chi-Wei Wang, Nicholas P. Carter |
FCCM | 2 |
| 2002 | The Design of the Amalgam Reconfigurable ClusterabstractAmalgam is a novel architecture for multifunction embedded systems. It integrates multiple reconfigurable and programmable processing resources (known as clusters) to achieve high-performance with low design effort on a variety of multimedia applications. The reconfigurable cluster (RClust) enables Amalgam to exploit the natural parallelism and operator granularities of a target application. The RClust contains a ring of reconfigurable logic interleaved with a banked register file to support Amalgam's register-based inter-cluster communication mechanism. This low-latency mechanism allows the RClust to coordinate with a programmable cluster (PClust) as a special purpose junctional unit implementing small custom operations. The relatively large size of the cluster, however, allows it to implement larger, more independent computational kernels. In this extended abstract, we describe the initial design of the RClust and present results from mapping several benchmarks to Amalgam architectures with and without RClust elements. Joshua D. Walstrom, Jeffrey J. Cook, Derek B. Gottlieb, Steve Ferrera, Chi-Wei Wang, Nicholas P. Carter |
FCCM | 3 |
| 2002 | Clustered programmable-reconfigurable processorsabstractIn order to pose a successful challenge to conventional processor architectures, reconfigurable computing systems must achieve significantly better performance than conventional programmable processors by both greatly reducing the number of clock cycles required to execute a wide range of applications and achieving high clock rates when implemented in deep-submicron fabrication technologies. In this paper, we describe the architecture of Amalgam, a clustered programmable-reconfigurable processor that integrates multiple conventional processors and blocks of reconfigurable logic onto a single chip. Amalgam's distributed architecture allows implementation at high clock rates by limiting the impact of wire delay on cycle time and delivers an average of 13.7/spl times/ speedup on our benchmark applications when compared to an equivalent architecture that contains only a single programmable processor. Derek B. Gottlieb, Jeffrey J. Cook, Joshua D. Walstrom, Steve Ferrera, Chi-Wei Wang, Nicholas P. Carter |
FPT | 1 |