VLDB 2026 Research / reviewers in the wild / expert
Robert P. Colwell
dblp:c/RobertPColwell · also Bob Colwell
· DBLP profile ↗
6ranked-venue papers
5as first author
0since 2021 · last 2013
0000-0001-6250-7497ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 5 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Processor architecture and microarchitecture · 88% Performance modeling and evaluation · 9% Memory systems · 3% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 100% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture › instruction-level parallelism
VLIW |
0.0 | 3 | 1990 | Architecture and implementation of a VLIW supercomputer · SC 1990 A VLIW Architecure for a Trace Scheduling Compiler · IEEE Trans. Computers 1988 A VLIW Architecture for a Trace Scheduling Compiler · ASPLOS 1987 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.0 | 3 | 1990 | Architecture and implementation of a VLIW supercomputer · SC 1990 A VLIW Architecure for a Trace Scheduling Compiler · IEEE Trans. Computers 1988 A VLIW Architecture for a Trace Scheduling Compiler · ASPLOS 1987 |
Compilers and program optimization
instruction scheduling |
0.0 | 1 | 1988 | A VLIW Architecure for a Trace Scheduling Compiler · IEEE Trans. Computers 1988 |
Performance modeling and evaluation
benchmarking |
0.0 | 1 | 1988 | Performance Effects of Architectural Complexity in the Intel 432 · ACM Trans. Comput. Syst. 1988 |
Processor architecture and microarchitecture › superscalar processor
wide-issue |
0.0 | 1 | 1988 | A VLIW Architecure for a Trace Scheduling Compiler · IEEE Trans. Computers 1988 |
Compilers and program optimization › instruction scheduling
trace scheduling |
0.0 | 1 | 1987 | A VLIW Architecture for a Trace Scheduling Compiler · ASPLOS 1987 |
Processor architecture and microarchitecture
instruction set architecture |
0.0 | 1 | 1986 | Fast Object-Oriented Procedure Calls: Lessons from the Intel 432 · ISCA 1986 |
Processor architecture and microarchitecture › instruction set architecture › high-level language architecture
object-oriented architecture |
0.0 | 1 | 1986 | Fast Object-Oriented Procedure Calls: Lessons from the Intel 432 · ISCA 1986 |
Memory systems
cache |
0.0 | 1 | 1988 | Performance Effects of Architectural Complexity in the Intel 432 · ACM Trans. Comput. Syst. 1988 |
Processor architecture and microarchitecture › instruction set architecture
tagged architecture |
0.0 | 1 | 1986 | Fast Object-Oriented Procedure Calls: Lessons from the Intel 432 · ISCA 1986 |
Methods — techniques the papers use, named apart from their topics
load/store architecture · 0.0compiler-controlled resource usage · 0.0trace scheduling · 0.0benchmarking · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | The chip design game at the end of Moore's law
Robert P. Colwell |
Hot Chips Symposium | 1 |
| 1990 | Architecture and implementation of a VLIW supercomputerabstractVery-long-instruction-word (VLIW) computers achieve high performance by exploiting the fine-grain parallelism present in sequential or vectorizable code. Multiflow's /200 and /300 VLIW systems yielded near-supercomputer performance by this means despite the relatively slow (65 ns) clocks. With its much faster clock period (15 ns) and architectural improvements, the new /500 system attains approximately 4-9* the performance of its predecessors. The authors describe the /500 architecture and implementation (i.e. TRACE/500), with special attention paid to the tradeoffs involved in designing very-high-speed VLIWs.> Robert P. Colwell, W. Eric Hall, Chandra S. Joshi, David B. Papworth, Paul K. Rodman, James E. Tornes |
SC | 1 |
| 1988 | A VLIW Architecure for a Trace Scheduling CompilerabstractA VLIW (very long instruction word) architecture machine called the TRACE has been built along with its companion Trace Scheduling compacting compiler. This machine has three hardware configurations, capable of executing 7, 14, or 28 operations simultaneously. The 'seven-wide' achieves a performance improvement of a factor of five or six for a wide range of scientific code, compared to machines of higher cost and fast chip implementation technology (such as the VAX 8700). The TRACE extends some basic reduced-instruction-set computer (RISC) precepts: the architecture is load/store, the microarchitecture is exposed to the compiler, there is no microcode, and there is almost no hardware devoted to synchronization, arbitration, or interlocking of any kind (the compiler has sole responsibility for run-time resource usage). The authors discuss the design of this machine and present some initial performance results.> Robert P. Colwell, Robert P. Nix, John J. O'Donnell, David B. Papworth, Paul K. Rodman |
IEEE Trans. Computers | 1 |
| 1988 | Performance Effects of Architectural Complexity in the Intel 432abstractThe Intel 432 is noteworthy as an architecture incorporating a large amount of functionality that most other systems perform by software. It has, in effect, “migrated” this functionality from the software into the microcode and hardware. The benefits of functional migration have recently been a subject of intense controversy, with critics claiming that a complex architecture is inherently less efficient than a simple architecture with good software support. This paper examines the performance impact of the incorporation of several kinds of functionality into the Intel 432. Among these are the addressing structure, the caches, instruction alignment, the buses, and the way that garbage collection is handled. A set of several benchmarks is used to quantify the performance effect of each of these decisions. The results indicate that the 432 could have been speeded up very significantly if a small number of implementation decisions had been made differently, and if incrementally better technology had been used in its construction. Even with these modifications, however, the 432 would still have only one-fourth to one times the speed of its contemporaries. These figures may represent the real cost of the 432's style of object-based programming environment. Robert P. Colwell, Edward F. Gehringer, E. Douglas Jensen |
ACM Trans. Comput. Syst. | 1 |
| 1987 | A VLIW Architecture for a Trace Scheduling CompilerabstractVery Long Instruction Word (VLIW) architectures were promised to deliver far more than the factor of two or three that current architectures achieve from overlapped execution. Using a new type of compiler which compacts ordinary sequential code into long instruction words, a VLIW machine was expected to provide from ten to thirty times the performance of a more conventional machine built of the same implementation technology.Multiflow Computer, Inc., has now built a VLIW called the TRACETM along with its companion Trace SchedulingTM compacting compiler. This new machine has fulfilled the performance promises that were made. Using many fast functional units in parallel, this machine extends some of the basic Reduced-Instruction-Set precepts: the architecture is load/store, the microarchitecture is exposed to the compiler, there is no microcode, and there is almost no hardware devoted to synchronization, arbitration, or interlocking of any kind (the compiler has sole responsibility for runtime resource usage).This paper discusses the design of this machine and presents some initial performance results. Robert P. Colwell, Robert P. Nix, John J. O'Donnell, David B. Papworth, Paul K. Rodman |
ASPLOS | 1 |
| 1986 | Fast Object-Oriented Procedure Calls: Lessons from the Intel 432abstractAs modular programming grows in importance, the efficiency of procedure calls assumes an ever more critical role in system performance. Meanwhile, software designers are becoming more aware of the benefits of object-oriented programming in structuring large software systems. But object-oriented programming requires a good deal of support, which can best be distributed between the compiler and architectural levels. A major part of this support relates to the execution of procedure calls. Must such support exact an unacceptable performance penalty? By considering the case of the Intel 432, a prominent object-oriented architecture, we argue that it need not. The 432 provided all the facilities needed to support object orientation. Though its procedure call was slow, the reasons were only tenuously related to object orientation. Most of the inefficiency could be removed in future designs by the adoption of a few new mechanisms: stack-based allocation of contexts, a memory-clearing coprocessor, and the use of multiple register sets to hold addressing information. These proposals offer the prospect of an object-oriented procedure call that can, on average, be performed nearly as fast as an ordinary unprotected procedure call. Edward F. Gehringer, Robert P. Colwell |
ISCA | 2 |