Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Robert P. Colwell

dblp:c/RobertPColwell · also Bob Colwell · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
0since 2021 · last 2013
0000-0001-6250-7497ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 5 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Processor architecture and microarchitecture · 88% Performance modeling and evaluation · 9% Memory systems · 3%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture › instruction-level parallelism
VLIW
0.031990
Architecture and implementation of a VLIW supercomputer · SC 1990
A VLIW Architecure for a Trace Scheduling Compiler · IEEE Trans. Computers 1988
A VLIW Architecture for a Trace Scheduling Compiler · ASPLOS 1987
Processor architecture and microarchitecture
instruction-level parallelism
0.031990
Architecture and implementation of a VLIW supercomputer · SC 1990
A VLIW Architecure for a Trace Scheduling Compiler · IEEE Trans. Computers 1988
A VLIW Architecture for a Trace Scheduling Compiler · ASPLOS 1987
Compilers and program optimization
instruction scheduling
0.011988
A VLIW Architecure for a Trace Scheduling Compiler · IEEE Trans. Computers 1988
Performance modeling and evaluation
benchmarking
0.011988
Performance Effects of Architectural Complexity in the Intel 432 · ACM Trans. Comput. Syst. 1988
Processor architecture and microarchitecture › superscalar processor
wide-issue
0.011988
A VLIW Architecure for a Trace Scheduling Compiler · IEEE Trans. Computers 1988
Compilers and program optimization › instruction scheduling
trace scheduling
0.011987
A VLIW Architecture for a Trace Scheduling Compiler · ASPLOS 1987
Processor architecture and microarchitecture
instruction set architecture
0.011986
Fast Object-Oriented Procedure Calls: Lessons from the Intel 432 · ISCA 1986
Processor architecture and microarchitecture › instruction set architecture › high-level language architecture
object-oriented architecture
0.011986
Fast Object-Oriented Procedure Calls: Lessons from the Intel 432 · ISCA 1986
Memory systems
cache
0.011988
Performance Effects of Architectural Complexity in the Intel 432 · ACM Trans. Comput. Syst. 1988
Processor architecture and microarchitecture › instruction set architecture
tagged architecture
0.011986
Fast Object-Oriented Procedure Calls: Lessons from the Intel 432 · ISCA 1986

Methods — techniques the papers use, named apart from their topics

load/store architecture · 0.0compiler-controlled resource usage · 0.0trace scheduling · 0.0benchmarking · 0.0
YearPublicationVenuePosition
2013 The chip design game at the end of Moore's law
Robert P. Colwell
Hot Chips Symposium1
1990 Architecture and implementation of a VLIW supercomputer
abstract
Very-long-instruction-word (VLIW) computers achieve high performance by exploiting the fine-grain parallelism present in sequential or vectorizable code. Multiflow's /200 and /300 VLIW systems yielded near-supercomputer performance by this means despite the relatively slow (65 ns) clocks. With its much faster clock period (15 ns) and architectural improvements, the new /500 system attains approximately 4-9* the performance of its predecessors. The authors describe the /500 architecture and implementation (i.e. TRACE/500), with special attention paid to the tradeoffs involved in designing very-high-speed VLIWs.>
Robert P. Colwell, W. Eric Hall, Chandra S. Joshi, David B. Papworth, Paul K. Rodman, James E. Tornes
SC1
1988 A VLIW Architecure for a Trace Scheduling Compiler
abstract
A VLIW (very long instruction word) architecture machine called the TRACE has been built along with its companion Trace Scheduling compacting compiler. This machine has three hardware configurations, capable of executing 7, 14, or 28 operations simultaneously. The 'seven-wide' achieves a performance improvement of a factor of five or six for a wide range of scientific code, compared to machines of higher cost and fast chip implementation technology (such as the VAX 8700). The TRACE extends some basic reduced-instruction-set computer (RISC) precepts: the architecture is load/store, the microarchitecture is exposed to the compiler, there is no microcode, and there is almost no hardware devoted to synchronization, arbitration, or interlocking of any kind (the compiler has sole responsibility for run-time resource usage). The authors discuss the design of this machine and present some initial performance results.>
Robert P. Colwell, Robert P. Nix, John J. O'Donnell, David B. Papworth, Paul K. Rodman
IEEE Trans. Computers1
1988 Performance Effects of Architectural Complexity in the Intel 432
abstract
The Intel 432 is noteworthy as an architecture incorporating a large amount of functionality that most other systems perform by software. It has, in effect, “migrated” this functionality from the software into the microcode and hardware. The benefits of functional migration have recently been a subject of intense controversy, with critics claiming that a complex architecture is inherently less efficient than a simple architecture with good software support. This paper examines the performance impact of the incorporation of several kinds of functionality into the Intel 432. Among these are the addressing structure, the caches, instruction alignment, the buses, and the way that garbage collection is handled. A set of several benchmarks is used to quantify the performance effect of each of these decisions. The results indicate that the 432 could have been speeded up very significantly if a small number of implementation decisions had been made differently, and if incrementally better technology had been used in its construction. Even with these modifications, however, the 432 would still have only one-fourth to one times the speed of its contemporaries. These figures may represent the real cost of the 432's style of object-based programming environment.
Robert P. Colwell, Edward F. Gehringer, E. Douglas Jensen
ACM Trans. Comput. Syst.1
1987 A VLIW Architecture for a Trace Scheduling Compiler
abstract
Very Long Instruction Word (VLIW) architectures were promised to deliver far more than the factor of two or three that current architectures achieve from overlapped execution. Using a new type of compiler which compacts ordinary sequential code into long instruction words, a VLIW machine was expected to provide from ten to thirty times the performance of a more conventional machine built of the same implementation technology.Multiflow Computer, Inc., has now built a VLIW called the TRACETM along with its companion Trace SchedulingTM compacting compiler. This new machine has fulfilled the performance promises that were made. Using many fast functional units in parallel, this machine extends some of the basic Reduced-Instruction-Set precepts: the architecture is load/store, the microarchitecture is exposed to the compiler, there is no microcode, and there is almost no hardware devoted to synchronization, arbitration, or interlocking of any kind (the compiler has sole responsibility for runtime resource usage).This paper discusses the design of this machine and presents some initial performance results.
Robert P. Colwell, Robert P. Nix, John J. O'Donnell, David B. Papworth, Paul K. Rodman
ASPLOS1
1986 Fast Object-Oriented Procedure Calls: Lessons from the Intel 432
abstract
As modular programming grows in importance, the efficiency of procedure calls assumes an ever more critical role in system performance. Meanwhile, software designers are becoming more aware of the benefits of object-oriented programming in structuring large software systems. But object-oriented programming requires a good deal of support, which can best be distributed between the compiler and architectural levels. A major part of this support relates to the execution of procedure calls. Must such support exact an unacceptable performance penalty? By considering the case of the Intel 432, a prominent object-oriented architecture, we argue that it need not. The 432 provided all the facilities needed to support object orientation. Though its procedure call was slow, the reasons were only tenuously related to object orientation. Most of the inefficiency could be removed in future designs by the adoption of a few new mechanisms: stack-based allocation of contexts, a memory-clearing coprocessor, and the use of multiple register sets to hold addressing information. These proposals offer the prospect of an object-oriented procedure call that can, on average, be performed nearly as fast as an ordinary unprotected procedure call.
Edward F. Gehringer, Robert P. Colwell
ISCA2