David M. Ungar

dblp:84/5937 · also Dave Ungar · DBLP profile ↗
← Back
32ranked-venue papers
8as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 30 · 8 first-authorSystems, architecture and hardware · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
21 papers
Programming languages and type systems · 41% Compilers and program optimization · 29% Runtime systems and virtual machines · 26%
Computer architecture, parallel and distributed computing, and storage systems
5 papers
Processor architecture and microarchitecture · 53% Electronic design automation · 20% Parallel and multicore computing · 15%
Human-computer interaction and pervasive computing
2 papers
Interaction techniques and input · 57% User interface design and tools · 43%

Topics — the 30 heaviest of 43, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Programming languages and type systems
language design
0.122004
Mirrors: design principles for meta-level facilities of object-oriented programming languages · OOPSLA 2004
Self: The Power of Simplicity · OOPSLA 1987
Compilers and program optimization
dynamic optimization
0.031996
Reconciling Responsiveness with Performance in Pure Object-Orieted Languages · ACM Trans. Program. Lang. Syst. 1996
A Third-Generation SELF Implementation: Reconsiling Responsiveness with Performance · OOPSLA 1994
Making Pure Object-Oriented Languages Practical · OOPSLA 1991
Programming languages and type systems
object-oriented programming
0.051995
Annotating Objects for Transport to Other Worlds · OOPSLA 1995
Panel: Inheritance: Can We Have Our Cake and Eat it, Too? · OOPSLA 1989
Self: The Power of Simplicity · OOPSLA 1987
Compilers and program optimization › compiler optimization › type-based optimization
dynamic dispatch optimization
0.031994
Optimizing Dynamically-Dispatched Calls with Run-Time Type Feedback · PLDI 1994
Iterative Type Analysis and Extended Message Splitting: Optimizing Dynamically-Typed Object-Oriented Programs · PLDI 1990
Customization: Optimizing Compiler Technology for SELF, A Dynamically-Typed Object-Oriented Programming Language · PLDI 1989
Runtime systems and virtual machines › virtual machine implementation
java virtual machine
0.011998
The New Crop of Java Virtual Machines (Panel) · OOPSLA 1998
Runtime systems and virtual machines
virtual machine implementation
0.021994
A Third-Generation SELF Implementation: Reconsiling Responsiveness with Performance · OOPSLA 1994
SOAR: Smalltalk Without Bytecodes · OOPSLA 1986
Compilers and program optimization › interprocedural optimization
inlining
0.011996
Reconciling Responsiveness with Performance in Pure Object-Orieted Languages · ACM Trans. Program. Lang. Syst. 1996
Compilers and program optimization
interprocedural optimization
0.011996
Reconciling Responsiveness with Performance in Pure Object-Orieted Languages · ACM Trans. Program. Lang. Syst. 1996
Runtime systems and virtual machines
garbage collection
0.021992
An Adaptive Tenuring Policy for Generation Scavengers · ACM Trans. Program. Lang. Syst. 1992
Tenuring Policies for Generation-Based Storage Reclamation · OOPSLA 1988
Runtime systems and virtual machines
object-oriented language implementation
0.021991
Making Pure Object-Oriented Languages Practical · OOPSLA 1991
An Efficient Implementation of SELF - a Dynamically-Typed Object-Oriented Language Based on Prototypes · OOPSLA 1989
Interaction techniques and input
direct manipulation
0.011995
The Self-4.0 User Interface: Manifesting a System-wide Vision of Concreteness, Uniformity and Flexibility · OOPSLA 1995
Compilers and program optimization › program specialization
type specialization
0.021990
Iterative Type Analysis and Extended Message Splitting: Optimizing Dynamically-Typed Object-Oriented Programs · PLDI 1990
Customization: Optimizing Compiler Technology for SELF, A Dynamically-Typed Object-Oriented Programming Language · PLDI 1989
Software maintenance and evolution › software reengineering
application extraction
0.011994
Sifting Out the Gold · OOPSLA 1994
Runtime systems and virtual machines
dynamic compilation
0.011994
A Third-Generation SELF Implementation: Reconsiling Responsiveness with Performance · OOPSLA 1994
Programming languages and type systems
type inference
0.011994
Sifting Out the Gold · OOPSLA 1994
User interface design and tools
animation in user interfaces
0.011993
Animation: From Cartoons to the User Interface · ACM Symposium on User Interface Software and Technology 1993
Programming languages and type systems › object-oriented programming
object-oriented languages
0.041994
Optimizing Dynamically-Dispatched Calls with Run-Time Type Feedback · PLDI 1994
A Third-Generation SELF Implementation: Reconsiling Responsiveness with Performance · OOPSLA 1994
Compiling Smalltalk-80 to a RISC · ASPLOS 1987
Debugging and program repair › software debugging
debugging optimized code
0.011992
Debugging Optimized Code with Dynamic Deoptimization · PLDI 1992
Runtime systems and virtual machines › dynamic compilation › just-in-time compilation
deoptimization
0.011992
Debugging Optimized Code with Dynamic Deoptimization · PLDI 1992
Runtime systems and virtual machines › garbage collection
generational garbage collection
0.011992
An Adaptive Tenuring Policy for Generation Scavengers · ACM Trans. Program. Lang. Syst. 1992
Processor architecture and microarchitecture
instruction set architecture
0.021987
Compiling Smalltalk-80 to a RISC · ASPLOS 1987
Architecture of SOAR: Smalltalk on a RISC · ISCA 1984
Processor architecture and microarchitecture › instruction set architecture
RISC
0.021987
Compiling Smalltalk-80 to a RISC · ASPLOS 1987
Architecture of SOAR: Smalltalk on a RISC · ISCA 1984
Runtime systems and virtual machines › dynamic compilation
just-in-time compilation
0.011998
The New Crop of Java Virtual Machines (Panel) · OOPSLA 1998
Compilers and program optimization › dynamic optimization
dynamic language optimization
0.011989
An Efficient Implementation of SELF - a Dynamically-Typed Object-Oriented Language Based on Prototypes · OOPSLA 1989
Programming languages and type systems
inheritance
0.011989
Panel: Inheritance: Can We Have Our Cake and Eat it, Too? · OOPSLA 1989
Parallel and multicore computing
parallel programming environment
0.011988
Multiprocessor Smalltalk: A Case Study of a Multiprocessor-Based Programming Environment · PLDI 1988
Programming languages and type systems › method dispatch
dynamic dispatch
0.011996
Reconciling Responsiveness with Performance in Pure Object-Orieted Languages · ACM Trans. Program. Lang. Syst. 1996
Compilers and program optimization
code generation
0.011987
Compiling Smalltalk-80 to a RISC · ASPLOS 1987
Programming languages and type systems › object-oriented programming
object-oriented language design
0.011987
Self: The Power of Simplicity · OOPSLA 1987
Compilers and program optimization
register allocation
0.011987
Compiling Smalltalk-80 to a RISC · ASPLOS 1987

Methods — techniques the papers use, named apart from their topics

motion blur · 0.0dissolve · 0.0dynamic compilation · 0.0profile-guided optimization · 0.0procedure inlining · 0.0type inference · 0.0type feedback · 0.0interpretation · 0.0inlining · 0.0on-demand recompilation · 0.0dynamic deoptimization · 0.0demographic feedback · 0.0serialization · 0.0replication · 0.0reorganization · 0.0register window allocation · 0.0instruction-level simulation · 0.0hashing · 0.0
YearPublicationVenuePosition
2017 Dynamic atomicity: optimizing swift memory management
abstract
Swift is a modern multi-paradigm programming language with an extensive developer community and open source ecosystem. Swift 3's memory management strategy is based on Automatic Reference Counting (ARC) augmented with unsafe APIs for manually-managed memory. We have seen ARC consume as much as 80% of program execution time. A significant portion of ARC's direct performance cost can be attributed to its use of atomic machine instructions to protect reference count updates from data races. Consequently, we have designed and implemented dynamic atomicity, an optimization which safely replaces atomic reference-counting operations with nonatomic ones where feasible. The optimization introduces a store barrier to detect possibly intra-thread references, compiler-generated recursive reference-tracers to find all affected objects, and a bit of state in each reference count to encode its atomicity requirements.
David M. Ungar, David Grove, Hubertus Franke
DLS1
2013 When spatial and temporal locality collide: the case of the missing cache hits
abstract
Even the simplest hardware, running the simplest programs, can behave in the strangest of ways. Tracking down the cause of a performance anomaly without the complete hardware reference of a processor is a prime example of black-box architectural exploration. When doubling the work of a simple benchmark program, that was run on a single core of Tilera's TILEPro64 processor, did not double the number of consumed cycles, a mystery was unveiled. After ruling out different levels of optimization for the two programs, a cycle-accurate simulation attributed the sub-optimal performance to an abnormally high number of L1 data cache misses. Further investigation showed that the processor stalled on every Read-After-Write instruction sequence when the following two conditions were met: 1) there are 0 or 1 instructions between the write and the read instruction and 2) the read and the write instructions target distinct memory locations that share an L1 cache line. We call this performance pitfall a RAW hiccup. We describe two countermeasures, memory padding and the explicit introduction of pipeline bubbles, that sidestep the RAW hiccup.
Mattias De Wael, David M. Ungar, Tom Van Cutsem
ICPE2
2009 Hosting an object heap on manycore hardware: an exploration
abstract
In order to construct a test-bed for investigating new programming paradigms for future systems (i.e. those with at least a thousand cores), we are building a Smalltalk virtual machine that attempts to efficiently use a collection of 56-on-chip caches of 64KB each to host a multi-megabyte object heap. In addition to the cost of inter-core communication, two hardware characteristics influenced our design: the absence of hardware-provided cache-coherence, and the inability to move a single object from one core's cache to another's without changing its address. Our design relies on an object table, and the exploitation of a user-managed caching regime for read-mostly objects. At almost every stage of our process, we obtained measurements in order to guide the evolution of our system.The architecture and performance characteristics of a manycore platform confound old intuitions by deviating from both traditional multicore systems and from distributed systems. The implementor confronts a wide variety of design choices, such as when to share address space, when to share memory as opposed to sending a message, and how to eke out the most performance from a memory system that is far more tightly integrated than a distributed system yet far less centralized than in a several-core system. Our system is far from complete, let alone optimal, but our experiences have helped us develop new intuitions needed to rise to the manycore software challenge.
David M. Ungar, Sam S. Adams
DLS1
2008 An assembler and disassembler framework for JavaTMprogrammers
Bernd Mathiske, Doug Simon, David M. Ungar
Sci. Comput. Program.3
2004 Mirrors: design principles for meta-level facilities of object-oriented programming languages
abstract
We identify three design principles for reflection and metaprogramming facilities in object oriented programming languages. Encapsulation: meta-level facilities must encapsulate their implementation. Stratification: meta-level facilities must be separated from base-level functionality. Ontological correspondence: the ontology of meta-level facilities should correspond to the ontology of the language they manipulate. Traditional/mainstream reflective architectures do not follow these precepts. In contrast, reflective APIs built around the concept of mirrors are characterized by adherence to these three principles. Consequently, mirror-based architectures have significant advantages with respect to distribution, deployment and general purpose metaprogramming.
Gilad Bracha, David M. Ungar
OOPSLA2
1998 The New Crop of Java Virtual Machines (Panel)
Lars Bak, John Duimovich, Jesse Fang, Scott Meyer, David M. Ungar
OOPSLA5
1996 Reconciling Responsiveness with Performance in Pure Object-Orieted Languages
abstract
Dynamically dispatched calls often limit the performance of object-oriented programs, since opject-oriented programming encourages factoring code into small, reusable units, thereby increasing the frequency of these expensive operations. Frequent calls not only slow down execution with the dispatch overhead per se, but more importantly they hinder optimization by limiting the range and effectiveness of standard global optimizations. In particular, dynamically dispatched calles prevent standard interprocedual optimizations that depend on the availability of a static call graph. The SELF implementation described here offers tow novel approaches to optimization. Type feedback speculatively inlines dynamically dispatched calls based on profile information that predicts likely receiver classes. Adaptive optimization reconciles optimizing compilation with interactive performance by incrementally optimizing only the frequently executed parts of a program. When combined, these two techniques result in a system that can execute programs significantly faster than previous systems while retaining much of the interactiveness of an interpreted system.
Urs Hölzle, David M. Ungar
ACM Trans. Program. Lang. Syst.2
1995 Do Object-Oriented Languages Need Special Hardware Support?
Urs Hölzle, David M. Ungar
ECOOP2
1995 Programming as an Experience: The Inspiration for Self
Randall B. Smith, David M. Ungar
ECOOP2
1995 The Self-4.0 User Interface: Manifesting a System-wide Vision of Concreteness, Uniformity and Flexibility
abstract
Manipulating programs is hard, while manipulating objects in the physical world is often easy. Several attributes of the physical world help make it comprehensible and manipulable: concreteness, uniformity, and flexibility. The Self programming system attempts to apply these attributes to the world within the computer. The semantics of the language, the efficiency and fidelity of its implementation, and the architecture of its user interface conspire to make the experience of constructing programs in Self immediate and tangible. We describe the mechanisms used to achieve this goal, and illustrate those mechanisms within the context of an extended programming task.
Randall B. Smith, John Maloney, David M. Ungar
OOPSLA3
1995 Annotating Objects for Transport to Other Worlds
abstract
In Self 4.0, people write programs by directly constructing webs of objects in a larger world of objects. But in order to save or share these programs, the objects must be moved to other worlds. However, a concrete, directly constructed program is incomplete, in particular missing five items of information: which module to use, whether to transport an actual value or a counterfactuaI initial value, whether to create a new object in the new world or to refer to an existing one, whether an object is immutable with respect to transportation, and whether an object should be created by a low-level, concrete expression or an abstract, type-specific expression. In Self 4.0, the programmer records this extra information in annotations and attributes. Any system that saves directly constructed programs will have to supply this missing information somehow.
David M. Ungar
OOPSLA1
1994 Sifting Out the Gold
abstract
Integrated, dynamically-typed object-oriented programming environments offer many advantages, but have trouble producing small, self-contained applications. Recent advances in type inference have made it possible to build an application extractor for Self. The extractor was able to extract a medium-sized application in a few minutes. The extracted application runs in a tenth the space of the original environment. Except for extracting reflection and sends with computed selectors, the extractor runs without human intervention and fully preserves the behavior of the application.
Ole Agesen, David M. Ungar
OOPSLA2
1994 A Third-Generation SELF Implementation: Reconsiling Responsiveness with Performance
abstract
Programming systems should be both responsive (to support rapid development) and efficient (to complete computations quickly). Pure object-oriented languages are harder to implement efficiently since they need optimization to achieve good performance. Unfortunately, optimization conflicts with interactive responsiveness because it tends to produce long compilation pauses, leading to unresponsive programming environments. Therefore, to achieve good responsiveness, existing exploratory programming environments such as the Smalltalk-80 environment rely on interpretation or non-optimizing dynamic compilation. But such systems pay a price for their interactiveness, since they may execute programs several times slower than an optimizing system.
Urs Hölzle, David M. Ungar
OOPSLA2
1994 Prototype-Based Languages: Object Lessons from Class-Free Programming (Panel)
Randall B. Smith, Mark Lentczner, Walter R. Smith, Antero Taivalsaari, David M. Ungar
OOPSLA5
1994 Optimizing Dynamically-Dispatched Calls with Run-Time Type Feedback
abstract
AbstrachObject-oriented programs are difficult to optimize because they execute many dynamically-dispatched calls.These calls cannot easily be eliminated because the compiler does not know which callee will be invoked at runtime.We have developed a simple technique that feeds back type information from the runtime system to the compiler.With this type feedback, the compiler can inline any dynamically-dispatched call.Our compiler drastically reduces the calI frequency of a suite of large SELF applications (by a factor of 3.6) and improves performance by a factor of 1.7.We believe that type feedback could significantly reduce call frequencies and improve performance for most other objectoriented languages (statically-typed or not) as well as for languages with type-dependent operations such as generic arithmetic.
Urs Hölzle, David M. Ungar
PLDI2
1993 Animation: From Cartoons to the User Interface
abstract
User interfaces are often based on static presentations, a model ill suited for conveying change. Consequently, events on the screen frequently startle and confuse users. Cartoon animation, in contrast, is exceedingly successful at engaging its audience; even the most bizarre events are easily comprehended. The Self user interface has served as a testbed for the application of cartoon animation techniques as a means of making the interface easier to understand and more pleasant to use. Attention to timing and transient detail allows Self objects to move solidly. Use of cartoon-style motion blur allows Self objects to move quickly and still maintain their comprehensibility. Self objects arrive and depart smoothly, without sudden materializations and disappearances, and they rise to the front of overlapping objects smoothly through the use of dissolve. Anticipating motion with a small contrary motion and pacing the middle of transitions faster than the endpoints results in smoother and c...
Bay-Wei Chang, David M. Ungar
ACM Symposium on User Interface Software and Technology2
1992 Debugging Optimized Code with Dynamic Deoptimization
abstract
SELF's debugging system provides complete source-level debugging (expected behavior) with globally optimized code. It shields the debugger from optimizations performed by the compiler by dynamically deoptimizing code on demand. Deoptimization only affects the procedure activations that are actively being debugged; all other code runs at full speed. Deoptimization requires the compiler to supply debugging information at discrete interrupt points; the compiler can still perform extensive optimizations between interrupt points without affecting debuggability. At the same time, the inability to interrupt between interrupt points is invisible to the user. Our debugging system also handles programming changes during debugging. Again, the system provides expected behavior: it is possible to change a running program and immediately observe the effects of the change. Dynamic deoptimization transforms old compiled code (which may contain inlined copies of the old version of the changed procedure) into new versions reflecting the current source-level state. To the best of our knowledge, SELF is the first practical system providing full expected behavior with globally optimized code.
Urs Hölzle, Craig Chambers, David M. Ungar
PLDI3
1992 An Adaptive Tenuring Policy for Generation Scavengers
abstract
One of the more promising automatic storage reclamation techniques, generation scavenging, suffers poor performance if many objects live for a fairly long time and then die. We have investigated the severity of this problem by simulating a two-generation scavenger using traces taken from actual 4-h sessions. There was a wide variation in the sample runs, with garbage-collection overhead ranging from insignificant, during three of the runs, to severe, during a single run. All runs demonstrated that performance could be improved with two techniques: segregating large bitmaps and strings, and adapting the scavenger's tenuring policy according to demographic feedback. We therefore incorporated these ideas into a commercial Smalltalk implementation. These two improvements deserve consideration for any storage reclamation strategy that utilizes a generation scavenger.
David M. Ungar, Frank Jackson
ACM Trans. Program. Lang. Syst.1
1991 Optimizing Dynamically-Typed Object-Oriented Languages With Polymorphic Inline Caches
Urs Hölzle, Craig Chambers, David M. Ungar
ECOOP3
1991 Making Pure Object-Oriented Languages Practical
abstract
In the past, object-oriented language designers and programmers have been forced to choose between pure message passing and performance. Last year, our SELF system achieved close to half the speed of optimized C but suffered from impractically long compile times. Two new optimization techniques, deferred compilation of uncommon cases and non-backtracking splitting using path objects, have improved compilation speed by more than an order of magnitude. SELF now compiles about as fast as an optimizing C compiler and runs at over half the speed of optimized C. This new level of performance may make pure object-oriented languages practical. 1 Introduction In the past, object-oriented language designers and programmers have been forced to choose between purity and performance. In a pure object-oriented language, all computation, even low-level operations like variable accessing, arithmetic, and array indexing, is performed by sending messages to objects. Although a message send may cost o...
Craig Chambers, David M. Ungar
OOPSLA2
1990 Iterative Type Analysis and Extended Message Splitting: Optimizing Dynamically-Typed Object-Oriented Programs
abstract
Object-oriented languages have suffered from poor performance caused by frequent and slow dynamically-bound procedure calls. The best way to speed up a procedure call is to compile it out, but dynamic binding of object-oriented procedure calls without static receiver type information precludes inlining. Iterative type analysis and extended message splitting are new compilation techniques that extract much of the necessary type information and make it possible to hoist run-time type tests out of loops.
Craig Chambers, David M. Ungar
PLDI2
1989 An Efficient Implementation of SELF - a Dynamically-Typed Object-Oriented Language Based on Prototypes
abstract
We have developed and implemented techniques that double the performance of dynamically-typed object-oriented languages. Our SELF implementation runs twice as fast as the fastest Smalltalk implementation, despite SELF's lack of classes and explicit variables.
Craig Chambers, David M. Ungar, Elgin Lee
OOPSLA2
1989 Panel: Inheritance: Can We Have Our Cake and Eat it, Too?
J. Eliot B. Moss, Ralf Johnson, Alan Snyder, David W. Stemple, David M. Ungar
OOPSLA5
1989 Customization: Optimizing Compiler Technology for SELF, A Dynamically-Typed Object-Oriented Programming Language
abstract
Dynamically-typed object-oriented languages please programmers, but their lack of static type information penalizes performance. Our new implementation tech-niques extract static type information from declaration-free programs. Our system compiles several copies of a given procedure, each customized for one receiver type, so that the type of the receiver is bound at compile time. The compiler predicts types that are statically unknown but likely, and inserts run-time type tests to verify its predictions. It splits calls, compiling a copy on each control path, optimized to the specific types on that path. Coupling these new techniques with compile-time message lookup, aggressive procedure inlining, and traditional optimizations has doubled the performance of dynamically-typed object-oriented languages. 1.
Craig Chambers, David M. Ungar
PLDI2
1988 Tenuring Policies for Generation-Based Storage Reclamation
abstract
One of the most promising automatic storage reclamation techniques, generation-based storage reclamation, suffers poor performance if many objects live for a fairly long time and then die. We have investigated the severity of this problem by simulating Generation Scavenging automatic storage reclamation from traces of actual four-hour sessions. There was a wide variation in the sample runs, with garbage-collection overhead ranging from insignificant, during the interactive runs, to severe, during a single non-interactive run. All runs demonstrated that performance could be improved with two techniques: segregating large bitmaps and strings, and mediating tenuring with demographic feedback. These two improvements deserve consideration for any generation-based storage reclamation strategy.
David M. Ungar, Frank Jackson
OOPSLA1
1988 Panel: Treaty of Orlando Revisited
David M. Ungar, Henry Lieberman, Lynn Andrea Stein, Daniel Halbert
OOPSLA1
1988 Multiprocessor Smalltalk: A Case Study of a Multiprocessor-Based Programming Environment
abstract
We have adapted an interactive programming system (Smalltalk) to a multiprocessor (the Firefly). The task was not as difficult as might be expected, thanks to the application of three basic strategies: serialization, replication, and reorganization. Serialization of access to resources disallows concurrent access. Replication provides multiple instances of resources when they cannot or should not be serialized. Reorganization allows us to restructure part of the system when the other two strategies cannot be applied.
Joseph Pallas, David M. Ungar
PLDI2
1987 Compiling Smalltalk-80 to a RISC
abstract
The Smalltalk On A RISC project at U. C. Berkeley proves that a high-level object-oriented language can attain high performance on a modified reduced instruction set architecture. The single most important optimization is the removal of a layer of interpretation, compiling the bytecoded virtual machine instructions into low-level, register-based, hardware instructions. This paper describes the compiler and how it was affected by SOAR architectural features. The compiler generates code of reasonable density and speed. Because of Smalltalk-80's semantics, relatively few optimizations are possible, but hardware and software mechanisms at runtime offset these limitations. Register allocation for an architecture with register windows comprises the major task of the compiler. Performance analysis suggests that SOAR is not simple enough; several hardware features could be efficiently replaced by instruction sequences constructed by the compiler.
William R. Bush, A. Dain Samples, David M. Ungar, Paul N. Hilfinger
ASPLOS3
1987 Self: The Power of Simplicity
abstract
Self is a new object-oriented language for exploratory programming based on a small number of simple and concrete ideas: prototypes, slots, and behavior. Prototypes combine inheritance and instantiation to provide a framework that is simpler and more flexible than most object-oriented languages. Slots unite variables and procedures into a single construct. This permits the inheritance hierarchy to take over the function of lexical scoping in conventional languages. Finally, because Self does not distinguish state from behavior, it narrows the gaps between ordinary objects, procedures, and closures. Self's simplicity and expressiveness offer new insights into object-oriented computation.
David M. Ungar, Randall B. Smith
OOPSLA1
1986 SOAR: Smalltalk Without Bytecodes
abstract
We have implemented Smalltalk-80 on an instruction-level simulator for a RISC microcomputer called SOAR. Measurements suggest that even a conventional computer can provide high performance for Smalltalk-80 by abandoning the 'Smalltalk Virtual Machine' in favor of compiling Smalltalk directly to SOAR machine code, linearizing the activation records on the machine stack, eliminating the object table, and replacing reference counting with a new technique called Generation Scavenging. In order to implement these techniques, we had to find new ways of hashing objects, accessing often-used objects, invoking blocks, referencing activation records, managing activation record stacks, and converting the virtual machine images.
A. Dain Samples, David M. Ungar, Paul N. Hilfinger
OOPSLA2
1984 Architecture of SOAR: Smalltalk on a RISC
abstract
Smalltalk on a RISC (SOAR) is a simple, Von Neumann computer that is designed to execute the Smalltalk-80 system much faster than existing VLSI microcomputers. The Smalltalk-80 system is a highly productive programming environment but poses tough challenges for implementors: dynamic data typing, a high level instruction set, frequent and expensive procedure calls, and object-oriented storage management. SOAR compiles programs to a low level, efficient instruction set. Parallel tag checks permit high performance for the simple common cases and cause traps to software routines for the complex cases. Parallel register initialization and multiple on-chip register windows speed procedure calls. Sophisticated software techniques relieve the hardware of the burden of managing objects. We have initial evaluations of the effectiveness of the SOAR architecture by compiling and simulating benchmarks, and will prove SOAR's feasibility by fabricating a 35,000-transistor SOAR chip. These early results suggest that a Reduced Instruction Set Computer can provide high performance in an exploratory programming environment.
David M. Ungar, Ricki Blau, Peter Foley, A. Dain Samples, David A. Patterson 0001
ISCA1
1982 Measurements of a VLSI design
abstract
This paper presents data about three facets of a recently-completed VLSI design containing 45000 transistors. The first set of data describes the mask-level features of the circuit, from which it is seen that almost all features have at least one small dimension. The second set of data analyzes the hierarchical cell structure used by the designers to specify the circuit. The measurements show that composite cells have a different structure from primitive cells, and that, outside of arrays, cells are rarely re-used. The third set of data concerns the usage of an interactive layout program during the circuit's design. In spite of the circuit's size, the most frequently invoked commands were all simple.
John K. Ousterhout, David M. Ungar
DAC2