Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

David Tarditi

dblp:91/6365 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
0since 2021 · last 2007
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 9 · 3 first-authorSystems, architecture and hardware · 4 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
7 papers
Programming languages and type systems · 42% Compilers and program optimization · 17% Runtime systems and virtual machines · 17%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Parallel and multicore computing · 70% Memory systems · 28% Performance modeling and evaluation · 2%
Network and information security
1 paper
Systems and software security · 100%

Topics — the 22 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Runtime systems and virtual machines
garbage collection
0.132006
Optimizing memory transactions · PLDI 2006
Memory System Performance of Programs with Intensive Heap Allocation · ACM Trans. Comput. Syst. 1995
Memory Subsystem Performance of Programs Using Copying Garbage Collection · POPL 1994
Systems and software security
operating system security
0.112007
Sealing OS processes to improve dependability and safety · EuroSys 2007
Operating systems › system security › operating system security › protection mechanism › isolation
process isolation
0.112007
Sealing OS processes to improve dependability and safety · EuroSys 2007
Programming languages and type systems › language implementation
typed intermediate language
0.122005
A simple typed intermediate language for object-oriented languages · POPL 2005
TIL: A Type-Directed Optimizing Compiler for ML · PLDI 1996
Concurrent programming › atomicity
atomic sections
0.112006
Optimizing memory transactions · PLDI 2006
Compilers and program optimization › code generation
GPU code generation
0.112006
Accelerator: using data parallelism to program GPUs for general-purpose uses · ASPLOS 2006
Concurrent programming
transactional memory
0.112006
Optimizing memory transactions · PLDI 2006
Parallel and multicore computing
data-parallel programming
0.112006
Accelerator: using data parallelism to program GPUs for general-purpose uses · ASPLOS 2006
Programming languages and type systems › method dispatch
dynamic dispatch
0.112005
A simple typed intermediate language for object-oriented languages · POPL 2005
Programming languages and type systems › language semantics › formal semantics
object-oriented language semantics
0.112005
A simple typed intermediate language for object-oriented languages · POPL 2005
Programming languages and type systems › type systems
subtyping
0.112005
A simple typed intermediate language for object-oriented languages · POPL 2005
Programming languages and type systems › language implementation
type-preserving compilation
0.112005
A simple typed intermediate language for object-oriented languages · POPL 2005
Programming languages and type systems
type systems
0.112005
A simple typed intermediate language for object-oriented languages · POPL 2005
Runtime systems and virtual machines › garbage collection
copying garbage collection
0.021995
Memory System Performance of Programs with Intensive Heap Allocation · ACM Trans. Comput. Syst. 1995
Memory Subsystem Performance of Programs Using Copying Garbage Collection · POPL 1994
Parallel and multicore computing
data parallelism
0.012006
Accelerator: using data parallelism to program GPUs for general-purpose uses · ASPLOS 2006
Parallel and multicore computing
parallel programming models
0.012006
Accelerator: using data parallelism to program GPUs for general-purpose uses · ASPLOS 2006
Compilers and program optimization › compiler construction
type-directed compilation
0.011996
TIL: A Type-Directed Optimizing Compiler for ML · PLDI 1996
Memory systems › cache
cache performance
0.011995
Memory System Performance of Programs with Intensive Heap Allocation · ACM Trans. Comput. Syst. 1995
Memory systems › memory management
heap allocation
0.011995
Memory System Performance of Programs with Intensive Heap Allocation · ACM Trans. Comput. Syst. 1995
Memory systems › cache
cache organization
0.011995
Memory System Performance of Programs with Intensive Heap Allocation · ACM Trans. Comput. Syst. 1995
Memory systems
memory hierarchy
0.011995
Memory System Performance of Programs with Intensive Heap Allocation · ACM Trans. Comput. Syst. 1995
Performance modeling and evaluation
workload characterization
0.011994
Memory Subsystem Performance of Programs Using Copying Garbage Collection · POPL 1994

Methods — techniques the papers use, named apart from their topics

sealing · 0.1runtime filtering · 0.1just-in-time compilation · 0.1direct access implementation · 0.1type soundness proof · 0.1decidable type checking · 0.1simulation · 0.0memory subsystem simulation · 0.0type-directed optimization · 0.0
YearPublicationVenuePosition
2007 Extending Object-Oriented Optimizations for Concurrent Programs
Kelly Heffner, David Tarditi, Michael D. Smith 0001
PACT2
2007 Sealing OS processes to improve dependability and safety
abstract
In most modern operating systems, a process is a hardware-protected abstraction for isolating code and data. This protection, however, is selective. Many common mechanisms---dynamic code loading, run-time code generation, shared memory, and intrusive system APIs---make the barrier between processes very permeable. This paper argues that this traditional open process architecture exacerbates the dependability and security weaknesses of modern systems.
Galen C. Hunt, Mark Aiken, Manuel Fähndrich, Chris Hawblitzel, Orion Hodson, James R. Larus, Steven Levi, Bjarne Steensgaard, David Tarditi, Ted Wobber
EuroSys9
2006 Accelerator: using data parallelism to program GPUs for general-purpose uses
abstract
GPUs are difficult to program for general-purpose uses. Programmers can either learn graphics APIs and convert their applications to use graphics pipeline operations or they can use stream programming abstractions of GPUs. We describe Accelerator, a system that uses data parallelism to program GPUs for general-purpose uses instead. Programmers use a conventional imperative programming language and a library that provides only high-level data-parallel operations. No aspects of GPUs are exposed to programmers. The library implementation compiles the data-parallel operations on the fly to optimized GPU pixel shader code and API calls.We describe the compilation techniques used to do this. We evaluate the effectiveness of using data parallelism to program GPUs by providing results for a set of compute-intensive benchmarks. We compare the performance of Accelerator versions of the benchmarks against hand-written pixel shaders. The speeds of the Accelerator versions are typically within 50% of the speeds of hand-written pixel shader code. Some benchmarks significantly outperform C versions on a CPU: they are up to 18 times faster than C code running on a CPU.
David Tarditi, Sidd Puri, Jose Oglesby
ASPLOS1
2006 Optimizing memory transactions
abstract
Atomic blocks allow programmers to delimit sections of code as 'atomic', leaving the language's implementation to enforce atomicity. Existing work has shown how to implement atomic blocks over word-based transactional memory that provides scalable multi-processor performance without requiring changes to the basic structure of objects in the heap. However, these implementations perform poorly because they interpose on all accesses to shared memory in the atomic block, redirecting updates to a thread-private log which must be searched by reads in the block and later reconciled with the heap when leaving the block.This paper takes a four-pronged approach to improving performance: (1) we introduce a new 'direct access' implementation that avoids searching thread-private logs, (2) we develop compiler optimizations to reduce the amount of logging (e.g. when a thread accesses the same data repeatedly in an atomic block), (3) we use runtime filtering to detect duplicate log entries that are missed statically, and (4) we present a series of GC-time techniques to compact the logs generated by long-running atomic blocks.Our implementation supports short-running scalable concurrent benchmarks with less than 50\% overhead over a non-thread-safe baseline. We support long atomic blocks containing millions of shared memory accesses with a 2.5-4.5x slowdown.
Tim Harris 0001, Mark Plesko, Avraham Shinnar, David Tarditi
PLDI4
2005 Broad New OS Research: Challenges and Opportunities
Galen C. Hunt, James R. Larus, David Tarditi, Ted Wobber
HotOS3
2005 A simple typed intermediate language for object-oriented languages
abstract
Traditional class and object encodings are difficult to use in practical type-preserving compilers because of the complexity of the encodings. We propose a simple typed intermediate language for compiling object-oriented languages and prove its soundness. The key ideas are to preserve lightweight notions of classes and objects instead of compiling them away and to separate name-based subclassing from structure-based subtyping. The language can express standard implementation techniques for both dynamic dispatch and runtime type tests. It has decidable type checking even with subtyping between quantified types with different bounds. Because of its simplicity, the language is a more suitable starting point for a practical type-preserving compiler than traditional encoding techniques.
David Tarditi
POPL2
2000 The Case for Profile-Directed Selection of Garbage Collectors
abstract
Many garbage-collected systems use a single garbage collection algorithm across all applications. It has long been known that this can produce poor performance on applications for which that collector is not well suited. In some systems, such as those that execute stand-alone compiled executables, an appropriate collector for each application can be selected from a pool of available collectors and tuned by using profile information. In a study of 20 benchmarks and several collectors, compiled with the Marmot optimizing Java-to-native compiler, for every collector there was at least one benchmark that would have been at least 15% faster with a more appropriate collector. The collectors are a copying collector, a generational copying collector, which is combined with each of 4 different write barriers, and the null collector, which allocates but never collects. A detailed analysis of storage management costs shows how they vary by application and collector.
Robert P. Fitzgerald, David Tarditi
ISMM2
2000 Compact Garbage Collection Tables
abstract
Garbage collection tables for finding pointers on the stack can be represented in 20-25% of the space previously reported. Live pointer information is often the same at many call sites because there are few pointers live across most call sites. This allows live pointer information to be represented compactly by a small index into a table of descriptions of pointer locations. The mapping from program counter values to those small indexes can be represented compactly using several techniques. The techniques a assign numbers to call sites and use those numbers to index an array of small indexes. One technique is to represent an array of return addresses by using a two-level table with 16-bit off-sets. Another technique is to use a sparse array of return addresses and interpolate the exact number via disassembly of the executable code.
David Tarditi
ISMM1
2000 Marmot: an optimizing compiler for Java
abstract
The Marmot system is a research platform for studying the implementation of high level programming languages. It currently comprises an optimizing native-code compiler, runtime system, and libraries for a large subset of Java. Marmot integrates well-known representation, optimization, code generation, and runtime techniques with a few Java-specific features to achieve competitive performance. This paper contains a description of the Marmot system design, along with highlights of our experience applying and adapting traditional implementation techniques to Java. A detailed performance evaluation assesses both Marmot's overall performance relative to other Java and C++ implementations, and the relative costs of various Java language features in Marmot-compiled code. Our experience with Marmot has demonstrated that well-known compilation techniques can produce very good performance for static Java applications – comparable or superior to other Java systems, and approaching that of C++ in some cases. Copyright © 2000 John Wiley & Sons, Ltd.
Robert P. Fitzgerald, Todd B. Knoblock, Erik Ruf, Bjarne Steensgaard, David Tarditi
Softw. Pract. Exp.5
1996 TIL: A Type-Directed Optimizing Compiler for ML
abstract
article Free Access Share on TIL: a type-directed optimizing compiler for ML Authors: D. Tarditi School of Computer Science, Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PA School of Computer Science, Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PAView Profile , G. Morrisett School of Computer Science, Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PA School of Computer Science, Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PAView Profile , P. Cheng School of Computer Science, Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PA School of Computer Science, Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PAView Profile , C. Stone School of Computer Science, Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PA School of Computer Science, Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PAView Profile , R. Harper School of Computer Science, Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PA School of Computer Science, Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PAView Profile , P. Lee School of Computer Science, Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PA School of Computer Science, Carnegie Mellon University, 5000 Forbes Avenue, Pittsburgh, PAView Profile Authors Info & Claims ACM SIGPLAN NoticesVolume 31Issue 5May 1996 pp 181–192https://doi.org/10.1145/249069.231414Online:01 May 1996Publication History 212citation719DownloadsMetricsTotal Citations212Total Downloads719Last 12 Months39Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
David Tarditi, J. Gregory Morrisett, Perry Cheng, Christopher A. Stone, Robert Harper 0001, Peter Lee 0001
PLDI1
1995 Memory System Performance of Programs with Intensive Heap Allocation
abstract
Heap allocation with copying garbage collection is a general storage management technique for programming languages. It is believed to have poor memory system performance. To investigate this, we conducted an in-depth study of the memory system performance of heap allocation for memory systems found on many machines. We studied the performance of mostly functional Standard ML programs which made heavy use of heap allocation. We found that most machines support heap allocation poorly. However, with the appropriate memory system organization, heap allocation can have good performance. The memory system property crucial for achieving good performance was the ability to allocate and initialize a new object into the cache without a penalty. This can be achieved by having subblock by placement with a subblock size of one word with a write-allocate policy, along with fast page-mode writes or a write buffer. For caches with subblock placement, the data cache overhead was under 9% for a 64K or larger data cache; without subblock placement the overhead was often higher than 50%.
Amer Diwan, David Tarditi, J. Eliot B. Moss
ACM Trans. Comput. Syst.2
1994 Memory Subsystem Performance of Programs Using Copying Garbage Collection
abstract
Heap allocation with copying garbage collection is believed to have poor memory subsystem performance. We conducted a study of the memory subsystem performance of heap allocation for memory subsystems found on many machines. We found that many machines support heap allocation poorly. However, with the appropriate memory subsystem organization, heap allocation can have good memory subsystem performance.
Amer Diwan, David Tarditi, J. Eliot B. Moss
POPL2