EDBT 2026 Demo / reviewers in the wild / expert
David R. Ditzel
dblp:98/1840
· DBLP profile ↗
9ranked-venue papers
6as first author
1since 2021 · last 2021
0000-0003-0944-3353ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 6 first-author · 1 since 2021Software engineering, systems software and programming languages · 8 · 5 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Processor architecture and microarchitecture · 60% Electronic design automation · 25% Parallel and multicore computing · 14% | |
| Software engineering, system software, and programming languages
5 papers |
Runtime systems and virtual machines · 59% Compilers and program optimization · 41% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
instruction set architecture |
0.2 | 7 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 An Analysis of SPARC and MIPS Instruction Set Utilization on the SPEC Benchmarks · ASPLOS 1991 Branch Folding in the CRISP Microprocessor: Reducing Branch Delay to Zero · ISCA 1987 |
Processor architecture and microarchitecture › computer arithmetic › floating-point arithmetic
fused multiply-add |
0.2 | 1 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 |
Electronic design automation
hardware/software co-design |
0.2 | 1 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 |
Runtime systems and virtual machines
binary translation |
0.1 | 1 | 2010 | A real system evaluation of hardware atomicity for software speculation · ASPLOS 2010 |
Compilers and program optimization
dynamic optimization |
0.1 | 1 | 2010 | A real system evaluation of hardware atomicity for software speculation · ASPLOS 2010 |
Parallel and multicore computing › transactional memory
hardware transactional memory |
0.1 | 1 | 2010 | A real system evaluation of hardware atomicity for software speculation · ASPLOS 2010 |
Runtime systems and virtual machines › binary translation
dynamic binary translation |
0.1 | 1 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 |
Processor architecture and microarchitecture
speculative execution |
0.0 | 1 | 2010 | A real system evaluation of hardware atomicity for software speculation · ASPLOS 2010 |
Processor architecture and microarchitecture › instruction set architecture
RISC |
0.0 | 1 | 1991 | An Analysis of SPARC and MIPS Instruction Set Utilization on the SPEC Benchmarks · ASPLOS 1991 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 1991 | An Analysis of SPARC and MIPS Instruction Set Utilization on the SPEC Benchmarks · ASPLOS 1991 |
Processor architecture and microarchitecture › branch handling
branch folding |
0.0 | 1 | 1987 | Branch Folding in the CRISP Microprocessor: Reducing Branch Delay to Zero · ISCA 1987 |
Processor architecture and microarchitecture
pipelining |
0.0 | 1 | 1987 | Branch Folding in the CRISP Microprocessor: Reducing Branch Delay to Zero · ISCA 1987 |
Compilers and program optimization
code generation |
0.0 | 1 | 1991 | An Analysis of SPARC and MIPS Instruction Set Utilization on the SPEC Benchmarks · ASPLOS 1991 |
Electronic design automation › high-level synthesis › resource binding
register allocation |
0.0 | 1 | 1982 | Register Allocation for Free: The C Machine Stack Cache · ASPLOS 1982 |
Processor architecture and microarchitecture › instruction set architecture
high-level language architecture |
0.0 | 1 | 1980 | Retrospective on High-Level Language Computer Architecture · ISCA 1980 |
Methods — techniques the papers use, named apart from their topics
speculative instruction-fusion optimization · 0.4cycle-accurate simulation · 0.4atomic region compiler abstraction · 0.2dynamic instruction counting · 0.0design tradeoff analysis · 0.0architecture-compiler co-design · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Accelerating ML Recommendation with over a Thousand RISC-V/Tensor Processors on Esperanto's ET-SoC-1 ChipabstractThe ET-SoC-1 has over a thousand RISC-V processors on a single TSMC 7nm chip, including: • 1088 energy-efficient ET-Minion 64-bit RISC-V in-order cores each with a vector/tensor unit • 4 high-performance ET-Maxion 64-bit RISC-V out-of-order cores • >160 million bytes of on-chip SRAM • Interfaces for large external memory with low-power LPDDR4x DRAM and eMMC FLASH • PCIe x8 Gen4 and other common I/O interfaces • Innovative low-power architecture and circuit techniques allows entire chip to • Compute at peak rates of 100 to 200 TOPS • Operate using under 20 watts for ML recommendation workloads David R. Ditzel, Roger Espasa, Nivard Aymerich, Allen Baum, Tom Berg, Jim Burr, Eric Hao, Jayesh Iyer, Miquel Izquierdo, Shankar Jayaratnam, Darren Jones, Chris Klingner, Stephen Lee, Marc Lupon, Grigorios Magklis, Bojan Maric, Rajib Nath, Mike Neilly, J. Duane Northcutt, Bill Orner, Jose Renau, Gerard Reves, Xavier Reves, Tom Riordan, Pedro Sanchez, Sridhar Samudrala, Guillem Sole, Raymond Tang, Tommy Thorn, Sebastia Tortella, Daniel Yau |
HCS | 1 |
| 2014 | Speculative hardware/software co-designed floating-point multiply-add fusionabstractA Fused Multiply-Add (FMA) instruction is currently available in many general-purpose processors. It increases performance by reducing latency of dependent operations and increases precision by computing the result as an indivisible operation with no intermediate rounding. However, since the arithmetic behavior of a single-rounding FMA operation is different than independent FP multiply followed by FP add instructions, some algorithms require significant revalidation and rewriting efforts to work as expected when they are compiled to operate with FMA--a cost that developers may not be willing to pay. Because of that, abundant legacy applications are not able to utilize FMA instructions. In this paper we propose a novel HW/SW collaborative technique that is able to efficiently execute workloads with increased utilization of FMA, by adding the option to get the same numerical result as separate FP multiply and FP add pairs. In particular, we extended the host ISA of a HW/SW co-designed processor with a new Combined Multiply-Add (CMA) instruction that performs an FMA operation with an intermediate rounding. This new instruction is used by a transparent dynamic translation software layer that uses a speculative instruction-fusion optimization to transform FP multiply and FP add sequences into CMA instructions. The FMA unit has been slightly modified to support both single-rounding and double-rounding fused instructions without increasing their latency and to provide a conservative fall-back path in case of mispeculation. Evaluation on a cycle-accurate timing simulator showed that CMA improved SPECfp performance by 6.3% and reduced executed instructions by 4.7%. Marc Lupon, Enric Gibert, Grigorios Magklis, Sridhar Samudrala, Raúl Martínez, Kyriakos Stavrou, David R. Ditzel |
ASPLOS | 7 |
| 2010 | A real system evaluation of hardware atomicity for software speculationabstractIn this paper we evaluate the atomic region compiler abstraction by incorporating it into a commercial system. We find that atomic regions are simple and intuitive to integrate into an x86 binary-translation system. Furthermore, doing so trivially enables additional optimization opportunities beyond that achievable by a high-performance dynamic optimizer, which already implements superblocks. Naveen Neelakantam, David R. Ditzel, Craig B. Zilles |
ASPLOS | 2 |
| 1991 | An Analysis of SPARC and MIPS Instruction Set Utilization on the SPEC BenchmarksabstractCopyright 1991, Association for Computing Machinery, Inc. Reprinted by permission. This paper originally appeared in Proceedings of the Fourth International Conference on Architecture Support for Programming Languages and Operating Systems (ASPLOS), Santa Clara, California, April 8-11, 1991. The dynamic instruction counts on MIPS and SPARC are compared using the SPEC benchmarks. MIPS typically executes more user-level instructions than SPARC. This difference can be counted for by architectural differences, compiler differences, and library differences. The most significant differences are that SPARC's double-precision floating point load/store is an architectural advantage in the SPEC floating point benchmarks while MIPS's compare-and-branch instruction is an architectural advantage in the SPEC integer benchmarks. After the differences in the two architectures are isolated, it appears that although MIPS and SPARC each have strengths and weaknesses in their compilers and library routines, the combined effect of compilers and library routines does not give either MIPS or SPARC a clear advantage in these areas. Robert F. Cmelik, Shing I. Kong, David R. Ditzel, Edmund J. Kelly |
ASPLOS | 3 |
| 1987 | Design Tradeoffs to Support the C Programming Language in the CRISP MicroprocessorabstractArticle Free Access Share on Design tradeoffs to support the C programming language in the CRISP microprocessor Authors: David R. Ditzel AT&T Bell Laboratories, Murray Hill, NJ AT&T Bell Laboratories, Murray Hill, NJSearch about this author , Hubert R. McLellan AT&T Bell Laboratories, Murray Hill, NJ AT&T Bell Laboratories, Murray Hill, NJSearch about this author , Alan D. Berenbaum AT&T Information Systems, Holmdel, NJ AT&T Information Systems, Holmdel, NJSearch about this author Authors Info & Claims ASPLOS II: Proceedings of the second international conference on Architectual support for programming languages and operating systemsOctober 1987 Pages 158–163https://doi.org/10.1145/36206.36198Online:01 October 1987Publication History 9citation608DownloadsMetricsTotal Citations9Total Downloads608Last 12 Months19Last 6 weeks7 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF David R. Ditzel, Hubert R. McLellan |
ASPLOS | 1 |
| 1987 | Branch Folding in the CRISP Microprocessor: Reducing Branch Delay to ZeroabstractA new method of implementing branch instructions is presented. This technique has been implemented in the CRISP Microprocessor. With a combination of hardware and software techniques the execution time cost for many branches can be effectively reduced to zero. Branches are folded into other instructions, making their execution as separate instructions unnecessary. Branch Folding can reduce the apparent number of instructions needed to execute a program by the number of branches in that program, as well as reducing or eliminating pipeline breakage. Statistics are presented demonstrating the effectiveness of Branch Folding and associated techniques used in the CRISP Microprocessor. David R. Ditzel, Hubert R. McLellan |
ISCA | 1 |
| 1987 | The Hardware Architecture of the CRISP MicroprocessorabstractArticle The hardware architecture of the CRISP microprocessor Share on Authors: D. R. Ditzel AT&T Bell Laboratories, Murray Hill, NJ AT&T Bell Laboratories, Murray Hill, NJView Profile , H. R. McLellan AT&T Bell Laboratories, Murray Hill, NJ AT&T Bell Laboratories, Murray Hill, NJView Profile , A. D. Berenbaum AT&T Information Systems, Holmdel, NJ AT&T Information Systems, Holmdel, NJView Profile Authors Info & Claims ISCA '87: Proceedings of the 14th annual international symposium on Computer architectureJune 1987 Pages 309–319https://doi.org/10.1145/30350.30385Published:01 June 1987 39citation493DownloadsMetricsTotal Citations39Total Downloads493Last 12 Months6Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access David R. Ditzel, Hubert R. McLellan, Alan D. Berenbaum |
ISCA | 1 |
| 1982 | Register Allocation for Free: The C Machine Stack CacheabstractThe Bell Labs C Machine project is investigating computer architectures to support the C programming language.1 One of the goals is to match an efficient architecture to the language and the compiler technology available. Measurements of different C programs show that roughly one out of every twenty instructions executed is either a procedure call or return.2 Procedure call overhead is therefore a very important consideration in the overall machine design. A second and related area of primary concern in overall machine efficiency is the register allocation strategy. While use of additional registers can offer considerable improvement in execution times, adding registers usually has the adverse effects of increasing the procedure call overhead due to register saving and creating an undue burden on the compiler. In this paper we describe a piece of the C Machine architecture which effectively eliminates the register allocation problem, and improves procedure calling by drastically reducing storage references required by traditional register saving. The technique can be generalized for other languages and architectures, though we will only directly address those issues involving the C language. David R. Ditzel, Hubert R. McLellan |
ASPLOS | 1 |
| 1980 | Retrospective on High-Level Language Computer ArchitectureabstractHigh-level language computers (HLLC) have attracted interest in the architectural and programming community during the last 15 years; proposals have been made for machines directed towards the execution of various languages such as ALGOL,1,2 APL,3,4,5 BASIC,6,7 COBOL,8,9 FORTRAN,10,ll LISP,12,13 PASCAL,14 PL/I,15,16,17 SNOBOL,18,19 and a host of specialized languages. Though numerous designs have been proposed, only a handful of high-level language computers have actually been implemented.4,7,9,20,21 In examining the goals and successes of high-level language computers, the authors have found that most designs suffer from fundamental problems stemming from a misunderstanding of the issues involved in the design, use, and implementation of cost-effective computer systems. It is the intent of this paper to identify and discuss several issues applicable to high-level language computer architecture, to provide a more concrete definition of high-level language computers, and to suggest a direction for high-level language computer architectures of the future. David R. Ditzel, David A. Patterson 0001 |
ISCA | 1 |